ML / AI ENGINEER
ABOUT SKETCHDECK.AI
SketchDeck.ai is transforming construction estimation through applied AI. Our platform uses computer vision and machine learning to turn complex construction drawings into structured information that estimators can review, correct, and use in their workflows.
ABOUT THE ROLE
We are looking for a Machine Learning / AI Engineer to design, train, evaluate, deploy, and improve the computer vision systems at the core of our products.
This is an applied engineering role with end-to-end ownership across datasets, experimentation, training, evaluation, inference, deployment, monitoring, and continuous improvement.
Structural drawings contain dense visual information, inconsistent conventions, complex geometry, text, symbols, and varying image quality. Success requires models that perform reliably across customers, drawing styles, and edge cases.
You will work with product, full-stack, platform, and ML engineers to turn model outputs into dependable production features. This role suits someone who enjoys the full applied ML lifecycle and measures success through production outcomes.
WHAT SUCCESS LOOKS LIKE
During your first year, you will help us:
- Improve the accuracy, precision, recall, and robustness of our production computer vision models.
- Build repeatable training and evaluation pipelines.
- Improve dataset quality, labeling workflows, augmentation strategies, and difficult-example management.
- Reduce the gap between offline model performance and customer outcomes.
- Improve inference performance, reliability, GPU utilization, and cost.
- Establish stronger model versioning, deployment, validation, monitoring, and rollback practices.
- Build diagnostics for false positives, false negatives, regressions, and customer-specific failures.
- Expand our use of OCR, geometric reasoning, document understanding, and multimodal AI where they create measurable value.
- Turn validated customer corrections into better training data and future model improvements.
KEY RESPONSIBILITIES
Computer Vision and Deep Learning
- Design, train, evaluate, and improve models for object detection, segmentation, classification, and document understanding.
- Develop production models using PyTorch, YOLO-family models, and other suitable architectures.
- Improve model calibration and generalization across customers and drawing types.
- Determine whether failures originate in data, labeling, preprocessing, training, inference, post-processing, or product workflows.
- Design controlled experiments that provide transparent evidence for technical decisions.
- Evaluate new architectures and foundation models against measurable product outcomes.
Datasets, Training, and Evaluation
- Build pipelines for dataset creation, annotation, validation, preprocessing, augmentation, training, and evaluation.
- Use hard-negative mining, representative sampling, and error analysis to improve training data.
- Establish dataset, experiment, configuration, and model versioning.
- Prevent data leakage, weak dataset splits, annotation inconsistencies, and misleading evaluations.
- Define evaluation frameworks based on product requirements.
- Track precision, recall, F1, confidence calibration, and class-level performance.
- Build regression datasets that represent key customers,
drawing types, edge cases, and known failures.
- Establish release criteria that prevent unacceptable regressions.
Production ML and Inference
- Design and maintain Python-based inference pipelines for batch and asynchronous workloads.
- Optimize inference latency, throughput, GPU utilization, memory use, and cost.
- Package and deploy models using Docker across AWS and GCP.
- Support model versioning, validation, progressive rollout, and rollback.
- Build resilient workflows that handle retries, duplicate work, timeouts, partial failures, and long-running jobs.
- Diagnose issues across model serving, queues, infrastructure, application services, and data pipelines.
ML Observability and Improvement
- Monitor model health, inference performance, failure rates, resource use, and production quality.
- Detect distribution changes and emerging failure patterns.
- Maintain lineage between production outputs, model versions, configurations, and datasets.
- Build feedback loops that turn validated customer corrections into future model improvements.
- Establish disciplined model release, regression testing, comparison, and production-validation practices.
Document AI and Geometric Reasoning
- Apply OpenCV and classical image processing when they provide a simpler or more reliable solution.
- Use OCR and document-understanding techniques to extract information from construction drawings.
- Develop solutions involving coordinate systems, transformations, spatial relationships, and geometry.
- Combine machine learning, deterministic algorithms, and domain rules when hybrid approaches produce better outcomes.
Product and Engineering Collaboration
- Integrate ML capabilities into customer-facing workflows with full-stack engineers.
- Define transparent contracts between inference and application services.
- Work with product and construction experts to understand how errors affect estimation workflows.
- Help design human-in-the-loop experiences for reviewing and correcting predictions.
- Contribute to architecture, planning, code review, documentation, and engineering standards.
- Mentor engineers and strengthen shared ML knowledge across the team.
ABOUT YOU
- You focus on solving customer problems rather than adopting technology for its novelty.
- You understand that production ML depends on data quality, evaluation, monitoring, maintainability, and reliable inference.
- You use evidence to select metrics, evaluate results, and guide model development.
- You can investigate failures across the entire ML and application pipeline.
- You treat datasets, configurations, experiments, artifacts, and evaluation results as reproducible engineering assets.
- You can independently navigate ambiguous problems while collaborating with engineers and domain experts.
- You own work through production and validation.
MUST-HAVE QUALIFICATIONS
- Four or more years of professional experience in applied machine learning, computer vision, deep learning, or a related field.
- Production experience with object detection, segmentation, classification, or document-understanding systems.
- Strong PyTorch and Python experience across model development and maintainable production software.
- Strong knowledge of image preprocessing, augmentation, annotation, evaluation, and error analysis.
- Experience with OpenCV or comparable image-processing libraries.
- Understanding of deep learning architectures, training dynamics, loss functions, optimization, regularization, and evaluation.
- Foundations in linear algebra, probability, statistics, optimization, and geometry.
- Experience building repeatable training and evaluation pipelines.
- Experience deploying and operating ML models in production.
- Experience with Docker, cloud infrastructure, and GPU workloads.
- Experience optimizing inference performance and diagnosing GPU, memory, latency, and throughput constraints.
- Understanding of precision, recall, confidence thresholds, class imbalance, and error tradeoffs.
- Strong written and verbal communication skills.
- Master’s or PhD in a relevant discipline, or equivalent professional experience.
PREFERRED QUALIFICATIONS
- Experience with YOLO-family models or comparable detection architectures.
- Experience with technical drawings, engineering documents, PDFs, maps, or other visually dense material.
- Experience with OCR, layout analysis, document understanding, or multimodal models.
- Experience with geometric reasoning, coordinate transformations, spatial relationships, or 3D computer vision.
- Experience building human-in-the-loop systems and retraining workflows based on production corrections.
- Experience with experiment tracking, model registries, dataset versioning, drift detection, and regression testing.
- Experience optimizing GPU workloads for performance and cost.
- Experience operating ML workloads across AWS or GCP.
- Experience with asynchronous job processing and distributed ML workloads.
- Experience with MongoDB, object storage, and large image or document datasets.
- Familiarity with MLOps, CI/CD, and automated model validation.
- Experience working in a growing company where engineers own work from experimentation through production.
OUR TECHNOLOGY
- Python, PyTorch, YOLO-family models, OpenCV, OCR, and GPU-based training and inference.
- Flask, MongoDB, Amazon DocumentDB, Redis, and large-scale PDF and image processing.
- Docker, Docker Compose, AWS, Google Cloud, Bitbucket Pipelines, Sentry, and Prometheus.
WHY SKETCHDECK.AI?
- Build AI systems that directly improve customer workflows.
- Solve complex computer vision problems using information-dense construction drawings.
- Own models from dataset development through production performance.
- Use customer interaction and corrections to improve future models.
- Influence our ML architecture, evaluation standards, tooling, and technical direction.
- Work with experienced engineering, product, and construction professionals.
- Join a team that values autonomy, accountability, practical engineering, and clear communication.
Ready to build the AI behind the future of construction estimation?
Apply at
[email protected]
📌 Machine Learning Engineer (Toronto)
🏢 SketchDeck.ai
📍 Toronto