LocationIn Canada, Mistplay follows a 2 days/week in-office hybrid model in Toronto (400 University Ave) & Montreal (1001 Blvd. Robert-Bourassa).ResponsibilitiesDesign, build, and operate machine and data infrastructure solutions for training models.Build real‑time inference systems to operate and serve models in a real‑time production environment.Develop high usability and accuracy feature platform capabilities for generating, backfilling and storing user‑level features.Create high‑accuracy low‑latency feature serving layers and preprocessing solutions to support online serving of models.Build platform abstractions and golden paths: Airflow DAG templates, CI/CD pipelines, CLI/SDKs, cookie‑cutter repositories that shepherd models from notebooks to production predictably.Implement end‑to‑end observability: data and feature freshness checks, drift/quality gates, model performance/latency SLOs, infrastructure health dashboards, tracing and alerting, incident response and post‑mortems.Partner with Security, SRE, and Data Engineering on private networking, policy‑as‑code, PII handling, least‑privilege IAM, and cost‑efficient architectures across environments.Evaluate, integrate, and rationalize platform tooling (e.G., MLflow registry, feature stores, serving gateways) and lead migrations with explicit change management and minimal downtime.Qualifications10+ years building and operating production‑grade ML/data platforms, focused on serving, reliability, and developer experience.Strong software engineering skills in Python, Go,
or Java; experience building resilient services, APIs, and automation tooling with high test coverage.Deep experience with inference solutions: endpoint configuration, containerization, model packaging, autoscaling, serverless vs. real‑time trade‑offs, MME, A/B and canary releases.Expertise in online feature store paradigms and underlying storage solutions for ML serving.Proven Terraform experience managing ML and data infrastructure end‑to‑end: modules, workspaces, drift detection, change reviews, and safe rollbacks; familiarity with GitOpspatterns.Airflow orchestration at scale: dependency modeling, sensors, retries, SLAs, backfills, DAG factories, integrations with registries, artifact stores, and Terraform pipelines.Familiarity with ML frameworks (scikit‑learn, XGBoost, PyTorch, TensorFlow) from a platform‑integration perspective.Observability for ML workflows: metrics, logs, traces, performance profiling, capacity planning, cost monitoring, and runbooks.Excellent communication and cross‑functional collaboration with Data Science, Data Engineering, DevOps, and Backend.BenefitsWe offer a range of perks including team lunches, game nights, company‑wide events, and more. Our culture is focused on growth, learning, and fostering an environment where everyone is encouraged to share ideas, push boundaries, and see their visions come to life.EEO StatementNous remercions tous les candidats. Le genre masculin a été utilisé dans le but d’alléger le texte. Nous souscrivons au principe de l’équité en matière d’emploi.#J-18808-Ljbffr
📌 Principal Platform Engineer, Ml (Toronto)
🏢 ODAIA
📍 Toronto