13 Sep
|
OpenTable
|
Toronto
Who you are
- 7+ years of professional software engineering experience, with a substantial portion spent building and operating machine learning systems in production
- Breadth of industry perspective. You have seen how ML systems are built and operated at more than one organization, and can speak to industry-standard practices, common reference architectures, and where the real tradeoffs lie
- Hands-on experience with a major cloud platform (AWS, GCP, or Azure) as a primary model-serving environment, including its managed services for deployment, scaling, and observability
- Deep engineering fundamentals: distributed systems, service and API design, concurrency, latency and throughput tradeoffs, testing discipline, and genuine production ownership including on-call
- Strong command of Python and proficiency in at least one strongly typed language (Java preferred)
- Demonstrated experience training, serving, and deploying ML models at production scale
- Production MLOps ownership: model and feature monitoring, drift and data-quality detection, retraining and promotion workflows, versioning, safe rollout and rollback, and incident response when a model misbehaves
- A track record of technical leadership: leading multi-quarter projects, influencing engineering decisions beyond your immediate team, and coordinating with Product Managers and other stakeholders
- Serving LLMs in production: inference infrastructure, GPU utilization, batching and caching strategies, quantization, and managing the latency/cost frontier (vLLM, TGI, TensorRT-LLM, or similar)
- Applied ML depth in ranking, recommendations, classification, NLP, RAG, and/or agentic systems
- Kubernetes in production at meaningful scale
- Experience developing ETL jobs (especially Spark) or data warehouse infrastructure
- Familiarity with A/B testing design and analysis best practices
- Experience introducing current tooling or platform capability to a team and driving adoption
- We do not expect experience with everything on this list - it is here so you know what you would be working with:
- Pipelines: Spark, Airflow, EMR, SageMaker, Snowflake, S3, Delta Lake
- ML: PyTorch, XGBoost / CatBoost, LLMs, LangChain, LangSmith
- Deployment: Docker, Kubernetes, Helm, Prometheus, Graphite / Grafana
- Infrastructure: Kafka, ElasticSearch, Postgres, MongoDB, Redis, Qdrant
- Build: Poetry, FastAPI, Flask, Gunicorn / Uvicorn, Spring, Maven, TeamCity
What the job involves
- The Data Science team at OpenTable supports a wide range of initiatives targeting diners, restaurants, and internal stakeholders. The team is expanding its capabilities across multiple areas, including building AI Agents to power restaurant search and discovery as well as AI-augmented products for restaurant management
- As a Staff Machine Learning Engineer, you will set the technical direction for how OpenTable builds, serves, and operates machine learning systems in production. You will partner with Machine Learning Scientists and engineers across the company to take models from experimentation to reliable, monitored production services --- and you will define the standards and patterns the rest of the team builds on
- This is a deliberately engineering-forward role. We are looking for an engineer who builds and operates production systems,
rather than a modeller who deploys occasionally. The strongest candidates will bring hard-won judgment from more than one organization about how mature ML teams actually work, and the ability to apply it here. This posting is for an existing vacancy
- Technical direction for how models are served, deployed, and monitored; the architecture, the patterns, and the tradeoffs behind them
- Production ML services that are high-throughput, low-latency, and observable, from design through operation
- Engineering standards for ML systems: testing, CI/CD, observability, alerting, rollback, and on-call practice
- Ambiguous, cross-team problems: scoping them with Product Managers and stakeholders, then ruthlessly prioritizing what the team actually builds
- Key Initiatives:
- AI Agents for restaurant discovery
- Personalized recommendations for diners
- Developing and serving high-throughput predictive models for strategic marketplace optimization initiatives
- Building and integrating tools into our agentic platform via LLM tool calls, MCP, and Agent-to-Agent protocols
- Multimodal understanding of restaurant content (text, images, geospatial)
- Creating an AI-powered platform for restaurant partners to gain insights into their business performance and diner demand
Benefits
- Work from (almost) anywhere - wherever you do your best work
- Mental health and well-being - company-paid therapy sessions through SpringHealth, company-paid subscription to HeadSpace, and company-wide weeks off a year so the whole team can recharge
- Generous parental leave
- Generous paid vacation + time off for your birthday
- Paid volunteer time
- Enriched learning and development opportunities - leadership development & access to thousands of on-demand e-learnings
📌 Staff Machine Learning Engineer (Toronto)
🏢 OpenTable
📍 Toronto