Your opportunity
Our client builds a high throughput system powering hundreds of billions of daily transactions, each completed within milliseconds, across globally distributed infrastructure designed for reliability and efficiency. They have been quietly bootstrapping and growing in line with revenue for over 2 decades. They maintain independence from external stakeholders, enabling them to chart their course, maintain a long-term perspective and build an enduring, sustainable business that currently employs ~600 team members globally.
Data is central to nearly every part of their platform. High-volume event and transactional data powers customer reporting, billing, marketplace analytics, experimentation, machine-learning workflows, and product experiences. The organization is evolving from centralized, batch-oriented reporting toward a platform-driven architecture combining batch, streaming, and asynchronous processing, with APIs becoming a primary interface to data.
Our client has been stubbornly racking and stacking infrastructure around the world for the duration of their existence, a habit that allows for them to price themselves at the cost of electricity while their competitors are mired in rising cloud infrastructure costs. This long-term, somewhat contrarian, thinking puts them in a position to offer stability and career longevity evidenced by robust advantages that include RRSP matching.
Machine learning is already in production and they have successfully shipped their first generation of ML-powered systems. They’re now moving into the next generation, at a point where growing transaction and data volumes leave considerable room for ML to influence the economics, efficiency and intelligence of the platform.
This is an opportunity to help define what that next generation looks like. You’ll operate across the entire path from raw data through feature generation, modelling, delivery, deployment, inference and production monitoring, helping establish architectures that allow sophisticated ML systems to operate reliably at extraordinary scale.
Because the client’s production ML capability is still maturing, you’ll have considerable influence over the architectures, patterns and technical approaches that underpin its next generation of ML systems. Rather than inheriting a completely settled way of doing things,
you’ll help determine what those systems should look like as the organization continues to expand its use of machine learning.
Key responsibilities
- ML architecture: Design and evolve end-to-end architectures for large-scale machine learning systems spanning data preparation, feature generation, model delivery, deployment, serving and production monitoring
- Data-centric systems: Design solutions for validating, transforming and generating features from extremely large datasets while ensuring consistently high-quality inputs for downstream model development
- Model delivery: Architect systems and patterns that allow models to move reliably from development into production environments
- Deployment & hosting: Develop effective approaches for deploying, hosting and operating ML models at scale, balancing performance, reliability, maintainability and infrastructure constraints
- Production ML: Establish robust approaches for managing and monitoring deployed models and the systems that support them
- Performance: Identify and pursue opportunities to optimize the performance of ML systems across data processing, model execution and serving infrastructure
- Technical design: Lead functional and process design across ML initiatives (including scenario design, flow mapping, prototyping, testing and defining appropriate support procedures)
- Engineering standards: Bring strong software engineering principles into the design of ML systems so that experimentation can translate into maintainable, production-grade software
- Cross-functional collaboration: Work across data science, machine learning, software engineering and other stakeholders to source the right data, establish technical requirements and define what success actually looks like
- Technical leadership: Provide architectural direction across complex ML initiatives and influence how the broader organization approaches machine learning as its use of the technology continues to mature
Your know-how
- You have a Masters or PhD in a STEM field or equivalent depth of experience gained as a Data Scientist, Machine Learning Engineer, ML Architect or in a closely related role
- You understand what it takes to move ML beyond experimentation and operate it reliably in production
- You have extensive knowledge of the infrastructure and pipelines that support production ML systems (aka MLOps)
- You have meaningful experience across several modelling approaches such as classification, regression, pattern recognition, recommendation systems, targeting systems or ranking systems
- You bring deep knowledge of mathematics, probability, statistics and algorithms and can apply that knowledge to practical engineering and product challenges
- You understand how to design the systems surrounding a model, including data preparation, feature generation, delivery, deployment, monitoring and operational support
- You have an interest in performance optimization and can reason about the trade-offs involved in operating ML systems at significant scale
- You understand software engineering best practices and can work effectively alongside advanced engineering teams building production systems
- You have strong business acumen and can work cross-functionally to source data, establish requirements, identify meaningful constraints and define measures of success
- You’re comfortable operating at “principal level” (aka defining architecture, navigating ambiguity, influencing technical direction and making decisions whose consequences extend beyond an individual project or team)
It’s a bonus if
- You have a software engineering background in addition to your machine learning expertise
- You have experience deploying ML models for rapid/real-time inference
- You’ve published articles, research or papers related to machine learning, statistics, optimization or adjacent fields
- You’ve worked in another environment where enormous data volumes, high transaction throughput and tight latency constraints materially change how ML systems need to be designed
Interested in learning more?
Please send your resume or LinkedIn profile URL to
[email protected] with “Principal Machine Learning Architect” as the subject line. One of our talent partners will be in contact shortly!
📌 Principal ML Architect, Open Internet (Canada)
🏢 Lutra
📍 Canada