27 Aug
|
Thomson Reuters
|
Winnipeg
27 Aug
Thomson Reuters
Winnipeg
About The RoleLead Inference Platform Engineer – specialized experience in machine learning/deep learning domains such as model compression, hardware‑aware model optimizations, hardware accelerators architecture, GPU/ASIC architecture, ML compilers, high‑performance computing, performance optimizations, numerics, or SW/HW co‑design.ResponsibilitiesOptimize LLMs and ML models for high‑performance inference using quantization, pruning, distillation, and hardware specific tuning.Deploy and scale inference workloads on GPUs across AWS, Azure, GCP and internal Kubernetes clusters, ensuring predictable performance during peak traffic.Implement routing and fail‑over strategies for OpenAI / Anthropic / Vertex AI traffic.Integrate models into production‑grade APIs supporting TR products and enterprise workflows.Develop highly optimized environments and eliminate performance bottlenecks to reduce latency.Collaborate with Platform Engineering teams (Landing Zones, Network, Storage, Compute, AI) to ensure inference workloads align with cloud‑native patterns.Build and optimize containerized inference pipelines using Kubernetes for large‑scale distributed workloads.Ensure compliance with TR’s AI standards for deployment, monitoring, governance, and drift detection.Profile inference performance, identify GPU/CPU bottlenecks, and optimize compute utilization across heterogeneous hardware.Implement observability and health monitoring for inference pipelines, ensuring reliability of enterprise AI services.Collaborate with platform teams to enhance capacity forecasting for AI workloads.Work with Product, Data Science, Architecture, and Enterprise AI teams to onboard new research models into production.Collaborate closely with AI engineers to invent new quantization techniques, improve numerical precision, and explore non‑standard architectures.Partner with Cloud Engineers (Azure, AWS, GCP)
to develop guardrails and automation that support inference workloads.Support the scale‑out of AI infrastructure during critical releases and global product rollouts.Required Skills & QualificationsStrong understanding of ML/LLM fundamentals and inference optimization techniques.Hands‑on experience with GPU programming (CUDA preferred), inference runtimes (TensorRT, ONNX Runtime), and deep learning frameworks (PyTorch / TensorFlow).Proficiency in Python and at least one systems language (C++ strongly preferred for performance‑critical inference paths).Experience deploying AI workloads to AWS, GCP, Azure and Kubernetes.Familiarity with vector search systems (OpenSearch vectors) and retrieval‑augmented generation pipelines.Knowledge of distributed systems, microservices, CI/CD, and cloud‑native architecture.Experience with AI networks such as CNNs, transformers, and diffusion model architectures and their performance characteristics.Understanding of GPU, multithreading, and/or other accelerators with vectorized instructions.Specialized experience in one or more of the following domains: Model compression, hardware‑aware model optimizations, hardware accelerators architecture, GPU/ASIC architecture, ML compilers, high‑performance computing, performance optimizations,
numerics and SW/HW co‑design.Preferred Qualifications3+ years production experience deploying ML/LLM models at scale.Experience managing GPU fleets or inference clusters across public cloud and container platforms.Experience supporting enterprise‑grade AI workloads in regulated or compliance‑heavy environments.What’s in it For You?Hybrid Work Model: Flexibility to work 2–3 days a week in the office or from anywhere.Flexibility & Work‑Life Balance: Policies to manage personal and professional responsibilities including up to 8 weeks away from the office per year.Career Development and Growth: Continuous learning, skill development, and leadership opportunities.Industry Competitive Advantages: Flexible vacation, two company‑wide mental‑health days, Headspace app access, retirement savings, tuition reimbursement, employee incentive programs, and support for mental, physical, and financial wellbeing.Culture: Inclusion, belonging, flexibility, and a strong corporate purpose.Social Impact: Paid volunteer days and ESG initiatives.Real‑World Impact: Contributing to justice, truth, and transparency worldwide.CompensationBase compensation range (Ontario, Canada): $140,000 CAD – $175,000 CAD. Base pay is determined by knowledge, skills, and experience, among other factors. Base pay is one part of a comprehensive Total Reward program which also includes flexible benefits and wellbeing programs. The role may also be eligible for an annual bonus based on performance.Equal Employment OpportunityThomson Reuters is an Equal Employment Opportunity Employer. We provide reasonable accommodations for qualified individuals with disabilities and sincerely held religious beliefs. The company welcomes employees regardless of race, color, sex/gender, pregnancy, gender identity and expression, national origin, religion, sexual orientation, disability, age, marital status, citizen status, veteran status, or any other protected classification under applicable law.#J-18808-Ljbffr
📌 Lead Inference Platform Support Engineer - Ai I - C$140,000 - C$175,000 A Year (Winnipeg)
🏢 Thomson Reuters
📍 Winnipeg