Design and build scalable, high throughput, and low latency distributed systems using ScalaBuild reusable components and services that serve various ML applications like Personalization, Search, Ads and ExplorationPartner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art.Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessaryMentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc.Collaborate with ML, Product,
and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi.RequirementsExperience designing and building scalable, distributed systems in any up-to-date backend language (e.G., Scala, Java, Python, Go, C++); experience with Scala or JVM based language is a plus.Solid experience with AWS or an equivalent cloud platformExperience building online microservices at scale with low latency servingExperience with both SQL (e.G. Postgres) and NoSQL databases (e.G. Cassandra), message brokers (e.G. Kafka), and caches (e.G. Redis)Experience with containerization technologies, such as Docker or KubernetesLed the response and resolution efforts for multiple major, large-scale incidents. #J-18808-Ljbffr
📌 Software Engineer, Ml Infra, Distributed Systems – Staff, Principal (Toronto)
🏢 Jobtailor
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.