Who you are
- No previous ML infrastructure experience is required for this role
- Availability for meetings and impromptu communication during Quora's "coordination hours" (Mon-Fri: 9am-3pm Pacific Time)
- A 2025 or 2026 graduate with or pursuing a B.S., M.S., or Ph.D. in Computer Science, Engineering, or a related technical field
- Genuine interest in large-scale distributed systems, infrastructure, and machine learning
- Knowledge of Python, Go, or C++, or the ability to learn them quickly
- A passion for learning and always improving yourself and the team around you
- Previous software engineering experience via an internship, work experience, open-source contribution, or coding competition
- Coursework or hands-on experience with ML frameworks such as PyTorch or TensorFlow
- Exposure to Kubernetes, Docker, or cloud technologies like AWS
- Experience with low-level performance work of any kind: profiling, benchmarking, optimization
- Passion for Quora's mission and goals
What the job involves
- Machine Learning is central to Quora's mission of growing the world's collective intelligence. We have 100+ Machine Learning models in production powering various product features. We use a variety of algorithms — everything from linear models to decision trees and deep neural networks. Our production models operate at a huge scale, serving hundreds of millions of people using Quora every month
- Our team owns Quora's ML platform and ranking infrastructure across four areas: serving reliability, ML engineer enablement and developer velocity, business impact, and cost efficiency.
We want to empower all ML engineers at Quora to be as impactful as they can be in solving different ML problems at scale
- As a Software Engineer (New Grad) on this team, you'll work at the intersection of machine learning, distributed systems, and GPU serving performance — and your work will have an enormous impact on Quora's long-term success
- You'll be joining a team of senior and staff engineers, learning this stack from the people who built it, with a dedicated mentor and solid technical guidance — and you'll be shipping to production in your first few weeks
- Stack: Python, Go, C++, PyTorch, Kubernetes/EKS, NVIDIA Triton, Ray, AWS
- Help build and maintain the core infrastructure that powers Quora's ML platform, ensuring high availability, scalability, and performance
- Build and improve the distributed systems that serve our ML models in production, from Large Recommendation Models (LRMs) to Large Language Models (LLMs)
- Work on GPU model serving, optimizing latency, throughput, and cost to support larger and more capable models
- Contribute to platform initiatives such as PyTorch-first standardization and ML ecosystem modernization
- Improve ML developer velocity by building tooling that helps ML engineers develop, test, and deploy models more efficiently
- Modernize our feature store so ML engineers can get new features into production faster
- Participate in the team's on-call rotation, helping resolve production issues as you grow your knowledge and ownership of the platform
📌 Software Engineer (Canada)
🏢 Quora
📍 Canada