07 Oct
|
Cerebras Systems
|
Toronto
07 Oct
Cerebras Systems
Toronto
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.
As a Senior Research Engineer on the Inference ML team atCerebrasSystems, you will adapt today's most advanced language and vision models to run efficiently on our flagshipCerebrasarchitecture.You'llwork alongside ML researchers and engineers to design, prototype,validate, andoptimizemodels, gaining end-to-end exposure tocutting-edgeinference research on the world's fastest AI accelerator. You will focus on pushing the frontier of speculative decoding , large-model pruning and compression , sparse attention , and sparsity-driven techniques to deliver low-latency, high-throughput inference at scale. Design, implement, andoptimizestate-of-the-arttransformer architectures for NLP and computer vision onCerebrashardware.
Research and prototype novel inference algorithms and model architectures that exploit the unique capabilities ofCerebrashardware, with emphasis on speculative decoding, pruning/compression, sparse attention, and sparsity . Develop diagnostic tooling or scripts to surface performance bottlenecks and guide optimization strategies for inference workloads. Collaborate across teams, including software, hardware, andproduct, to drive projects frominceptionthrough delivery.
Bachelor’s degree in Computer Science, Software Engineering, Computer Engineering, Electrical Engineering,
or a related technical field AND 7+ years of ML software development experience, OR Master’s degree in Computer Scienceor related technical field AND 4+ years of software development experience, OR PhD in Computer Science or related technical field with 2+ years of relevant research or industry experience, OR 4+ years of experience testing,maintaining, or launching software products, including 2+ years of experience with software design and architecture. ~3+ years of experience in software development focused on machine learning (e.g., deep learning, large language models, or computer vision). ~ Strong programming skills in Python and/or C++. ~ Experience with Generative AI and Machine Learning systems. ~ Evidence of research impact in machine learning, such as publications at top conferences (NeurIPS, ICLR, ICML, ACL, EMNLP, MLSys) or comparable contributions to widely used open-source projects or high-quality preprints. Master’s degree or PhD in Computer Science, Computer Engineering, or a related technical field.
Experience independently driving complex ML or inference projects from prototype to production-quality implementations. Hands-on experience with relevant ML frameworks such as PyTorch, Transformers,vLLM, orSGLang .
Experience with large language models, mixture-of-experts models, multimodal learning, or AI agents.
Experience with speculative decoding , neural network pruning and compression , sparse attention , quantization ,
sparsity , post-training techniques, and inference-focused evaluations. Familiarity with large-scale model training and deployment, including performance and cost trade-offs in production systems. Proficiencywith at least one major ML framework ( PyTorch, Transformers,vLLM, orSGLang ).
Deep understanding of transformer-based models in language and/or vision domains, withdemonstratedexperience implementing andoptimizingthem. Proven ability to translate research ideas into robust code: implementing new model variants, training strategies, and evaluation workflows end-to-end. Strong foundationin performance optimization on specialized hardware (e.g., Deep understanding of modern ML architectures and strong intuition foroptimizingtheir performance, particularly for inference workloads using sparse attention, pruning/compression, and speculative decoding .
Collaborative approach with humility, eagerness to help colleagues, and commitment to team success. Genuine passion for AI and a drive to push the limits of inference performance. People who are serious about software make their own hardware.
At Cerebras, we have built a breakthrough architecture that is unlocking recent opportunities for the AI industry. Build a breakthrough AI platform beyond the constraints of the GPU. # Publish and open source their cutting-edge AI research. # Work on one of the fastest AI supercomputers in the world. # Our simple, non-corporate work culture that respects individual beliefs. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.
This website or its third-party tools process personal data.
📌 Senior Research Engineer - Inference ML (Toronto)
🏢 Cerebras Systems
📍 Toronto