Join Cerebras Systems as the AI Inference Systems Reliability Lead to build the most reliable AI service in the industry. Drive innovative strategies and implement critical reliability systems in a collaborative environment.
This hands-on role requires a seasoned engineer with over 7 years in reliability engineering for large-scale distributed systems. You will be responsible for defining service-level objectives, leading incident management, and developing advanced reliability mechanisms. If you have strong programming skills in popular languages and a deep understanding of SLOs, this position is for you.
Key Responsibilities
- Define reliability goals and align engineering efforts
- Design and develop fault detection and recovery systems
- Lead incident management and root-cause analysis
- Architect reliable systems with observability features
- Create chaos testing and load simulation tools
Requirements
- Bachelor's or master's degree in a relevant field
- 7+ years of experience in distributed systems
- Robust programming experience in Python, C++, or Go
- Knowledge of reliability architectures and practices
- Excellent communication and leadership abilities
Join Cerebras to leverage your expertise in ensuring reliable AI systems and driving significant technological advancements.
📌 AI Inference Systems Reliability Lead (Ahuntsic)
🏢 Cerebras Systems
📍 Ahuntsic
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.