12 Aug
|
Cerebras Systems
|
Ahuntsic
12 Aug
Cerebras Systems
Ahuntsic
Join Cerebras Systems as the AI Inference Systems Reliability Lead to build the most reliable AI service in the industry. Drive creative strategies and implement critical reliability systems in a cooperative environment.
This hands-on role requires a seasoned engineer with over 7 years in reliability engineering for large-scale distributed systems. You will be responsible for defining service-level objectives, leading incident management, and developing advanced reliability mechanisms. If you have strong programming skills in popular languages and a deep understanding of SLOs, this position is for you.
Key Responsibilities
Define reliability goals and align engineering efforts
Design and develop fault detection and recovery systems
Lead incident management and root-cause analysis
Architect reliable systems with observability features
Create chaos testing and load simulation tools
Requirements
Bachelor's or master's degree in a relevant field
7+ years of experience in distributed systems
Robust programming experience in Python, C++, or Go
Knowledge of reliability architectures and practices
Excellent communication and leadership abilities
Join Cerebras to leverage your expertise in ensuring reliable AI systems and driving significant technological advancements.
📌 Ai Inference Systems Reliability Lead Ahuntsic
🏢 Cerebras Systems
📍 Ahuntsic