Join Cerebras Systems as the AI Inference Systems Reliability Lead to build the most reliable AI service in the industry. Drive innovative strategies and implement critical reliability systems in a collaborative setting.
This hands-on role requires a seasoned engineer with over 7 years in reliability engineering for large-scale distributed systems. You will be responsible for defining service-level objectives, leading incident management, and developing advanced reliability mechanisms. If you have strong programming skills in popular languages and a deep understanding of SLOs, this position is for you.
Key Responsibilities:
• Define reliability goals and align engineering efforts
• Design and develop fault detection and recovery systems
• Lead incident management and root-cause analysis
• Architect reliable systems with observability features
• Create chaos testing and load simulation tools
Requirements:
• Bachelor's or master's degree in a relevant field
• 7+ years of experience in distributed systems
• Strong programming experience in Python, C++, or Go
• Knowledge of reliability architectures and practices
• Excellent communication and leadership abilities
Join Cerebras to leverage your expertise in ensuring reliable AI systems and driving significant technological advancements.
#J-18808-Ljbffr
📌 AI Inference Systems Reliability Lead (Ottawa)
🏢 Cerebras Systems
📍 Ottawa
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.