Explore a high-performance role at Cerebras as an AI Reliability Engineer, specializing in automation and operational execution. This position prioritizes the delivery of industry-leading AI infrastructures.Cerebras Systems is looking for an enthusiastic SRE who will manage and optimize operations for one of the fastest AI inference services globally. You’ll take ownership of production systems and learn directly from experienced Staff SREs. This role involves building self-service capabilities and enhancing the reliability of operations through comprehensive telemetry.Key Responsibilities:
- Manage operational tasks, including cluster upgrades and capacity changes
- Develop self-service CD pipelines using Kubernetes and other tools
- Automate workflows to minimize operational effort
- Extend observability solutions for large-scale operations
- Collaborate with development teams on reliability practicesRequirements:
- 2-4+ years of SRE experience with an operations focus
- Experience with Kubernetes and cloud technologies
- Proficient in Python or Go for tool development
- Familiarity with Prometheus and observability workflows
- Experience in GitOps or building delivery pipelines highly valuedBecome a key player in driving AI infrastructure reliability at Cerebras.#J-18808-Ljbffr
📌 Ai Reliability Engineer At Cerebras (Ottawa)
🏢 Cerebras Systems
📍 Ottawa
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.