04 Sep
|
Cerebras Systems
|
Lower Sackville
04 Sep
Cerebras Systems
Lower Sackville
Cerebras Systems seeks a Senior Site Reliability Engineer to enhance the performance of AI inference services. Focus on architecting solutions, mentoring, and driving automation in high-tech environments. As a Senior SRE, you will lead efforts to refine and optimize operational workflows for a groundbreaking AI platform.
Collaborating with engineers, product teams, and external stakeholders, you will help establish robust self-service infrastructure and reliability practices. This role emphasizes eliminating toil and deploying creative solutions across multi-datacenter operations. Key Responsibilities:
- Define and implement software delivery and reliability strategies
- Architect self-service tools for product teams and operators
- Evolve reliability practices for inference workloads
- Mentor mid-level SREs and support incident escalations
- Measure impact through deployment velocity and SLO compliance
Requirements:
- 8+ years in SRE or infrastructure engineering
- Expertise in operating large scale clusters
- Proven experience with CI/CD or GitOps systems
- Hands-on with observability tools like Prometheus
- Ability to lead complex projects and communicate effectively
Lead the transformation of AI infrastructure reliability with Cerebras, ensuring top-tier service and innovation.
📌 Senior Site Reliability Engineer at Cerebras (Lower Sackville)
🏢 Cerebras Systems
📍 Lower Sackville