Cerebras Systems seeks a Senior Site Reliability Engineer to enhance the performance of AI inference services. Focus on architecting solutions, mentoring, and driving automation in high-tech environments.
As a Senior SRE, you will lead efforts to refine and optimize operational workflows for a groundbreaking AI platform. Collaborating with engineers, product teams, and external stakeholders, you will help establish robust self-service infrastructure and reliability practices. This role emphasizes eliminating toil and deploying cutting-edge solutions across multi-datacenter operations.
Key Responsibilities:
• Define and implement software delivery and reliability strategies • Architect self-service tools for product teams and operators • Evolve reliability practices for inference workloads • Mentor mid-level SREs and support incident escalations • Measure impact through deployment velocity and SLO compliance
Requirements: • 8+ years in SRE or infrastructure engineering • Expertise in operating large scale clusters • Proven experience with CI/CD or GitOps systems • Hands-on with observability tools like Prometheus • Ability to lead complex projects and communicate effectively
Lead the transformation of AI infrastructure reliability with Cerebras, ensuring top-tier service and innovation. #J-18808-Ljbffr
📌 Senior Site Reliability Engineer at Cerebras (Lower Sackville)
🏢 Cerebras Systems
📍 Lower Sackville
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.