03 Sep
|
Cerebras Systems
|
Nova Scotia
03 Sep
Cerebras Systems
Nova Scotia
Cerebras Systems seeks a Senior Site Reliability Engineer to enhance the performance of AI inference services. Focus on architecting solutions, mentoring, and driving automation in high-tech environments.
As a Senior SRE, you will lead efforts to refine and optimize operational workflows for a groundbreaking AI platform. Collaborating with engineers, product teams, and external stakeholders, you will help establish robust self-service infrastructure and reliability practices. This role emphasizes eliminating toil and deploying creative solutions across multi-datacenter operations.
Key Responsibilities:
• Define and implement software delivery and reliability strategies
• Architect self-service tools for product teams and operators
• Evolve reliability practices for inference workloads
• Mentor mid-level SREs and support incident escalations
• Measure impact through deployment velocity and SLO compliance
Requirements:
• 8+ years in SRE or infrastructure engineering
• Expertise in operating large scale clusters
• Proven experience with CI/CD or GitOps systems
• Hands-on with observability tools like Prometheus
• Ability to lead complex projects and communicate effectively
Lead the transformation of AI infrastructure reliability with Cerebras, ensuring top-tier service and innovation.
#J-18808-Ljbffr
📌 Senior Site Reliability Engineer at Cerebras (Nova Scotia)
🏢 Cerebras Systems
📍 Nova Scotia