We are looking for senior Site Reliability Engineer with hands on experience in design, analysis, development and troubleshooting of highly-distributed large-scale production systems.
- Site Reliability Engineer, Observability Engineer - Work with cloud-native open-source tools to improve the visibility and reliability of self-hosted Kubernetes platform.
- Robust automation skills and enterprise-grade systems thinking
- Someone with prior experience in developing, debugging, and deploying enterprise applications, including AI/ML-powered services.
- Ownership of reliability, uptime, system security, cost, operations, capacity, resiliency and performance-analysis thereof
- Define, monitor and report on service level indicators for applications workloads
- Support on-call rotations for operational duties that have not been addressed with automation, with an eye for correcting issues that result in on-call alarms
- Candidate should have experience in working on developing AI systems or AI models.
- Candidate must have experience in Kubernetes, Terraform and managing AI infrastructure.