Cohere is seeking a Site Reliability Engineer to enhance AI platform performance. Contribute to deploying robust machine learning systems in a remote-friendly setting.
As part of the Model Serving team at Cohere, you will develop and operate a platform for large language model API endpoints. This role requires collaborating with various teams to ensure efficient deployment and optimized performance in high-availability environments. You’ll engage with customers to tailor specific deployments that meet their needs, marking a significant impact in AI technology.
Key Responsibilities:
• Build self-service systems for managing services
• Automate Kubernetes deployments for language models
• Ensure environment observability and resilience
• Participate in on-call rotation for SLO adherence
• Foster relationships with internal teams for feedback integration
Requirements:
• 5+ years in production infrastructure engineering
• Experience with Kubernetes and large distributed systems
• Background in GCP, Azure, AWS, or similar
• Excellent troubleshooting and collaboration skills
• Familiarity with GPUs and distributed system technologies
Drive infrastructure excellence and build impactful AI systems as a Site Reliability Engineer at Cohere.
#J-18808-Ljbffr
📌 Site Reliability Engineer at Cohere (Toronto)
🏢 Visa Hunt
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.