Cohere invites applications for a Site Reliability Engineer position, focused on deploying advanced AI solutions. Embrace the challenge of high-availability environments while working remotely.
In this role with the Model Serving team, you will play a crucial part in developing automated systems that facilitate the deployment of large language models. Your collaboration with diverse teams will drive improvements in service reliability and performance, ensuring that cutting-edge AI applications are delivered effectively. Create creative solutions tailored to client needs through effective API endpoint management.
Key Responsibilities:
• Automate and manage service operations • Support Kubernetes infrastructure for language models • Drive resilience and observability in systems • Commit to SLOs with regular on-call rotation • Build cross-team relationships to enhance development
Requirements: • Over 5 years of experience in scalable infrastructure • Expertise in Kubernetes and distributed engineering • Familiarity with major cloud platforms (GCP, AWS, etc.) • Excellent collaboration and problem-solving capabilities • Understanding of resource management and accelerators
Be a part of Cohere’s mission to transform AI experiences for enterprises. #J-18808-Ljbffr
📌 Cohere Site Reliability Engineer Opportunity (Toronto)
🏢 Visa Hunt
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.