Join Cohere as a Site Reliability Engineer and shape the future of AI systems. Work on high-performance infrastructure in distributed environments in a team-oriented manner.Cohere's Model Serving team is on the lookout for an experienced individual who thrives in building scalable AI platforms. This role emphasizes the deployment and optimization of cutting-edge language models, providing efficient API services. Your interaction with customers will be key to customizing deployments that suit diverse needs while ensuring sustainability and high performance.Key Responsibilities:Develop systems to automate service managementManage Kubernetes deployment operations effectivelyEnhance observability within various environmentsParticipate actively in the on-call duty rosterCollaborate with internal developers to improve infrastructureRequirements:Minimum of 5 years in production systems engineeringProficiency in Kubernetes and multi-cloud solutionsSolid experience with Linux-based environmentsStrong troubleshooting and teamwork skillsKnowledge of performance characteristics of acceleratorsJoin an innovative team committed to building advanced AI technologies with Cohere.#J-18808-Ljbffr
📌 Experienced Site Reliability Engineer Role (Toronto)
🏢 Visa Hunt
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.