03 Oct
|
F. Hoffmann-La Roche
|
Mississauga
03 Oct
F. Hoffmann-La Roche
Mississauga
At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally.
The Position Kubernetes Reliability
Engineer A healthier future. This role combines software and systems engineering to optimize systems, increase efficiency, and eliminate operational work through automation.
The Opportunity: You will be part of the global CaaS infrastructure team at a leading healthcare company, working with members across different regions. The team's mandate is to deliver, maintain, and continuously improve a highly available Kubernetes platform across hybrid cloud deployments, including on-premise data centers and public clouds like AWS. In this role, you will apply software engineering principles to operations to build and run massively distributed, fault-tolerant systems, focusing heavily on automation, security, and observability.
Scope: Engages in and improves autonomously the whole lifecycle of platforms and services—from inception and design through deployment, operation, and retirement. Applies software engineering principles to build and manage large-scale IT infrastructure products, abstracting away complexity by providing self-service tools and APIs for developers. Designs, implements, and maintains CI/CD pipelines and develops self-healing features.
Focus on capacity planning and launch reviews for services before they go live. Evaluates promising solutions via Proof of Concept (PoCs) and feasibility studies across multiple areas, and serves as an internal escalation point for major incidents Stakeholder Management: Acts as a bridge between engineering and operations. Communicates and presents complex information and potential solutions to cross-functional teams and the business in non-technical terms.
Represents the organization as a prime contact on initiatives and interacts with senior internal and external personnel. Mentors and shares DevOps culture, guiding developers on how to create and deploy cloud-native applications Impact/Strategy: Provides technical leadership and direction for small-to-medium sized initiatives (projects, lifecycle work, PoCs). Ensures solutions comply with Quality/Regulatory standards and that designs adhere to the organization’s Technical Architecture Framework (TAF) policies and directions.
Assists in planning technology projects, estimating engineering resources, dependencies, risks and timelines for successful delivery Complexity: Business/Technical ability: Applies extensive cloud native technical expertise, acting as a recognized expert in Kubernetes and maintaining in-depth knowledge across related cloud native technologies (containers, AWS, etc.). Bachelor’s degree in Computer Science, Mathematics, Physics or related field, and 2-5 years of relevant experience. Master’s degree: Technical Skills Scripting & Software Engineering: Proficiency in scripting and programming languages, primarily Python, Bash, or Go, including experience with test automation (e.g., pytest) and APIs deployment and management.
CI/CD Tools: Expert knowledge of implementing software delivery pipelines using tools (e.g., Systems & Networking: Robust understanding of Linux operating systems and core networking principles, including DNS, load balancing, firewalls, routing, and service meshes.
Experience configuring logging, metrics, and monitoring tools, specifically focusing on setting up alerts based on symptoms rather than waiting for system outages.
Cloud Infrastructure: Experience with public cloud platforms, with a strong preference for AWS, specifically involving managed services for compute, networking, security, and identity (e.g., You have a proven experience applying best practices in an always-up, always-available service environment utilizing Scrum and Agile methodologies You demonstrate a deep understanding of Technical Architecture Frameworks (TAF) and Quality/Regulatory compliance standards You demonstrated a strong team-oriented mindset with the ability to function independently with low supervision and navigate ambiguity. You are highly fluent in oral and written English communication skills are required. You demonstrate strong customer & delivery focus with the ability to act as an analyst, seamlessly transforming complex stakeholder needs into actionable technical requirements.
You possess strong practice of sustainable incident response, including managing ITSM processes and leading audit evidence collection. Relocation benefits are not available for this position. We use artificial intelligence to screen, assess or select applicants for this role.
Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.
Roche is an Equal Opportunity Employer. We believe it’s urgent to deliver medical solutions right now – even as we develop innovations for the future. We commit ourselves to scientific rigor, unassailable ethics, and access to medical innovations for all.
We are Roche. #
📌 System Reliability Engineer - F/H (Mississauga)
🏢 F. Hoffmann-La Roche
📍 Mississauga