Sr. Site Reliability Engineer, IE&O (London)

Sr. Site Reliability Engineer, IE&O (London)

03 Aug
|
McCain Foods
|
London

03 Aug

McCain Foods

London

Position Title: Sr. Site Reliability Engineer, IE&O;
Position Type: Regular - Full-Time

Requisition ID: 42659

McCainers are ambitious, curious, and passionate about creating exceptional work experiences together. With a customer-first mindset, we make doing business with McCain easy. Our Global Technology team's goal is to leverage technology and data to drive profitable growth, enhance customer experience, and further our purpose of celebrating real connections through delicious, planet-friendly food.

McCain has embarked on an ambitious digital transformation across our business from Agriculture to Manufacturing and commercial capabilities to enhance our customer obsession. As part of this transformation, we are making significant investments in our digital platforms, technology transformations, and building a data-driven culture.

About The Role

The Site Reliability Engineer will help improve the reliability, availability, performance, and operability of McCain's critical technology platforms. This role will design resilient cloud-native systems, embed observability into applications and infrastructure, automate operational workflows, and help scale SRE practices across engineering and platform teams.

As part of the global SRE practice, this role will also help build and drive McCain's AIOps capabilities by improving telemetry, alert quality, event correlation, incident automation, and proactive reliability insights.

What You'll Be Doing

- Design, build, and improve reliable, scalable, and secure systems across Azure cloud and hybrid environments.
- Embed observability into applications and platforms using metrics, logs, traces, dashboards, alerts, and Open Telemetry standards.
- Build and drive AIOps capabilities by improving alert quality, event correlation, incident enrichment, noise reduction, automated triage, and operational automation.
- Partner with engineering teams to define SLOs, SLIs, Error Budgets, production readiness standards, and reliability scorecards.
- Build automation to reduce toil across infrastructure,



deployments, incident response, monitoring, and operational workflows.
- Use Infrastructure as Code, CI/CD pipelines, scripting, and self-healing patterns to improve reliability and delivery speed.
- Support incident response, root cause analysis, postmortems, escalation workflows, and continuous reliability improvements.
- Troubleshoot complex issues across application, infrastructure, cloud, network, database, and integration layers.
- Build reusable SRE playbooks, standards, templates, and automation patterns for broader enterprise adoption.
- Collaborate with developers, platform teams, operations teams, vendors, and stakeholders to improve system reliability and operational maturity.

What You'll Need To Be Successful

- 9+ years of experience in software engineering, platform engineering, cloud engineering, Dev Ops, production engineering, or site reliability engineering.
- Strong hands‐on experience with Azure, Kubernetes, containers, APIs, distributed systems, and modern deployment patterns.
- Strong scripting or software engineering experience using Python, Go, Power Shell, Bash, Java, or similar languages.
- Experience with observability, including metrics, logs, traces, dashboards, alerts, Open Telemetry, and telemetry-driven reliability practices.
- Experience with Infrastructure as Code, CI/CD, automation, and deployment tooling such as Terraform, Bicep, Git Hub Actions, Azure Dev Ops, or similar technologies.
- Good understanding of SLOs, SLIs, Error Budgets, resiliency patterns, incident management, production readiness, and capacity planning.
- Strong troubleshooting, communication, and stakeholder influencing skills.




- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related technical field. Azure certifications are preferred.
- Experience with AIOps, event correlation, alert enrichment, noise reduction, automated triage, or incident automation.
- Experience using AI-assisted capabilities for incident triage, root cause analysis, knowledge management, operational automation, or engineering productivity.
- Experience building self-service platforms, reusable automation frameworks, golden paths, or internal developer platforms.

About The Team

Reporting to the Sr. Engineering Manager, SRE and Observ, you will be part of a team consisting of Site Reliability Engineers, AI Ops Engineers, and Software Engineers.

Perks

Health coverage (medical, dental, vision, prescription drug), retirement savings benefits, and leave support including medical, family, and bereavement. Well‐being programs include vacation and holidays, company-supported volunteering time, and mental health resources.

Compensation Package

$102,700.00 - $137,000.00 CAD annually + Bonus Eligibility

Job Family

Information Technology

Location(s)

CA - Canada: Ontario : Toronto

Company

McCain Foods (Canada)

Equal Opportunity Statement

McCain Foods is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, religion, color, national origin, sex, age, veteran status, disability, or any other protected characteristic under applicable law.

Accessibility Statement

If you require an accommodation throughout the recruitment process, please let us know, and we will work with you to find appropriate solutions.

Privacy Statement

By submitting personal data or information to McCain, you agree it will be handled in accordance with McCain's Global Privacy Policy and Global Employee Privacy Policy. McCain leverages AI in the hiring process, though all final decisions are made by humans.

#J-18808-Ljbffr

📌 Sr. Site Reliability Engineer, IE&O (London)
🏢 McCain Foods
📍 London

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sr. site reliability engineer, ie&o (london) / london