25 Sep
|
SoTalent
|
Canada
Site Reliability Engineer, Infrastructure Platforms
? Location: Canada (Remote)
? Industry: IT Services and IT Consulting
? Work Setting: Remote
Are you passionate about building highly reliable, scalable, and automated infrastructure that powers mission-critical applications and services? We are seeking a Site Reliability Engineer (SRE) to help design, operate, and continuously improve large-scale production environments. In this role, you will combine software engineering expertise with operational excellence to enhance system reliability, reduce operational toil, and drive infrastructure automation across cloud-native platforms.
Key Responsibilities
- Design, build, and maintain reliable, scalable, and highly available production systems and services.
- Develop automation and tooling that eliminate manual processes and improve operational efficiency.
- Manage, deploy, and troubleshoot containerized workloads within Kubernetes environments.
- Build and maintain Infrastructure as Code (IaC) solutions to support consistent, repeatable, and secure infrastructure deployments.
- Implement and optimize CI/CD and GitOps workflows to enable safe, automated software delivery.
- Participate in on-call rotations, monitor production environments, respond to alerts, and resolve service incidents.
- Improve observability through metrics, logging, monitoring, alerting, and service-level objectives (SLOs).
- Lead and contribute to incident response activities, root cause analysis, post-incident reviews, and preventative improvements.
- Document architecture decisions, operational procedures, troubleshooting guides, and best practices.
- Collaborate with engineering teams to improve service reliability, performance, scalability, and operational readiness.
- Drive continuous improvement initiatives focused on resilience, automation, efficiency, and platform stability.
Required Qualifications
- Experience supporting and maintaining large-scale production systems with a focus on reliability, performance, and operational excellence.
- Strong software engineering background with the ability to read, analyze, debug, and troubleshoot application code.
- Experience designing and building infrastructure automation solutions rather than solely administering existing tools.
- Hands-on expertise with Infrastructure as Code technologies and cloud infrastructure management.
- Strong experience with Kubernetes and cloud-native technologies, including deployment, scaling, and operational management.
- Experience working with at least one major public cloud platform such as AWS, Google Cloud Platform (GCP), or Azure.
- Knowledge of CI/CD pipelines, GitOps methodologies, and automated deployment practices.
- Experience implementing observability solutions, including monitoring, logging, alerting, metrics, SLIs, and SLOs.
- Ability to diagnose and resolve complex production issues in high-pressure environments.
- Experience participating in incident response, root cause investigations, and operational reviews.
- Solid problem-solving, analytical, and troubleshooting skills.
- Excellent written and verbal communication skills, with the ability to work effectively in distributed and remote-first environments.
Preferred Qualifications
- Experience building custom infrastructure tools, automation frameworks, operators, controllers, or platform services.
- Strong programming skills in languages such as Go, Python, Ruby, Java, or similar.
- Experience developing and managing Kubernetes operators, controllers, or advanced platform automation.
- Knowledge of distributed systems architecture, scalability patterns, and resilience engineering principles.
- Experience with cloud security, networking, and platform engineering best practices.
- Familiarity with service reliability engineering concepts, error budgets, capacity planning, and performance optimization.
- Experience leveraging AI-enabled tools and automation to improve productivity, operational efficiency, and engineering workflows.
- Demonstrated ability to lead reliability initiatives and influence technical direction across teams.
- Experience mentoring engineers and establishing operational best practices.
📌 Site Reliability Engineer (Canada)
🏢 SoTalent
📍 Canada