28 Aug
|
TOTEM Recruteur de talent
|
Toronto
28 Aug
TOTEM Recruteur de talent
Toronto
Employment Status: Permanent Schedule: 40 hours/week 100% remote work Job Description We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms.
Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infrastructure supporting critical cloud-based services.
This role combines hands-on engineering with operational leadership, giving you direct ownership of system availability, scalability, and incident response.
You will be involved throughout the service lifecycle, from architecture and launch preparation to production monitoring and continuous improvement.
Responsibilities Partner with engineering teams during system design, capacity planning, launch readiness, and production deployment.
Monitor and improve service availability, latency, performance, and overall system health.
Identify recurring operational issues and implement sustainable solutions that improve scalability and resilience.
Define and evolve observability practices, including dashboards, alerts, SLOs, and SLIs.
Build and maintain automated infrastructure using Terraform and CI/CD pipelines.
Develop automation and operational tooling to reduce manual intervention and support self-healing systems.
Coordinate incident response and act as Incident Commander during critical production events.
Facilitate blameless post-incident reviews and ensure that corrective actions are completed.
Use AI-assisted engineering tools responsibly to accelerate development and operational workflows.
Maintain clear technical documentation and contribute to the continuous improvement of SRE practices.
Required profile Significant experience in Site Reliability Engineering, Cloud Engineering, Dev
Ops, or a similar infrastructure-focused role.
Experience supporting complex or large-scale SaaS environments with high availability requirements.
Strong hands-on knowledge of AWS services and architecture, including multi-account environments, VPC, EC2, and EKS.
Proven experience operating and troubleshooting Kubernetes environments at scale.
Strong knowledge of Infrastructure as Code, particularly Terraform.
Experience building or maintaining CI/CD pipelines using Git
Lab, Jenkins, or comparable tools.
Experience with enterprise observability platforms such as Datadog, Prometheus, Grafana, or equivalent solutions.
Robust scripting skills using Python, Bash, or a similar language.
Direct experience participating in on-call rotations, coordinating incident response, and conducting post-incident reviews.
Familiarity with Java or .NET application environments is considered an asset.
Ability to communicate clearly and collaborate with development, infrastructure, security, and operations teams.
Must be legally authorized to work in Canada.
What to Expect Remote-first work environment within Canada.
Occasional visits to a local office or participation in in-person meetings may be required, representing less than 10% of the role.
Participation in a scheduled on-call rotation is required.
Does this opportunity sound like a good fit for you? Apply now through our website or by sending your resume to .
Thank you for your interest in this position; only candidates who meet our clients requirements will be contacted.
The masculine gender is used as a neutral form. #totemtech
📌 Site Reliability Engineer (Toronto)
🏢 TOTEM Recruteur de talent
📍 Toronto