12 Sep
|
Kaseya
|
Toronto
Kaseya is hiring a Site Reliability Engineer to keep our production systems healthy as we scale. You'll own the reliability of services that thousands of MSPs depend on every day. That means defining the SLOs we hold ourselves to, leading incidents when they happen, and building the automation that keeps things stable as we ship. The work is hands on, the on call rotation is real, and the environment runs heavily on AWS. If you treat reliability as a product instead of a chore, you'll fit in well here.What You'll DoSet, monitor, and enforce SLOs, SLIs, and error budgets that keep our systems reliableLead incident response, troubleshooting, and blameless postmortems that produce real fixesBuild and maintain automated deployment, configuration management, and infrastructure provisioning using Infrastructure as CodeManage cloud and hybrid infrastructure with Terraform or CloudFormation, balancing cost, scalability, and resilienceImprove observability across systems through proactive monitoring, alerting, and dashboards that surface issues earlyPartner with development teams to bake reliability into the SDLC, including deployment automation, capacity planning, and chaos engineeringCut operational toil through automation, systems that recover themselves, and engineering solutions that scaleSupport containerized and serverless workloads so they stay highly available and fault tolerant in productionStay current on SRE, cloud,
and observability practices and bring what works back to the teamRequired Qualifications4 to 5 years of AWS production experienceIaC ownership with Terraform or CloudFormation, including state managementAWS ECS production experience (or strong Kubernetes background willing to ramp)Active on call rotation with incidents led and postmortems writtenWorking fluency with SLOs, SLIs, and error budgets in productionPreferred QualificationsKubernetes production experienceBroader observability tooling (Datadog, Dynatrace, CloudWatch, Elasticsearch/Kibana)Chaos engineeringAWS Lambda or serverless workloadsAnsible, Chef, or PuppetDevSecOps work (vulnerability scanning, secrets management, SOC2 or ISO 27001)Production database support (RDS, PostgreSQL, MySQL)Open source contributions or public technical portfolioThe expected annual base salary for this role is CAD $115,000 to CAD $130,000. Final offer will depend on experience, skills, and internal equity. This posting is for an existing vacancy.Additional informationKaseya provides equal employment chance to all employees and applicants without regard to race, religion, age, ancestry, gender, sex, sexual orientation, national origin, citizenship status, physical or mental disability, veteran status, marital status, or any other characteristic protected by applicable law.#J-18808-Ljbffr
📌 Site Reliability Engineer Markham, Ontario - C$115,000 - C$130,000 A Year (Toronto)
🏢 Kaseya
📍 Toronto