Lead DevOps Engineer, Platform (Canada)

Lead DevOps Engineer, Platform (Canada)

29 Aug
|
MedirAI
|
Canada

29 Aug

MedirAI

Canada

MedirAI is building a clinical AI platform for European healthcare. We are a team of 10 with our first clinical pilot running this autumn, and we are building the production infrastructure for it right now. You would be one of two engineers on that platform, and you would own it.

You would report to our CTO. This is a hands-on role: there is no team beneath you initially, and nothing between you and production except a code review and a validation gate. If that sounds like the valuable part, read on.

What you would own

- The AWS accounts and the Kubernetes estate the platform runs on, across multiple availability zones, with separate application and database node pools and enough headroom to drain a node without deadlocking the databases underneath it.
- A GitOps delivery path. Everything reaches production declaratively from the main branch; nobody applies by hand, including you.
- The Postgres clusters running on Kubernetes, their backups, and their restore drills. Clinical data lands here, so restores are practised and timed, not assumed.
- The inference path. This is an AI-first platform: model serving and the gateway in front of it are production surfaces with their own failure modes, and more machine learning workloads are coming onto this estate. Accelerated node pools, capacity and the spend that comes with them are yours.
- A European data boundary operated from Canada. Where data is processed constrains nearly every choice we make — this is the most interesting constraint in the job, not the most annoying one.
- Observability and alerting: metrics, logs, and an alert set a human actually reads.




- The written record — decision records and runbooks that also feed our ISO 27001 work. We decide in writing here.

What you need to have done

- Three or more years supporting Kubernetes in production. Not "has deployed to Kubernetes" — you ran a cluster other people depended on, through upgrades, node drains, evictions and capacity crunches, and you were the one paged when it broke. Five years is welcome and paid accordingly.
- Production experience on a major public cloud at account level: networking, identity and key management, not just compute. AWS preferred.
- Infrastructure as code, reviewed by other people. Terraform preferred, in a plan-review-apply discipline rather than a console.
- Declarative git-driven delivery — Argo CD, Flux or a comparable reconciler — and a real understanding of configuration drift.
- Restored a production database from backup, for real, in a drill or an incident.
- Owned monitoring and alerting rather than only consuming it.
- Carried a production pager, and can walk an incident end to end: what broke, what you did, what changed afterwards.
- Solid Linux and networking fundamentals: DNS, TLS, ingress and load balancers, and a method for debugging "this service cannot reach that one".
- Clear written English, and genuine comfort being one of two.

Nice to have





Managed Kubernetes on AWS · Postgres on Kubernetes via an operator such as CloudNativePG or Patroni · building under a data-residency constraint · regulated-industry experience and audit evidence for ISO 27001 or SOC 2 · secrets management with External Secrets, Vault or a hosted equivalent · cost discipline at small scale · hosting or gating model inference inside a region boundary.

Not required

No certifications. No experience managing a large team — this job is hands-on. No model training or data science background: you run the infrastructure the models need, not the models themselves. No clinical knowledge either — that is learnable here.

On-call

This role carries the production pager, in a rotation of two engineers — one week on, one week off — with best-effort response outside working hours for severity-one incidents only. Our CTO is the break-glass escalation above the rota, so you are never the last line on your own. The rota widens as the team grows, and helping make that case is part of the job. On-call is paid as a standby allowance on top of salary, and time lost to a night incident is taken back the following day. We would rather tell you this now than have you discover it in month two.

How to apply

Send a CV or a LinkedIn profile to [email protected], with a few lines on one cluster you were responsible for: what it ran, who depended on it, and the worst night you had with it. That is the whole application — no cover letter, no take-home. We read every one and reply.

Rémunération : 115 000,00$ à 160 000,00$ par an

Lieu du poste : Télétravail

📌 Lead DevOps Engineer, Platform (Canada)
🏢 MedirAI
📍 Canada

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead devops engineer, platform (canada) / canada