13 Sep
|
Jobtailor
|
Toronto
- Design, build and operate infrastructure for real-time systems carrying millions of concurrent connections and billions of monthly API requests
- Drive Kubernetes end to end, including cluster architecture, workload design and migration of existing services
- Re-architect workloads for the AWS-to-GCP migration for cost and performance
- Own cloud cost and efficiency work using real spend and utilisation data
- Write production Go and Python for internal services, platform tooling and automation
- Lead post-migration tuning and capacity planning
- Collaborate with backend, video and moderation engineers on system design, reliability targets and cross-service tradeoffs
- Participate in on-call, incident response and root cause analysis, turning findings into durable fixes
- Own infrastructure projects end to end on a small senior team
Requirements
- 5+ years in infrastructure, platform, DevOps or SRE engineering, with clear depth in infrastructure over application development
- A software engineering background; built systems, not only configured them
- Production coding experience in Go or Python; scripting-only backgrounds are not a fit
- Kubernetes experience at meaningful production scale, including cluster strategy, workload design, migration leadership, and post-migration cost and efficiency tuning
- Personally led cloud cost or efficiency optimisation on AWS or GCP, with a measurable outcome
- Direct experience running high-scale, high-load production systems
- Strong cloud fundamentals across networking, compute, storage and IAM
- Comfortable leading projects and reviewing PRs in a small team
- AI tooling already in your engineering workflow; applied use,
not familiarity
- Both AWS and GCP, migration experience between providers
- PostgreSQL at scale, including sharding, replication strategy, partitioning tradeoffs, ideally self-hosted
- Experience with real-time systems such as WebSockets, WebRTC, streaming or persistent-connection workloads
- Experience with CockroachDB, Redis, Terraform, and a Prometheus-based observability stack
- Experience at an API-first or infrastructure company at scaleup stage
- Open source contributions to infrastructure or platform tooling
- Writing or talks on cloud, platform or distributed systems
- Formal FinOps practice or ownership of cloud commitment and reservation strategy
- Work on developer-facing API or SDK products
Core Competencies
Demonstrates expertise in designing and operating high-scale infrastructure for real-time systems, with a solid focus on cloud cost optimization and efficiency. Proficient in production coding with Go and Python, and experienced in Kubernetes and cloud migration strategies across AWS and GCP.
Highest-signal resume keywords
- Kubernetes Cluster Architecture
- Production Coding in Go
- AWS and GCP Migration Experience
- Cloud Cost Optimization
- PostgreSQL at Scale
ATS Optimization Keywords
Hard Skills
- Go Programming
- Python Programming
- Kubernetes
- AWS
- GCP
- PostgreSQL
- CockroachDB
- Redis
- Terraform
- Real-Time Systems
Soft Skills
- Project Leadership
- Collaboration
- Incident Response
Industry Keywords
- Infrastructure Engineering
- DevOps
- SRE Engineering
- Cloud Fundamentals
- FinOps
Tools & Technologies
- Prometheus
- WebSockets
- WebRTC
- Streaming Workloads
- API Development
📌 Senior Software Engineer, Infrastructure (Toronto)
🏢 Jobtailor
📍 Toronto