12 Sep
|
Jobtailor
|
Toronto
- Design, build and operate infrastructure for real-time systems carrying millions of concurrent connections and billions of monthly API requests - Drive Kubernetes end to end, including cluster architecture, workload design and migration of existing services - Re-architect workloads for the AWS-to-GCP migration for cost and performance - Own cloud cost and efficiency work using real spend and utilisation data - Write production Go and Python for internal services, platform tooling and automation - Lead post-migration tuning and capacity planning - Collaborate with backend, video and moderation engineers on system design, reliability targets and cross-service tradeoffs - Participate in on-call, incident response and root cause analysis, turning findings into durable fixes - Own infrastructure projects end to end on a small senior team Requirements - 5+ years in infrastructure, platform, DevOps or SRE engineering, with clear depth in infrastructure over application development - A software engineering background; built systems, not only configured them - Production coding experience in Go or Python; scripting-only backgrounds are not a fit - Kubernetes experience at meaningful production scale, including cluster strategy, workload design, migration leadership, and post-migration cost and efficiency tuning - Personally led cloud cost or efficiency optimisation on AWS or GCP, with a measurable outcome - Direct experience running high-scale, high-load production systems - Strong cloud fundamentals across networking, compute, storage and IAM - Comfortable leading projects and reviewing PRs in a small team - AI tooling already in your engineering workflow; applied use,
not familiarity - Both AWS and GCP, migration experience between providers - PostgreSQL at scale, including sharding, replication strategy, partitioning tradeoffs, ideally self-hosted - Experience with real-time systems such as WebSockets, WebRTC, streaming or persistent-connection workloads - Experience with CockroachDB, Redis, Terraform, and a Prometheus-based observability stack - Experience at an API-first or infrastructure company at scaleup stage - Open source contributions to infrastructure or platform tooling - Writing or talks on cloud, platform or distributed systems - Formal FinOps practice or ownership of cloud commitment and reservation strategy - Work on developer-facing API or SDK products Core Competencies Demonstrates expertise in designing and operating high-scale infrastructure for real-time systems, with a robust focus on cloud cost optimization and efficiency. Proficient in production coding with Go and Python, and experienced in Kubernetes and cloud migration strategies across AWS and GCP. Highest-signal resume keywords - Kubernetes Cluster Architecture - Production Coding in Go - AWS and GCP Migration Experience - Cloud Cost Optimization - PostgreSQL at Scale ATS Optimization Keywords Hard Skills - Go Programming - Python Programming - Kubernetes - AWS - GCP - PostgreSQL - CockroachDB - Redis - Terraform - Real-Time Systems Soft Skills - Project Leadership - Collaboration - Incident Response Industry Keywords - Infrastructure Engineering - DevOps - SRE Engineering - Cloud Fundamentals - FinOps Tools & Technologies - Prometheus - WebSockets - WebRTC - Streaming Workloads - API Development
📌 Senior Software Engineer, Infrastructure (Toronto)
🏢 Jobtailor
📍 Toronto