01 Aug
|
Jobtailor
|
Toronto
- Own technical direction and architecture evolution of a live, revenue-generating platform - Scale throughput and reliability while customers onboard, including zero-downtime changes to running systems - Harden the platform: tenant isolation, rate limiting, security, data integrity - Establish observability, SLOs, and incident practices as the customer base grows - Lead performance and capacity work ahead of demand, not behind it - Raise engineering standards through design reviews and mentorship - Partner with Product to balance platform investment against feature delivery - Identify and retire scaling risks and technical debt before they compound Requirements - 10+ years building and operating cloud-native backend or platform systems - Owned the scaling of a production SaaS or data platform through rapid customer growth: hardening, re-architecture under load, and maturing an early product to enterprise grade - Deep hands-on experience with high-throughput distributed systems: event streaming or message queues, async processing, idempotency, retries, backpressure - Strong performance engineering: profiling, load testing, horizontal scaling, capacity planning - Production reliability ownership: observability, SLOs, incident management, fault-tolerant design - Multi-tenant SaaS experience, including tenant isolation and noisy-neighbor problems - Proficiency in Python and JavaScript/TypeScript (Next.js) - Hands-on daily coder with founder-level ownership; comfortable in low-process,
high-ambiguity environments - Built or scaled an integration platform, ETL/data pipeline product, or API platform (connectors, transformation pipelines, webhooks) (Nice to Have) - Kafka or similar streaming infrastructure at scale (Nice to Have) - Supply chain or logistics systems exposure (WMS, OMS, TMS, EDI integrations) (Nice to Have) - AI-assisted engineering workflows; AI is deeply embedded in how we build (Nice to Have) - Experience shipping AI/LLM-driven capabilities in production (rag, generation pipeline, agentic workflows) (Nice to Have) Core Competencies Demonstrates expertise in scaling cloud-native backend systems, ensuring production reliability, and implementing performance engineering practices. Proficient in Python and JavaScript/TypeScript, with a robust focus on multi-tenant SaaS architecture and observability. Highest-signal resume keywords - Cloud-Native Backend Systems - Performance Engineering - Production Reliability Ownership - Multi-Tenant SaaS Experience - High-Throughput Distributed Systems ATS Optimization Keywords Hard Skills - Python - JavaScript - TypeScript - Event Streaming - Message Queues - Async Processing - Load Testing - Capacity Planning - Fault-Tolerant Design - Tenant Isolation Soft Skills - Mentorship - Leadership - Collaboration Industry Keywords - SaaS - Data Platform - Observability - Incident Management - Scaling Risks - Technical Debt Tools & Technologies - Kafka - ETL - API Platforms - Integration Platforms - AI/LLM-Driven Capabilities
📌 Principal Engineer (Toronto)
🏢 Jobtailor
📍 Toronto