13 Aug
|
Palona AI
|
Toronto
Palona's AI agents operate continuously in production, handle real-time guest interactions, integrate with restaurant systems, and face sharp traffic peaks.
Infrastructure is therefore part of the product: latency, reliability, deployment safety, observability, security, and cost directly shape the guest and operator experience. We are looking for an Infrastructure Engineer who combines cloud and reliability depth with strong software engineering judgment. You will build and operate the platform beneath Palona's AI products, improve how engineers ship, and turn production signals into durable system improvements.
This is not a ticket-driven IT or operations role. You will write production code, design systems, automate repetitive work, and own outcomes across the full service lifecycle. Our current setting includes Python services, Docker, AWS and selected Azure services, ECS and Lambda workloads, API Gateway, load balancers, relational data systems, OpenTofu/Terraform, Datadog, and CI/CD automation.
We value the ability to learn and make sound tradeoffs more than exact tool-for-tool matching.
What you will own: Design, build, and evolve secure, scalable cloud infrastructure for real-time AI services and customer-facing applications
Improve service reliability through clear SLOs, actionable observability, capacity planning, failure testing, and pragmatic incident prevention
Build deployment and release systems that make production changes fast, repeatable, auditable, and safe
Own infrastructure as code, environment consistency, and reusable platform patterns across development, staging, and production
Partner with product and AI engineers on architecture, performance, data flows, and operational readiness for new capabilities
Diagnose complex distributed-system failures across application, network, database, model-provider, and third-party integration boundaries
Reduce infrastructure and model-serving cost without compromising customer experience or engineering velocity
Strengthen secrets management, access controls, backup and recovery, vulnerability management, and other practical security foundations
Build internal tooling and paved paths that let engineers ship and operate services with less manual work
Participate in incident response and turn incidents into better systems, automation, documentation, and engineering judgment Requirements 3+ years industrial experience in relevant technical domain
Strong software engineering fundamentals and experience building or operating production distributed systems
Hands-on experience with a major cloud platform; AWS experience is especially relevant
Experience with containers, infrastructure as code, CI/CD, monitoring, alerting, and production debugging
Ability to write reliable automation and services in Python or another modern programming language
Sound judgment around availability, latency, scalability, security, and cost tradeoffs A track record of taking ambiguous operational problems from diagnosis through durable resolution
Clear communication during architecture reviews, launches, and incidents
AI-native working habits and curiosity about the operational behavior of LLM- and agent-powered systems Benefits Competitive Salary and Stock Option Plan
Medical, dental, vision, retirement, leave, and disability benefits as applicable
Family Leave
Short Term & Long Term Disability
Paid time off and company holidays
Learning and development support
📌 AI Infrastructure Engineer (Toronto)
🏢 Palona AI
📍 Toronto