Senior DevOps Engineer- Platform (Toronto)

Senior DevOps Engineer- Platform (Toronto)

03 Aug
|
Tru
|
Toronto

03 Aug

Tru

Toronto

Tru Inc is seeking a Senior DevOps / Site Reliability Engineer . We have one (1) permanent full time position available for an immediate start. At Tru, we are more than just a technology agency—we are strategic partners committed to driving digital transformation for enterprise-level clients.

As experts in digital experience platforms, e-commerce engines, and innovative solutions, we help businesses thrive in the digital era. We are looking for a Senior DevOps / Site Reliability Engineer to help build, operate, monitor, and scale our new enterprise AI platform. This role will be responsible for the DevOps and reliability capabilities supporting the platform across all environments, including Development, QA, and Production.

The environment is currently focused and manageable, consisting of approximately 10–30 containers, but is expected to grow as the AI platform expands in 2027 and beyond. The most important part of the role will be creating intelligent dashboards, monitoring, alerting, automated scaling, and AI-assisted platform management. The platform is hosted primarily in Microsoft Azure with limited exposure to AWS.

Occasional travel to client sites or workshops may also be necessary.

Azure

Infrastructure and Platform Deployment

- Design, deploy, configure and maintain infrastructure within Microsoft Azure. Deploy and manage virtual machines, containers, Kubernetes clusters, networking, storage and supporting platform services. Establish repeatable and reliable deployment processes using Infrastructure as Code and CI/CD automation. Manage container lifecycle, configuration, secrets, networking, storage and application dependencies. Monitor container and cluster health, resource consumption, capacity and performance. Troubleshoot deployment, networking, configuration, and runtime issues. Establish appropriate standards for container deployment and Kubernetes operations.



Performance, Load Management, and Scaling:
- Monitor platform demand, workload patterns, resource utilization and application performance. Configure horizontal and vertical scaling policies for containers and supporting infrastructure. Conduct capacity planning and identify potential performance bottlenecks before they affect production. Help introduce predictive or AI-assisted scaling and platform management capabilities. Build dashboards covering availability, performance, capacity, errors, latency, traffic, container health, AI workloads and service dependencies. Continuously improve dashboards so that issues, trends and risks can be quickly identified. Advanced dashboard design and dashboard-building experience is a core requirement for this role. Monitoring and Alerting:
- Implement monitoring and alerting across infrastructure, applications, containers, integrations and AI platform services. Configure actionable alerts that identify real production risks while minimizing unnecessary alert noise. Establish thresholds, anomaly detection, health checks, synthetic monitoring and automated remediation where appropriate. Support the stability, availability and operational readiness of the production AI platform. Investigate and resolve platform, deployment, infrastructure, monitoring and performance issues. Participate in root-cause analysis and implement preventative improvements. Help prepare the operating model for increased platform adoption and support requirements expected in 2027.



Strong professional experience in DevOps, Site Reliability Engineering, cloud infrastructure or platform engineering. Advanced hands‑on experience with Microsoft Azure. Advanced experience designing and building dashboards using Grafana, Kibana, Azure Monitor, Application Insights or comparable tools. Strong experience implementing monitoring, observability, logging, alerting and operational health checks.

Experience managing application load, infrastructure capacity, performance and automated scaling.

Experience with CI/CD pipelines and automated application deployment. Strong troubleshooting skills across applications, containers, infrastructure, networking and cloud services. Ability to work independently while collaborating closely with developers, architects, AI engineers and platform stakeholders.

Experience supporting AI, machine learning, data or high-compute platforms.

Experience monitoring AI models, inference services, token usage, GPU workloads, API consumption, queues or model performance.

Experience implementing automated remediation, predictive monitoring or AI-assisted platform operations. Familiarity with AWS services and cloud operations. The role offers the opportunity to establish the platform correctly from the beginning, introduce modern DevOps and SRE practices, and experiment with intelligent monitoring, automated scaling, advanced dashboards, and AI-assisted platform management.

We take pride in building a culture that stands out for its courage, entrepreneurial culture, diversity, and passion for people. We are proud to offer competitive salaries along with a 100% employer-paid benefits package with a remote work- from- home arrangement. Quality as Standard Join us and be part of a team that makes a tangible impact in the digital world! #

📌 Senior DevOps Engineer- Platform (Toronto)
🏢 Tru
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior devops engineer- platform (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: senior devops engineer- platform (toronto) / toronto