16 Sep
|
Intellibus
|
Toronto
16 Sep
Intellibus
Toronto
Imagine working at Intellibus to engineer platforms that impact billions of lives around the world.
Our Platform Engineering
Team is looking for experienced DevOps / SRE leaders who can help build highly reliable, observable, and automated infrastructure supporting mission-critical applications. We are looking for hands-on technical leaders with deep experience in Datadog, infrastructure automation, configuration management, Terraform/Chef, Bash/Shell scripting, cloud infrastructure, and production systems. The ideal candidate will also have a strong understanding of Java-based applications and Java coding, as this role will work closely with Java engineering teams and distributed application platforms.
Lead DevOps/SRE initiatives across mission-critical environments. Automate infrastructure provisioning and configuration using Terraform, Chef, or similar tools. Manage deployment, configuration, and environment automation across development, QA, and production.
Troubleshoot complex production, infrastructure, networking, and application issues. Improve system availability, scalability, performance, and reliability. Support CI/CD pipelines and automated deployments.
Work closely with Java engineering teams to understand application behavior, performance, dependencies, and production issues.
Analyze
Java applications from an operational perspective, including JVM behavior, memory, CPU, threads, logs, and application performance. Participate in incident response, root-cause analysis, and post-mortems.
Provide technical leadership and mentor other DevOps/SRE engineers.
Core DevOps / SRE Responsibilities Monitor applications, infrastructure, services, APIs, and databases. Configure APM, logs, metrics, traces, and service-level monitoring. Identify performance and reliability issues before they impact clients.
Automate configuration management and environment provisioning. Manage infrastructure consistency and configuration drift. Automate deployments, operational processes, monitoring, and infrastructure tasks.
Python scripting is a plus. This is not a Java Developer position, but candidates must have a strong understanding of Java-based applications. Read and understand Java code.
Troubleshoot
Java application issues from an infrastructure/SRE perspective. Understand JVM, memory, CPU, threads, garbage collection, and application performance. Work effectively with Java/Spring Boot engineering teams. 12+ years of experience in DevOps, SRE, Infrastructure Engineering, Platform Engineering, or related roles. ~8+ years of hands-on DevOps/SRE leadership experience. ~ Solid understanding of Java applications and Java coding. ~ Experience with CI/CD and deployment automation. ~ Strong understanding of networking fundamentals. ~ Experience with monitoring, logging, alerting, and observability. ~ Experience leading technical initiatives and mentoring engineers. ~ AWS Python Performance engineering Schedule a 15 min Video Call with someone from our Team ~30-45 min Final & Technical Video Interview ~
📌 Lead Site Reliability Engineer (SRE) - DevOps & Observability (Toronto)
🏢 Intellibus
📍 Toronto