21 Aug
|
Aptonet
|
Toronto
Senior Application Support / Systems Engineer
Job Title: Senior Associate Technology L2
Client: Leading Retail / eCommerce Company
Location: Toronto, Ontario
Work Arrangement: Hybrid – 2–3 days per week onsite
Contract Duration: 6 months initially
Target Start Date: Immediately
Rate: $105–$115 CAD/hour
Position Summary
We are seeking a Senior Application Support / Systems Engineer to provide hands-on production support for cloud-native, Kubernetes-based applications within a large-scale retail environment.
The successful candidate will be responsible for maintaining application health and reliability, configuring monitoring and alerting solutions, responding to and triaging production incidents, and troubleshooting performance, latency, and availability issues.
This role requires a strong combination of application support, cloud engineering, observability, incident management, CI/CD, and production troubleshooting experience. The ideal candidate will have hands-on expertise with Google Cloud Platform (GCP), Google Kubernetes Engine (GKE), Kubernetes, Jenkins, and modern observability tools .
Key Responsibilities
- Monitor application health, infrastructure signals, service availability, performance, and production workloads across cloud-based environments.
- Configure, maintain, and continuously improve observability solutions, dashboards, alerts, metrics, logs, and distributed tracing.
- Utilize monitoring and observability platforms such as Grafana, Prometheus, Datadog, Dynatrace, Splunk , or similar technologies.
- Support Kubernetes and GKE environments , including troubleshooting workloads, deployments, services, networking, and resource-related issues.
- Troubleshoot production incidents involving application failures, latency, performance degradation, integration issues, and service availability.
- Participate in and drive incident triage, escalation, root-cause analysis, and resolution activities.
- Collaborate with development, DevOps, infrastructure, and client-facing teams to resolve complex production issues.
- Support production deployments and releases, validating application health before, during, and after deployment activities.
- Work with Jenkins and CI/CD pipelines to support reliable software delivery and investigate deployment or pipeline failures.
- Analyze logs, metrics, traces, application behavior,
and system dependencies to identify and resolve issues within distributed environments.
- Identify recurring operational problems and recommend automation, monitoring, configuration, or process improvements to increase system reliability.
- Contribute to operational documentation, troubleshooting procedures, knowledge sharing, and continuous improvement initiatives.
- Maintain clear and timely communication with technical stakeholders throughout production incidents and resolution efforts.
Required Qualifications & Technical Skills
- Strong professional experience in Senior Application Support, Production Support, Systems Engineering, Site Reliability Engineering , or a related discipline.
- Strong hands-on experience supporting production applications and distributed systems .
- Extensive experience with Google Cloud Platform (GCP) – mandatory .
- Hands-on experience with Google Kubernetes Engine (GKE) and Kubernetes-based production environments.
- Robust understanding of observability and monitoring concepts, including:
- Metrics
- Logs
- Alerting
- Dashboards
- Distributed tracing
- Hands-on experience with one or more observability platforms such as Grafana, Prometheus, Datadog, Dynatrace, Splunk , or similar.
- Experience with Jenkins and CI/CD processes.
- Demonstrated experience supporting production deployments and troubleshooting deployment-related issues.
- Strong experience with incident management and production troubleshooting .
- Ability to diagnose application performance, latency, availability, and reliability issues across distributed environments.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration skills.
- Ability to work effectively in a fast-paced production workplace where application availability and reliability are critical.
Preferred Qualifications
- Experience supporting large-scale retail or eCommerce platforms .
- Experience working with microservices architectures and cloud-native applications.
- Experience with multiple observability platforms and the ability to determine the appropriate monitoring strategy for different applications and services.
- Experience with distributed tracing and troubleshooting complex service-to-service dependencies.
- Experience working in highly available, business-critical production environments.
- Familiarity with operational automation and practices designed to improve reliability and incident response.
- Experience working collaboratively with software development, DevOps, infrastructure, and platform engineering teams.
Core Technical Skills
Required:
- GCP
- GKE
- Kubernetes
- Jenkins
- Production Deployments
- Application Support
- Production Troubleshooting
- Incident Management
- Observability / Monitoring
- Distributed Systems
Preferred:
- Grafana
- Prometheus
- Datadog
- Dynatrace
- Splunk
- Distributed Tracing
- Microservices
- CI/CD
- Cloud-Native Applications
- Retail / eCommerce Platforms
Work Arrangement & Candidate Requirements This is a hybrid position based in Toronto, Ontario , with an expectation of working onsite at the office 2–3 days per week .
Occasional local visits to the client site may also be required.
Candidates who are unable to complete the interview process onsite with Publicis Sapient must be prepared to retrieve their company laptop from the Publicis Sapient Toronto office on their start date .
Interview Process The expected interview process consists of:
1. 1–2 internal interview rounds with Publicis Sapient.
2. Final client interview with the end client.
What Success Looks Like The successful candidate will help ensure stable, observable, and reliable production operations by quickly identifying issues, coordinating effective incident resolution, maintaining strong visibility into application health, and supporting safe and reliable production deployments. The ideal engineer will bring a proactive production-support mindset , strong GCP and Kubernetes expertise, and the ability to use observability data to quickly identify root causes. They will also contribute to continuous improvement by identifying recurring issues and implementing solutions that improve application reliability, performance, and operational efficiency.
📌 Application Support Engineer (Toronto)
🏢 Aptonet
📍 Toronto