29 Aug
|
Aptonet
|
Winnipeg
Job Title:
Senior Associate Technology L2
Contract Duration:
6 months initially
Target Start Date:
Immediately
Position Summary We are seeking a
Senior Application Support / Systems Engineer
to provide hands‑on production support for cloud‑native, Kubernetes‑based applications within a large‑scale retail environment.
The successful candidate will be responsible for maintaining application health and reliability, configuring monitoring and alerting solutions, responding to and triaging production incidents, and troubleshooting performance, latency, and availability issues.
This role requires a strong combination of application support, cloud engineering, observability, incident management, CI/CD, and production troubleshooting experience. The ideal candidate will have hands‑on expertise with
Google Cloud Platform (GCP), Google Kubernetes Engine (GKE), Kubernetes, Jenkins, and modern observability tools .
Key Responsibilities
Monitor application health, infrastructure signals, service availability, performance, and production workloads across cloud‑based environments.
Configure, maintain, and continuously improve observability solutions, dashboards, alerts, metrics, logs, and distributed tracing.
Utilize monitoring and observability platforms such as
Grafana, Prometheus, Datadog, Dynatrace, Splunk , or similar technologies.
Support
Kubernetes and GKE environments , including troubleshooting workloads, deployments, services, networking, and resource‑related issues.
Troubleshoot production incidents involving application failures, latency, performance degradation, integration issues, and service availability.
Participate in and drive incident triage, escalation, root‑cause analysis, and resolution activities.
Collaborate with development, DevOps, infrastructure, and client‑facing teams to resolve complex production issues.
Support production deployments and releases, validating application health before, during, and after deployment activities.
Work with
Jenkins and CI/CD pipelines
to support reliable software delivery and investigate deployment or pipeline failures.
Analyze logs, metrics, traces, application behavior, and system dependencies to identify and resolve issues within distributed environments.
Identify recurring operational problems and recommend automation, monitoring, configuration, or process improvements to increase system reliability.
Contribute to operational documentation, troubleshooting procedures, knowledge sharing, and continuous improvement initiatives.
Maintain clear and timely communication with technical stakeholders throughout production incidents and resolution efforts.
Required Qualifications & Technical Skills
Strong professional experience in
Senior Application Support, Production Support, Systems Engineering, Site Reliability Engineering , or a related discipline.
Strong hands‑on experience supporting
production applications and distributed systems .
Extensive experience with
Google Cloud Platform (GCP) – mandatory .
Hands‑on experience with
Google Kubernetes Engine (GKE)
and Kubernetes‑based production environments.
Strong understanding of observability and monitoring concepts, including:
Metrics
Logs
Alerting
Distributed tracing
Hands‑on experience with one or more observability platforms such as
Grafana, Prometheus, Datadog, Dynatrace, Splunk , or similar.
Experience with
Jenkins
and CI/CD processes.
Demonstrated experience supporting
production deployments
and troubleshooting deployment‑related issues.
Strong experience with
incident management and production troubleshooting .
Ability to diagnose application
performance, latency, availability, and reliability issues
across distributed environments.
Strong analytical and problem‑solving skills.
Excellent communication and collaboration skills.
Ability to work effectively in a fast‑paced production environment where application availability and reliability are critical.
Preferred Qualifications
Experience supporting
large‑scale retail or eCommerce platforms .
Experience working with
microservices architectures
and cloud‑native applications.
Experience with multiple observability platforms and the ability to determine the appropriate monitoring strategy for different applications and services.
Experience with
distributed tracing
and troubleshooting complex service‑to‑service dependencies.
Experience working in highly available, business‑critical production environments.
Familiarity with operational automation and practices designed to improve reliability and incident response.
Experience working collaboratively with software development, DevOps, infrastructure, and platform engineering teams.
Core Technical Skills Required:
GCP
GKE
Jenkins
Production Deployments
Application Support
Production Troubleshooting
Observability / Monitoring
Distributed Systems
Preferred:
Grafana
Dynatrace
Splunk
Microservices
CI/CD
Cloud‑Native Applications
Retail / eCommerce Platforms
Work Arrangement & Candidate Requirements This is a
hybrid position based in Toronto, Ontario , with an expectation of working onsite at the office
2-3 days per week .
Occasional local visits to the client site may also be required.
Candidates who are unable to complete the interview process onsite with Publicis Sapient must be prepared to
retrieve their company laptop from the Publicis Sapient Toronto office on their start date .
Interview Process The expected interview process consists of:
1-2 internal interview rounds
with Publicis Sapient.
Final client interview
with the end client.
What Success Looks Like The successful candidate will help ensure stable, observable, and reliable production operations by quickly identifying issues, coordinating effective incident resolution, maintaining strong visibility into application health, and supporting safe and reliable production deployments.
The ideal engineer will bring a
proactive production‑support mindset , robust GCP and Kubernetes expertise, and the ability to use observability data to quickly identify root causes. They will also contribute to continuous improvement by identifying recurring issues and implementing solutions that improve application reliability, performance, and operational efficiency.
#J-18808-Ljbffr
📌 Application Support Engineer (Winnipeg)
🏢 Aptonet
📍 Winnipeg