31 Aug
|
HCL Technologies
|
Mississauga
31 Aug
HCL Technologies
Mississauga
Job Title: Chaos Engineer Experience: 5+ Years Key Responsibilities Design and execute resiliency and chaos testing scenarios across cloud, Kubernetes, APIs, microservices, and distributed environments. Identify resilience gaps, SPOFs, and operational risks and drive remediation.
Validate High
Availability (HA), Disaster Recovery (DR), failover, auto-healing, and recovery processes. Automate testing and reliability validation using scripting and cloud-native tools. Monitor application and infrastructure behavior using observability platforms such as AppDynamics, Prometheus, and Grafana. Collaborate with DevOps, SRE, Infrastructure, and Application teams to improve system reliability. Define and track reliability metrics including SLA, SLO, SLI, MTTR, RTO, and RPO.
Required Skills Strong experience in Java/Python, Microservices, APIs, Kubernetes, Docker, and Cloud Platforms (Azure/GCP/AWS). Hands-on experience with monitoring and observability tools such as Dynatrace metrics and observability, AppDynamics, Prometheus, and Grafana. Knowledge of Reliability Engineering, Chaos Testing, Incident Analysis, and Resiliency Validation. Failure-as-a-Service platforms to achieve resiliency in infrastructure failures, network and application failures.
Chaos
Testing mechanism and tools such as AWS-FIS, Lambda testing, On-premise , OpenShift testing. Monitoring tools test/scenario capture and report creation through standardized templates and hypothesis formation.
Experience with automation, CI/CD, and cloud-native architectures. Excellent troubleshooting, analytical, and communication skills.
Key Responsibilities Job Title: Chaos Engineer Experience: 5+ Years Key Responsibilities Design and execute resiliency and chaos testing scenarios across cloud, Kubernetes, APIs, microservices, and distributed environments. Identify resilience gaps, SPOFs, and operational risks and drive remediation.
Validate High
Availability (HA), Disaster Recovery (DR), failover, auto-healing, and recovery processes. Automate testing and reliability validation using scripting and cloud-native tools. Monitor application and infrastructure behavior using observability platforms such as AppDynamics, Prometheus, and Grafana. Collaborate with DevOps, SRE, Infrastructure, and Application teams to improve system reliability. Define and track reliability metrics including SLA, SLO, SLI, MTTR, RTO, and RPO.
Required Skills Strong experience in Java/Python, Microservices, APIs, Kubernetes, Docker, and Cloud Platforms (Azure/GCP/AWS).
Hands-on experience with monitoring and observability tools such as Dynatrace metrics and observability, AppDynamics, Prometheus, and Grafana. Knowledge of Reliability Engineering, Chaos Testing, Incident Analysis, and Resiliency Validation. Failure-as-a-Service platforms to achieve resiliency in infrastructure failures, network and application failures.
Chaos
Testing mechanism and tools such as AWS-FIS, Lambda testing, On-premise , OpenShift testing. Monitoring tools test/scenario capture and report creation through standardized templates and hypothesis formation.
Experience with automation, CI/CD, and cloud-native architectures. Excellent troubleshooting, analytical, and communication skills.
Skill Requirements Job Title: Chaos Engineer Experience: 5+ Years Key Responsibilities Design and execute resiliency and chaos testing scenarios across cloud, Kubernetes, APIs, microservices, and distributed environments. Identify resilience gaps, SPOFs, and operational risks and drive remediation.
Validate High
Availability (HA), Disaster Recovery (DR), failover, auto-healing, and recovery processes. Automate testing and reliability validation using scripting and cloud-native tools. Monitor application and infrastructure behavior using observability platforms such as AppDynamics, Prometheus, and Grafana. Collaborate with DevOps, SRE, Infrastructure, and Application teams to improve system reliability. Define and track reliability metrics including SLA, SLO, SLI, MTTR, RTO, and RPO.
Required Skills Solid experience in Java/Python, Microservices, APIs, Kubernetes, Docker, and Cloud Platforms (Azure/GCP/AWS). Hands-on experience with monitoring and observability tools such as Dynatrace metrics and observability, AppDynamics, Prometheus, and Grafana. Knowledge of Reliability Engineering, Chaos Testing, Incident Analysis, and Resiliency Validation. Failure-as-a-Service platforms to achieve resiliency in infrastructure failures, network and application failures.
Chaos
Testing mechanism and tools such as AWS-FIS, Lambda testing, On-premise , OpenShift testing. Monitoring tools test/scenario capture and report creation through standardized templates and hypothesis formation.
Experience with automation,
CI/CD, and cloud-native architectures. Excellent troubleshooting, analytical, and communication skills.
Other Requirements Job Title: Chaos Engineer Experience: 5+ Years Key Responsibilities Design and execute resiliency and chaos testing scenarios across cloud, Kubernetes, APIs, microservices, and distributed environments. Identify resilience gaps, SPOFs, and operational risks and drive remediation.
Validate High
Availability (HA), Disaster Recovery (DR), failover, auto-healing, and recovery processes. Automate testing and reliability validation using scripting and cloud-native tools. Monitor application and infrastructure behavior using observability platforms such as AppDynamics, Prometheus, and Grafana. Collaborate with DevOps, SRE, Infrastructure, and Application teams to improve system reliability. Define and track reliability metrics including SLA, SLO, SLI, MTTR, RTO, and RPO.
Required Skills Strong experience in Java/Python, Microservices, APIs, Kubernetes, Docker, and Cloud Platforms (Azure/GCP/AWS). Hands-on experience with monitoring and observability tools such as Dynatrace metrics and observability, AppDynamics, Prometheus, and Grafana. Knowledge of Reliability Engineering, Chaos Testing, Incident Analysis, and Resiliency Validation. Failure-as-a-Service platforms to achieve resiliency in infrastructure failures, network and application failures.
Chaos
Testing mechanism and tools such as AWS-FIS, Lambda testing, On-premise , OpenShift testing. Monitoring tools test/scenario capture and report creation through standardized templates and hypothesis formation.
Experience with automation, CI/CD, and cloud-native architectures. Excellent troubleshooting, analytical, and communication skills.
At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.
HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026totaled $14.8billion.
📌 SME - Kubernetes, Terraform (Mississauga)
🏢 HCL Technologies
📍 Mississauga