22 Sep
|
HCL Technologies
|
Ontario
22 Sep
HCL Technologies
Ontario
AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams. Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards. Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.
Key Responsibilities AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams.
Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards. Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.
Skill Requirements AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams. Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards.
Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.
Other Requirements AWS SRE Admin: The SRE Administrator is responsible for maintaining the reliability, availability, performance, and operational stability of applications and infrastructure hosted on AWS. The role focuses on monitoring, incident management, platform administration, observability, automation, and operational excellence to ensure seamless business services across the technology estate. The AWS SRE Administrator works closely with application, cloud, infrastructure, security, and operations teams to proactively identify issues, optimize platform performance, and drive continuous service improvements. Key Responsibilities: Monitor AWS-hosted applications, infrastructure, and services to ensure high availability and performance. Perform proactive health checks, alert monitoring, incident triage, and operational support activities. Manage AWS services including EC2, EKS, ECS, Lambda, RDS, S3, Route 53, CloudWatch, and related platform components. Investigate production incidents, perform root cause analysis, and coordinate resolution with support and engineering teams. Administer observability platforms and monitoring tools, including dashboard maintenance, alert tuning, and reporting. Support production releases by performing deployment validation, smoke testing, and post-change monitoring activities. Develop and maintain operational runbooks, standard operating procedures, and knowledge documentation. Automate repetitive operational tasks using AWS native services, scripting, and infrastructure-as-code tools. Monitor platform capacity, utilization trends, and system reliability metrics to support capacity planning. Collaborate with security, infrastructure, and application teams to ensure compliance with operational and security standards. Generate operational reports, SLA/KPI dashboards, and service health updates for stakeholders. Drive continuous improvement initiatives focused on reliability, automation, operational efficiency, and reduction of manual effort.
At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.
HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026totaled $14.8billion.
#J-18808-Ljbffr
📌 SME - Kubernetes, Terraform (Ontario)
🏢 HCL Technologies
📍 Ontario