Ops, Site Reliability Engineering (SRE) & Dynatrace Required Skills Strong experience as a Platform Engineer with expertise in Dev
Ops and Site Reliability Engineering (SRE).
Experience designing, implementing, automating, and supporting enterprise-scale platform infrastructure.
Strong knowledge of high availability, reliability, scalability, and performance engineering for mission-critical applications.
Hands-on experience with Dynatrace, including: Monitoring Dashboard creation Synthetic monitoring Observability Automation Incident management Strong understanding of SRE best practices and cloud/platform engineering.
Ability to drive operational excellence through automation, monitoring, reliability engineering, and continuous improvement initiatives.
Key Responsibilities Design, build, and maintain highly available, scalable, and resilient platform infrastructure.
Implement contemporary Platform Engineering and Site Reliability Engineering (SRE) practices across enterprise applications.
Define and maintain:
Service Level Indicators (SLIs) Service Level Objectives (SLOs) Error Budgets Drive initiatives focused on: Reliability Availability Capacity planning Performance optimization Operational excellence Support production environments and participate in on-call rotations when required.
Observability & Monitoring Lead the implementation and administration of enterprise monitoring and observability solutions.
Develop and maintain Dynatrace monitoring strategies for complex distributed systems.
Create and manage: Dynatrace dashboards Alerts Management Zones Reporting solutions Implement proactive monitoring for: Infrastructure Middleware Applications Databases APIs Cloud services Configure and optimize: Anomaly detection Problem management Root cause analysis Dynatrace Expertise Hands-on experience with: Dynatrace One