Senior Infrastructure Engineer (Linux & Cloud Automation) (Stone Mills)

Senior Infrastructure Engineer (Linux & Cloud Automation) (Stone Mills)

26 Sep
|
Manx Telecom Enterprise
|
Stone Mills

26 Sep

Manx Telecom Enterprise

Stone Mills

Applicant Portal

:

Job Details: Senior Infrastructure Engineer (Linux &

- Cloud Automation)

Full details of the job.

Vacancy Name Senior Infrastructure Engineer (Linux &

- Cloud Automation) Vacancy No VN970 Function Syn - Delivery &
- Operations Work Location One St Peter's Square, Manchester, M2 3DE Basis Permanent Full Time/Part Time Full time Employment Duration - Hours Per Week 35.00 Drivers Licence Required No Benefits Competitive Salary and Advantages package will be provided to the successful candidate. About the Role

Position Overview

As a Senior Infrastructure Engineer (Linux &

- Cloud Automation) at Synapse360, you are responsible for the operational delivery and automation of Linux-based infrastructure across managed service customers.

You will own the end-to-end patching lifecycle for Linux server estates using Ansible, and manage Azure Update Manager policies for Linux VMs across Azure. You will manage and remediate backup operations for Linux workloads via Azure Backup, implement infrastructure-as-code deployments using Bicep and ARM templates, maintain CI/CD pipelines via Azure DevOps, and handle escalations from the Infrastructure team.

This is a demanding, hands-on engineering role requiring strong proficiency in Linux administration, automation tooling, Microsoft Azure, and DevOps practices.

You will participate in the weekly on-call rota and must be capable of independently resolving P2 incidents and supporting the Infrastructure Lead during P1 major incident response.

What We Expect

- Minimum of 5–7 years of experience in enterprise infrastructure engineering, ideally within managed services or a multi-customer MSP environment.
- Demonstrated ownership of large-scale Linux patching programmes using Ansible, including scheduling, deployment, validation, remediation, and compliance reporting.
- Hands-on experience with Ansible — playbook development and execution for patching, configuration management, and compliance enforcement.
- Proficiency in Microsoft Azure — specifically Linux VM management, Azure Update Manager, Azure Backup, Recovery Services Vaults, NSG/UDR configuration, and Azure networking.
- Strong Linux administration skills across RHEL, CentOS, Ubuntu, and Oracle Linux — package management, service management, log analysis, and security hardening.
- Experience with Infrastructure-as-Code using Bicep and/or ARM templates for repeatable Azure deployments.
- Familiarity with Azure DevOps —
- CI/CD pipelines, repos, and release management for infrastructure automation.
- Willingness to participate in a 1-in-5 weekly on-call rotation providing 24×7 P1 incident response.

Areas of Responsibility

- Patching: Own the monthly patching cycle for Linux servers. Deploy patches via Ansible playbooks. Manage patch scheduling, pre-patch snapshots, deployment waves, post-patch validation, and remediation. Achieve >95% patching compliance SLA.

Manage Azure Update

Manager policies for the Linux Azure estate — configure maintenance windows, deploy OS-level patches, validate compliance, and remediate failures.

- Backup Management: Manage Azure Backup policies for Linux workloads, monitor Recovery Services Vaults, perform restore testing, and maintain backup compliance. Develop and maintain scripted backup validation routines. Target >98% backup success rate.
- Infrastructure-as-Code &

- Automation: Develop and maintain Bicep and ARM templates for Azure resource deployments. Build and manage Azure DevOps CI/CD pipelines for infrastructure provisioning and configuration.

Maintain

Ansible playbooks for configuration management, compliance enforcement, and operational automation.

- VM Lifecycle Management: Provision, configure, snapshot, resize, and decommission Linux virtual machines on Azure (IaaS). Manage VM images and ensure compliance with customer standards.
- Azure Key Vault &

- Secrets Management: Manage secrets,



certificates, and keys within Azure Key Vault.

Implement certificate rotation and integrate Key Vault with automated deployments.

- Change Management: Prepare and implement standard and normal RFCs. Document changes fully including risk assessment, rollback procedures, and post-implementation reviews. Attend weekly CAB as required.
- T1 Escalation Handling: Receive and resolve Linux and cloud automation incidents escalated by T1 engineers.

Provide guidance and knowledge transfer to T1 team members.

- ServiceNow &
- Reporting: Maintain accurate ticket records, contribute to SLA reporting, and ensure all work is logged against the correct customer and category.

On-Call Commitment This role participates in a shared five-person weekly on-call rotation.

Each engineer is primary on-call for one week in every five. During on-call periods, you are expected to respond to P1 critical alerts within 15 minutes (24×7), independently resolve P2 incidents, and support the Infrastructure Lead during major incident response. On-call compensation is provided in line with Synapse360's standard on-call policy.

Ideal Candidate Characteristics

Important Attributes

- Patching Ownership: You are the patching authority for Linux estates. Full lifecycle ownership — scheduling, risk assessment, Ansible playbook deployment, validation, remediation, and compliance reporting.

Patching is a critical SLA metric (>95% compliance) and must be treated as a controlled change.

- Automation &
- DevOps Mindset: You must be comfortable building and maintaining Ansible playbooks, Bicep templates, and Azure DevOps pipelines. The role demands a strong drive toward automation, repeatability, and infrastructure-as-code principles.
- Backup Operations: You own the remediation of Azure Backup failures for Linux workloads and scripted backup validation.

You must understand backup policies and Recovery Services Vault configurations to troubleshoot and maintain compliance.

- Change Management Discipline: You will prepare and implement standard and normal RFCs, ensuring full documentation, risk assessment, rollback plans, and post-implementation validation. You will attend and present at the weekly CAB as required.
- Escalation Handling: You are the escalation point for T1 engineers on Linux and cloud automation issues. You must be able to receive a partially triaged incident and drive it to resolution without unnecessary re-escalation to Infrastructure Lead.

Areas of Responsibility
- Monitoring &
- Alert Management: Manage and triage alerts from LogicMonitor and Azure Monitor across environments. Correlate alerts, close false positives, and escalate genuine incidents.

Maintain monitoring dashboards and ensure alert thresholds remain appropriate.

- ServiceNow Ticket Management: Own the ticket lifecycle for incoming incidents and service requests. Triage, categorise, prioritise, and either resolve or escalate. Maintain SLA compliance for response and update targets.

Expected to handle 15–25 tickets per day across environments.

- Backup Validation: Perform daily Veeam Backup &
- Replication job status checks (~1,290 servers) and Azure Backup status validation. Log failures, attempt basic remediation (re-run failed jobs), and escalate persistent failures to Senior Engineers.
- Patching Support: Assist Senior Engineers during monthly patch cycles. This includes pre-patch checks, server reboots, and post-patch validation across WSUS/SCCM (Windows) and Ansible-managed (Linux) estates.





Also assist with Azure Update Manager deployment validation and failed-patch reporting.

- VM Troubleshooting: Perform basic troubleshooting on Azure VMs (connectivity, disk, performance) and vSphere VMs (console access, snapshot management, resource contention). Independently resolve P3/P4 VM issues.
- Windows Server Administration: Basic administration including service restarts, event log analysis, disk space management, user access troubleshooting, and DNS/DHCP validation.
- Incident First Response (On-Call): During on-call periods, act as the first responder for all P1 alerts. Perform initial triage, engage vendor support if needed, communicate status updates, and escalate to Senior Engineers/Infrastructure Lead within 30 minutes if resolution is not achievable.
- Documentation &
- Knowledge Base:Maintain and update operational runbooks, known-error records, and knowledge base articles.

Contribute to process improvement initiatives.

Reports to: Head of Technical Services On-Call Commitment his role participates in a shared five-person weekly on-call rotation.

Each engineer is primary on-call for one week in every five. During on-call periods, you are expected to respond to P1 critical alerts within 15 minutes (24×7), perform initial triage and diagnostics, and escalate to Senior Engineers/Infrastructure Lead if the incident cannot be resolved within 30 minutes.

On-call compensation is provided in line with Synapse360’s standard on-call policy.

Experience

Required Technical Skills &

- Experience Technology Area Required Skill Level Linux Administration (RHEL/CentOS/Ubuntu/Oracle Linux) Proficient — package management, service management, log analysis, security hardening Ansible Proficient — playbook development and execution for patching, configuration management, and compliance Bash Scripting Proficient —
- Linux automation, log parsing, maintenance scripts, backup validation Bicep / ARM Templates Proficient — template development for repeatable Azure infrastructure deployments Azure DevOps Proficient —
- CI/CD pipelines, repos, release management for infrastructure automation Azure VMs / Azure Networking Proficient —
- Linux VM lifecycle, NSGs, UDRs, VPN Gateways Azure Update Manager Proficient — policy configuration for Linux VMs, maintenance windows, compliance reporting Azure Backup Proficient — policy management for Linux workloads, restore testing, vault monitoring Azure Key Vault Proficient — secret management, certificate rotation, integration with automated deployments Azure Arc Proficient — hybrid server governance, policy enforcement, inventory management ServiceNow Proficient — incident, change, and request management Networking (TCP/IP, VLANs, DNS) Proficient — troubleshooting, VLAN configuration, DNS resolution ITIL Processes Proficient — incident, change, problem, and release management PowerShell Awareness to Proficient —
- Azure automation, cross-platform scripting Desirable Skills

- Container fundamentals (Docker, Kubernetes awareness)
- Terraform awareness for multi-cloud IaC
- Azure Policy and Governance frameworks
- Azure Monitor and Log Analytics advanced configuration
- Git branching strategies and version control best practices
- Python scripting for automation and tooling
- Experience with LogicMonitor advanced configuration (custom datasources, escalation chains)

Education

Qualifications

Required

- Microsoft AZ-104 (Azure Administrator Associate) or equivalent demonstrated Azure experience

Desirable

- Microsoft AZ-400 (Azure DevOps Engineer Expert) or equivalent
- Ansible Fundamentals certification or equivalent training
- CompTIA Linux+ or RHCSA (Red Hat Certified System Administrator)
- CompTIA Security+ or equivalent security awareness certification

Close Date 31 Oct 2026 Please click here to read and review our Privacy Policy.

📌 Senior Infrastructure Engineer (Linux & Cloud Automation) (Stone Mills)
🏢 Manx Telecom Enterprise
📍 Stone Mills

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior infrastructure engineer (linux & cloud automation) (stone mills) / stone mills

Subscribe to this job alert:

Get the latest job offers by email for: senior infrastructure engineer (linux & cloud automation) (stone mills) / stone mills