16 Sep
|
RiverMeadow - Workload Mobility Platform
|
Canada
16 Sep
RiverMeadow - Workload Mobility Platform
Canada
We are looking for a highly experienced Senior Operations Engineer to join our team and support the automation, security, reliability, and operational stability of our infrastructure, CI/CD pipelines, and infrastructure automation.
You will work as part of a 3-person Operations team. This role is intended for a strong hands-on engineer with deep experience in Linux systems, virtual machines, networking, cloud platforms, automation, CI/CD, and production operations. The ideal candidate should be able to work independently, design and deploy new services, troubleshoot complex issues, improve operational processes, and maintain reliable infrastructure environments.
Key Responsibilities
- Maintain, operate, and improve production infrastructure environments in AWS.
- Administer Linux-based systems, services, and networking components across traditional data center components and public clouds.
- Support and maintain CI/CD pipelines for development teams and build automation for infrastructure operations.
- Support cloud and on-premises virtualization platforms.
- Improve infrastructure security, system hardening, monitoring, logging, and alerting practices.
- Investigate and resolve production incidents, including root cause analysis and follow-up improvements.
- Work closely with the Operations team, team lead or manager, and engineering teams to support application deployment, release processes, and operational requirements.
- Prepare and maintain clear infrastructure documentation, operational procedures, and runbooks.
Required Skills and Experience
Core Infrastructure and Systems
- Strong Linux administration and troubleshooting experience.
- Experience installing, configuring, and maintaining production-grade infrastructure environments in a data center.
- Experience deploying, maintaining, and monitoring production environments in public clouds.
- Deep understanding of networking concepts and protocols, including TCP/IP, DNS, VPN, SSL/TLS, HTTP/HTTPS, routing, firewalls, and load balancing.
- Ability to analyze system behavior, diagnose performance issues, and resolve complex infrastructure problems.
Security and Reliability
- Strong understanding of infrastructure security, hardening, and operational best practices.
- Experience with monitoring, alerting, and centralized logging systems.
- Experience handling production incidents in a structured and responsible manner.
- Ability to improve system reliability, availability, and operational stability over time.
Cloud, Virtualization, and Platforms
We use the following platforms for our applications, both to support our production environment and for building platform features to support customer migrations to and from cloud and on-prem virtualized environments. Our core product runs in AWS,
but it enables VM migrations to and from public and private clouds.
Public Cloud, core IaaS, minimal PaaS
- Amazon Web Services
Private Cloud
- VMware
Automation and DevOps
- Demonstrated drive to identify and automate repetitive infrastructure and operational tasks.
- Experience supporting CI/CD pipelines in a development environment.
- Experience with configuration management tools such as Chef.
- Solid scripting and automation skills using Bash, Python, or similar languages.
- Ability to automate repetitive infrastructure, deployment, and operational tasks.
- Understanding of release management and environment promotion workflows.
Development Workflow
- Strong Git and GitHub knowledge.
- Understanding of up-to-date DevOps workflows, branching strategies, pull requests, and code review processes.
- Experience working with development teams in production-oriented environments.
Nice to Have Experience The following items are part of the job, but can be learned by the right candidate
- Infrastructure as Code experience with tools such as Terraform, Ansible, or similar.
- Experience with high-availability and scalable production systems.
- Knowledge of observability stacks such as ELK or OpenSearch.
- Experience in fast-paced production environments.
- Experience with backup, disaster recovery, and incident response processes.
- Google Cloud Platform
- Microsoft Azure
- Red Hat OpenShift Virtualization for hosting VMs
- OpenStack
- HPE Morpheus Essentials with HPE VM Essentials hypervisor
- Azure Local
- Nutanix
- Microsoft Hyper-V
📌 Senior Operations Engineer (Canada)
🏢 RiverMeadow - Workload Mobility Platform
📍 Canada