Senior DevOps Engineer HPC / EDA / SLURM- #26-24242 (British Columbia)

Senior DevOps Engineer HPC / EDA / SLURM- #26-24242 (British Columbia)

27 Sep
|
Fives DyAG
|
British Columbia

27 Sep

Fives DyAG

British Columbia

Duration: 12 Months Contract
Job Description

Client is seeking a Senior DevOps Engineer supporting our High Performance Computing (HPC) and Electronic Design Automation (EDA) infrastructure team. This role will work directly within the IT Datacenter (ITDC) organization and is expected to operate independently at a senior level with minimal ramp-up time. The ideal candidate brings strong hands-on experience with Linux HPC environments, infrastructure automation, SLURM workload management, datacenter migration execution, and enterprise identity integration.

Responsibilities
HPC / EDA Platform Operations

Support and administer SLURM-based HPC compute environments, including partition configuration and migration planning

Plan and execute HPC/EDA compute and storage infrastructure migrations across datacenters

Develop migration strategies and evaluate implementation options, risks, dependencies, and operational tradeoffs

Author formal Method of Procedure (MOP) documents and runbooks for infrastructure changes and service cutovers

Coordinate cross-functionally with EDA/SPE teams, storage teams, and IDAM to deliver coordinated platform changes

Verify storage volumes, application access, and service continuity following migrations or infrastructure changes

Define HPC storage service tiers and gather performance and capacity requirements for EDA workloads

Linux Systems Engineering & OS Deployment

Administer SUSE Linux Enterprise Server (SLES) 12 and SLES 15 systems in a production HPC environment

Build and maintain custom Linux OS images and installation media using Kiwi NG and related tooling

Enable and maintain bare-metal provisioning workflows via RackN / Digital Rebar Provision for SLES and ESXi deployments

Provision and configure VMware vSphere virtual machines for HPC service workloads

Troubleshoot production issues in Linux HPC operational scripts, services, and system daemons

SLURM > SLES > RackN

Automation & Infrastructure as Code

Develop, maintain, and extend Ansible playbooks and roles for Linux system setup, authentication, and platform configuration

Ensure multi-version Ansible playbook compatibility across SLES 12 and SLES 15

Manage Git repositories and Artifactory artifact storage; migrate large binaries and configuration artifacts out of source control





Contribute GitHub pull requests, conduct code reviews, and manage inner-source infrastructure repositories

Drive production environment changes through change management workflows using ServiceNow

Identity & Access Management

Integrate and configure enterprise identity systems including Okta, Active Directory, LDAP, NIS, VAS, and SSSD for Linux/HPC environments

Currently using NIS and VAS, but experience with any are OK

Audit and reconcile Linux user and group identity data (UID/GID) across multiple directory and authentication domains

Validate authentication methods and access behavior across HPC compute and storage environments

Extend SSSD-based corporate authentication to new compute environments and author corresponding Ansible automation

Monitoring, Logging & Operational Readiness

Assess and implement log management strategies, including evaluation of Splunk integration for HPC system logs

Investigate and remediate operational issues in production Linux services (VNC, NIS, AutoFS, Zabbix, etc.)

Produce technical documentation, architecture diagrams, implementation guides, and end-user instructions in Confluence

Experience

5+ years of experience in a DevOps, Platform Engineering, or Linux Systems Engineering role

Hands-on HPC cluster administration experience, including SLURM or equivalent workload managers

Demonstrated experience supporting EDA or scientific computing environments

Strong Ansible automation skills with production-grade playbook and role development

Experience with bare-metal provisioning tools (RackN, Cobbler, or equivalent)

Proven ability to plan and execute datacenter or infrastructure migrations with minimal disruption

Familiarity with enterprise Linux identity and authentication stacks (SSSD, LDAP, AD, NIS, Okta)

Experience with NetApp or comparable enterprise storage platforms in HPC contexts





Ability to author formal technical documentation (MOPs, runbooks, architecture diagrams)

Strong written and verbal communication skills; capable of coordinating across multiple teams.

Core Technical Skills

HPC / EDA Platforms: SLURM, HPC compute/storage administration, EDA infrastructure, datacenter migrations

Linux / OS: SLES 12, SLES 15, ESXi 8.0, Kiwi NG ISO creation, Linux system services, acct/pacct, dracut

Provisioning / Automation: Ansible (playbooks, roles, multi-version), RackN / Digital Rebar Provision, VMware vSphere

Identity / Auth: SSSD, Okta, Active Directory, LDAP, NIS, VAS, UID/GID auditing, cross-domain identity management

Storage / Filesystems: NetApp SVM, NFS, AutoFS, RootSquash, storage tier design, IOPS/capacity planning

DevOps / Source Control: Git, GitHub, Artifactory, inner-source repository management

Monitoring / Logging: Splunk integration, Zabbix, operational script hardening, log management

Scripting / Languages: Ansible (YAML), Python, Perl (debugging), Bash

ITSM / Documentation: ServiceNow (change requests), MOP authoring, Confluence, Jira, technical diagramming.

Preferred Qualifications

Experience with SUSE Linux Enterprise Server (SLES) 12 and/or 15 in an enterprise setting

Familiarity with RackN / Digital Rebar Provision for bare-metal OS deployment

Hands-on experience with Kiwi NG or similar tools for custom OS image creation

Knowledge of VMware vSphere for HPC support VM provisioning

Experience migrating configuration artifacts and binaries to Artifactory

Background in semiconductor, storage, or high-tech manufacturing IT environments.

Education

Bachelors Degree

About US Tech Solutions
US Tech Solutions is a global staff augmentation firm providing a wide range of talent on-demand and total workforce solutions. To know more about US Tech Solutions, please visit www.ustechsolutions.com.

US Tech Solutions is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

AI Statement
By applying, you acknowledge that AI-assisted tools may be used during hiring.

#J-18808-Ljbffr

📌 Senior DevOps Engineer HPC / EDA / SLURM- #26-24242 (British Columbia)
🏢 Fives DyAG
📍 British Columbia

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior devops engineer hpc / eda / slurm- #26-24242 (british columbia) / british columbia