25 Sep
|
Arcadion
|
Riviere-des-Prairies—Pointe-aux-Trembles
25 Sep
Arcadion
Riviere-des-Prairies—Pointe-aux-Trembles
SLURM HPC Architect / AdministratorLocation: Remote (Canada, U.S., or Europe Preferred)Company: Cylix Applied IntelligenceEmployment Type: Full-Time or ContractAbout the RoleCylix Applied Intelligence is seeking an experienced SLURM HPC Architect / Administrator to design, deploy, and operate high-performance computing (HPC) clusters supporting AI training, large‐scale inference, scientific computing, and enterprise workloads.This role will focus on building and managing enterprise‐grade HPC environments powered by GPU and CPU compute clusters, leveraging SLURM as the core workload orchestration and resource scheduling platform.You will work closely with AI engineers, infrastructure teams, and enterprise clients to deliver scalable, reliable, and high‐performance compute environments across on‐premise, hybrid, and cloud platforms.Key ResponsibilitiesHPC Cluster Architecture and DesignDesign and implement SLURM‐based HPC cluster architecturesArchitect scalable CPU and GPU compute environmentsDefine cluster topology including compute, storage, login, and management nodesDesign high‐availability SLURM controller configurationsImplement cluster segmentation, partitioning, and resource allocation strategiesSLURM Deployment and AdministrationInstall, configure, and manage SLURM workload manager environmentsConfigure SLURM partitions, queues, QoS policies, and scheduling policiesManage job scheduling optimization and fair‐share policiesImplement accounting, usage tracking, and reporting systemsMaintain SLURM cluster health, stability, and performanceGPU Cluster and AI Infrastructure ManagementConfigure GPU scheduling and allocation policiesSupport GPU resource management including:NVIDIA A100, H100, L40,
and similar accelerator platformsMIG partitioning and GPU isolationMulti‐tenant GPU resource allocationOptimize cluster performance for AI training and inference workloadsInfrastructure Automation and OperationsAutomate cluster deployment and configuration using:Ansible, Terraform, or similar toolsShell scripting and PythonImplement monitoring, alerting, and performance tracking systemsSupport cluster lifecycle management, upgrades, and expansionStorage and Filesystem IntegrationIntegrate HPC clusters with high-performance storage systems including:NFSLustreBeeGFSGPFS / Spectrum ScaleOptimize I/O performance and storage architectureUser and Workload SupportSupport enterprise and research users with job scheduling and optimizationTroubleshoot job failures and performance issuesAssist engineering teams in optimizing workloads for HPC environmentsRequired Qualifications3+ years experience administering HPC clustersStrong experience with SLURM workload managerStrong Linux system administration experience (Ubuntu, Rocky Linux, RHEL, or similar)Experience with HPC cluster architecture and deploymentExperience with shell scripting and automationExperience with:Cluster resource managementMulti-node distributed computing environmentsSSH, networking,
and Linux system internalsPreferred QualificationsExperience managing GPU-based HPC clustersExperience supporting AI / ML workloadsExperience with NVIDIA GPU platforms and driversExperience with:CUDA environmentsNVIDIA MIG configurationGPU scheduling optimizationExperience with configuration management tools:AnsibleTerraformPuppet or ChefExperience with monitoring tools such as:PrometheusGrafanaNode exporterSLURM accounting toolsNice to HaveExperience with large-scale enterprise or cloud HPC environmentsExperience deploying HPC environments in cloud platforms such as:AWSAzurePrivate cloud environmentsExperience with containerized HPC workloads:DockerSingularity / ApptainerExperience integrating SLURM with Kubernetes or hybrid orchestration systemsExample Projects You Will Work OnDeployment of enterprise AI GPU clustersMulti-tenant SLURM cluster architecture designGPU scheduling optimization for AI workloadsHPC infrastructure for large-scale inference and model trainingHybrid HPC environments spanning data center and cloudHPC cluster performance optimization and scalingTechnology EnvironmentYou will work with:SLURM Workload ManagerLinux (Ubuntu, Rocky Linux, RHEL)NVIDIA GPU platforms (A100, H100, L40)High-performance storage systemsHPC networking (InfiniBand, high-speed Ethernet)Automation tools (Ansible, Terraform)Monitoring tools (Prometheus, Grafana)Container environments (Docker, Apptainer)What We OfferCompetitive compensationRemote-first environmentOpportunity to work with cutting-edge HPC and AI infrastructureExposure to enterprise-scale AI and compute environmentsFlexible employment structure (Full-Time or Contract)Prospect to architect next-generation HPC environmentsAbout Cylix Applied IntelligenceCylix Applied Intelligence builds enterprise AI infrastructure and high-performance computing environments supporting advanced AI workloads, intelligent automation, and enterprise-scale compute platforms.
📌 Slurm Hpc Architect / Administrator (Riviere-des-Prairies—Pointe-aux-Trembles)
🏢 Arcadion
📍 Riviere-des-Prairies—Pointe-aux-Trembles