06 Sep
|
HCLTech
|
Montreal
We are looking for a Senior Storage Site Reliability Engineer (SRE) to join a high-performing infrastructure team responsible for designing, automating, and maintaining large-scale enterprise storage platforms. This role is ideal for someone with strong Linux administration, storage technologies, and automation expertise who enjoys solving complex reliability and performance challenges.
Hybrid work - 3 days in office in Montreal
What You'll Do
✅ Manage and support enterprise storage environments, including GPFS and NAS technologies
✅ Automate operational processes and infrastructure tasks using Python
✅ Monitor, troubleshoot, and improve storage platform reliability and performance
✅ Support Linux-based environments and distributed storage systems
✅ Collaborate with network, cloud, and infrastructure teams to resolve complex issues
✅ Drive continuous improvement through automation, tooling, and operational excellence
What We're Looking For
✔ Strong experience with Linux Administration and/or NAS technologies (File Systems, SMB)
✔ Knowledge of Software Defined Storage and Cloud technologies
✔ Experience with storage and Unix protocols including NFS, SMB, Fibre Channel, and iSCSI
✔ Solid Python scripting skills for automation and tooling development
✔ Solid understanding of TCP/IP networking, DNS, CIFS, and related technologies
✔ Experience supporting large-scale infrastructure in production environments
✔ Strong troubleshooting and problem-solving skills
Top Skills
? GPFS (IBM Storage Scale)
? Linux Administration
? Python Automation
? NAS Storage (NFS/SMB)
? Software Defined Storage (SDS)
? TCP/IP Networking
? Site Reliability Engineering (SRE)
Nous recherchons un(e) Ingénieur(e) principal(e)
SRE Stockage pour rejoindre une équipe infrastructure performante responsable de la gestion, de l'automatisation et de la fiabilité de plateformes de stockage d'entreprise à grande échelle. Ce rôle convient parfaitement à une personne passionnée par Linux, les technologies de stockage et l'automatisation.
En mode hybride - 3 jours en présentiel au bureau à Montréal
Responsabilités
✅ Gérer et soutenir les environnements de stockage d'entreprise, incluant GPFS et les technologies NAS
✅ Automatiser les processus opérationnels et les tâches d'infrastructure à l'aide de Python
✅ Surveiller, diagnostiquer et optimiser la performance et la fiabilité des plateformes de stockage
✅ Assurer le support des environnements Linux et des systèmes de stockage distribués
✅ Collaborer avec les équipes réseau, infonuagique et infrastructure
✅ Contribuer à l'amélioration continue grâce à l'automatisation et aux meilleures pratiques SRE
Compétences recherchées
✔ Excellente expérience en administration Linux et/ou technologies NAS (systèmes de fichiers, SMB)
✔ Connaissance du stockage défini par logiciel (SDS) et des technologies Cloud
✔ Maîtrise des protocoles de stockage et Unix : NFS, SMB, Fibre Channel, iSCSI
✔ Excellentes compétences en Python pour l'automatisation et le développement d'outils
✔ Bonne compréhension des technologies TCP/IP, DNS, CIFS et des réseaux d'entreprise
✔ Expérience dans des environnements critiques à grande échelle
✔ Fortes capacités d'analyse et de résolution de problèmes
Compétences clés
? GPFS (IBM Storage Scale)
? Linux Administration
? Python Automation
? NAS Storage (NFS/SMB)
? Software Defined Storage (SDS)
? TCP/IP Networking
? Site Reliability Engineering (SRE)
📌 Storage Site Reliability Engineer (GPFS / Storage Ops) | Ingénieur(e) SRE Stockage (GPFS / Opérations de stockage)
🏢 HCLTech
📍 Montreal