Join the Data Infrastructure team at •••••• as a Software Engineer focusing on large-scale storage solutions. Contribute to developing a unified storage layer for AI model training.
In this key role, you'll design and build a distributed storage system for model training and evaluation at petabyte scale. Collaborate with researchers and infrastructure teams to address challenges in data reading, writing, and movement. Your work will directly influence performance metrics like GPU utilization, throughput, and consistency.
Key Responsibilities:
• Design and operate distributed storage systems for AI models
• Manage Kubernetes clusters handling petabyte-scale data
• Collaborate with teams to optimize data throughput and latency
• Resolve consistency and I/O issues with large datasets
• Enhance system reliability and performance for AI workloads
Requirements:
• Solid storage fundamentals, including data lifecycle management
• Experience coding in Python and/or Go
• Familiarity with Kubernetes and stateful systems
• Background in cloud object storage (e.g., S3)
• Knowledge of parallel filesystems is a bonus
Drive innovation in AI storage solutions while building your expertise in an evolving technology landscape.
#J-18808-Ljbffr
📌 Data Infrastructure Engineer at Tech Firm (Ontario)
🏢 Talanto
📍 Ontario
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.