29 Aug
|
Tekishub Consulting Services
|
Canada
29 Aug
Tekishub Consulting Services
Canada
Role: Architect - Platform Engineer
Experience Level: 10+ yrs
Work Location: 100% Remote -East/Canada [ET & CT]
Role Overview:
We are looking for a highly skilled Architect - Platform Engineer to design, optimize, and scale infrastructure for GenAI and LLM workloads. This role is ideal for someone with deep hands-on experience in GPU profiling, distributed training, and high-performance compute environments.
You’ll play a key role in building out GenAI platform foundations, supporting production-grade deployments, and partnering closely with data science, MLOps, and application teams to bring cutting-edge AI solutions to life.
Key Responsibilities:
● Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments
● Perform GPU profiling, benchmarking, and performance optimization for distributed training workloads
● Manage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environments
● Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS, etc.)
● Collaborate with cross-functional teams to deploy models in research and production environments
● Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps)
● Develop reusable infrastructure templates using tools like Terraform and Helm
● Contribute to internal innovation (PoCs, workshops) and support client-facing delivery engagements Basic Qualifications:
● Strong experience with Slurm and distributed training environments
● Hands-on expertise with Red Hat OpenShift and/or Kubernetes
● Deep knowledge of the NVIDIA GPU ecosystem (CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT)
● Robust foundation in Linux systems, performance tuning, and multi-GPU optimization
● Experience deploying GenAI workloads (LLM fine-tuning, RAG pipelines, multi-modal systems)
● Familiarity with Infrastructure-as-Code tools (Terraform, Ansible)
● Experience with cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters
Other Qualifications (OQs):
● Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containers
● Knowledge of LLMOps frameworks and MLOps integration
● Familiarity with vector databases and retrieval systems for RAG architectures
● Comfortable working in client-facing environments and collaborating with AI solution teams
Pay: $140,000.00-$150,000.00 per year
Experience:
- NVIDIA GPU: 1 year (required)
- AI/ML: 1 year (required)
Work Location: Remote
📌 Platform Engineer (Canada)
🏢 Tekishub Consulting Services
📍 Canada