Harness the power of AI at NVIDIA as a Senior Software Engineer focusing on AI inference systems. Architect solutions that enhance model performance in a adaptable hybrid environment. We're seeking a talented professional to lead the optimization of GPU kernels and inference frameworks like vLLM at NVIDIA.
This role involves critical contributions to benchmarking methodologies and containerized large-scale deployments, requiring deep knowledge in ML systems and performance optimization. Key Responsibilities:
- Engineer features for vLLM to leverage NVIDIA hardware
- Develop and benchmark GPU kernels with optimization techniques
- Define efficient methodologies for inference benchmarking
- Schedule and orchestrate large-scale inference on GPU clusters
- Integrate research findings into NVIDIA software products
Requirements:
- Bachelor’s degree with 7+ years or equivalent education
- Proficient in Python and C/C++, with Go or Rust knowledge
- Experience with GPU programming and ML frameworks
- Skilled in Docker and Kubernetes for containerization
- Strong debugging and communication capabilities
Step into a role where your engineering skills will have a tangible impact on AI systems at NVIDIA.
📌 Senior Software Engineer, AI Systems Optimization (Toronto)
🏢 NVIDIA Gruppe
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.