02 Aug
|
huaweicanada
|
Vancouver
02 Aug
huaweicanada
Vancouver
About the team
The Advanced Computing and Storage Lab, part of the Vancouver Research Centre, explores adaptive computing system architectures to address future flexible and variable application loads. It helps ensure stability and quality of training clusters, constructs dynamic cluster configuration strategy solvers, and develops precision control systems to create stable and efficient computing power clusters. The lab focuses on industry AI application scenarios such as large model training and inference, using technologies like low-precision training, multi-modal training, and reinforcement learning to analyze bottlenecks and design optimization solutions that improve training and inference performance and usability.
About the job
- Aim to advance performance, efficiency, and usability of AI systems on the Ascend platform for key industry AI scenarios such as large model training and inference. Responsibilities include low-precision training, multimodal optimization, reinforcement learning, and training resource optimization to address system bottlenecks and deliver next-generation AI capabilities.
- Design and develop optimization solutions for AI training and inference systems, focusing on FP8 optimization, RL-driven training agents, multimodal reinforcement learning, or next-generation multimodal understanding and generation.
- Combine AI algorithm requirements with system-level architectural optimization in computing, I/O, scheduling, and precision control to improve performance.
- Build stable, efficient AI training clusters using dynamic configuration and precision control to ensure scalability and reliability.
- Develop software frameworks, operator libraries, acceleration libraries, and system-level optimizations for NPU platforms to accelerate large-model AI training.
- Drive innovation in optimizing large-model training and inference with low-precision training, parallel strategy tuning, and reinforcement learning.
- Keep up with the latest research in AI computing cluster architecture design, training acceleration, and inference acceleration to strengthen competitiveness of AI computing cluster systems.
The target annual compensation (based on 2080 hours per year) ranges from $56,000 to $79,000 depending on education, experience, and demonstrated expertise.
About the ideal candidate
- Currently pursuing a degree in Computer Science, Computer Engineering, or related fields (artificial intelligence, software, automation, electronics, communications, robotics, etc.).
- Familiar with common model structures of large models (e.g., Deepseek and Llama) and have basic experience in large-model training and inference optimization in LLM, MoE, multimodality, etc.
- Familiar with hardware architecture and programming systems of AI accelerators (e.g., GPUs/NPUs) and experience optimizing AI systems with coordinated software and hardware cores.
- Assets include any of the following:
- Solid programming foundation in Python, C, or C++, good architectural design and programming habits.
- Ability to work independently, good communication, willingness to cooperate, eagerness to learn recent technologies, and ability to summarize and share; hands-on practice.
- Experience in developing AI training frameworks and AI reasoning engines, or algorithm hardware-related experience.
- Strong research capabilities in new technologies and architectures, ability to track and gain insights into cutting-edge AI technologies, and leadership in system architecture innovation.
Additional Information:
Huawei Canada is committed to a fair, inclusive, and accessible recruitment process. If you require accommodation during any stage of the hiring process, please let us know and we will work with you to meet your needs.
All applications for this position are reviewed directly by our hiring team; we do not use artificial intelligence tools to screen or select candidates.
#J-18808-Ljbffr
📌 Co-op Researcher - AI Computing System (Vancouver)
🏢 huaweicanada
📍 Vancouver