About the role You will develop low-level code to maximize the efficiency of machine learning workloads on our custom RISC-V hardware. This position focuses on tuning compute kernels to ensure optimal performance at the instruction and memory levels. We are hiring across multiple seniority levels to join our Toronto-based engineering team.
Key facts Location: Toronto, Ontario, Canada
Engagement: Hybrid
Team: Acceleration Kernel Development
What you'll do Write and optimize compute kernels for parallel machine learning tasks.
Profile and tune performance to address latency, memory, and bandwidth constraints.
Collaborate with machine learning engineers to integrate kernel optimizations into production pipelines.
Maintain and debug the low-level software stack to ensure reliability and speed.
Work alongside hardware engineers to push the limits of our custom architectures.
Requirements Proficiency in C and C++ for building high-performance, efficient code.
Experience in developing and tuning kernels for parallel or high-performance computing.
Ability to analyze and improve instruction-level code execution.
Solid problem-solving skills and a focus on precision in software development.
Skills & tools C/C++
Parallel algorithms
Performance profiling and debugging
Machine learning workload optimization
RISC-V architecture concepts
Practical notes Compensation ranges from $100,000 to $500,000, including base salary and variable components, determined by experience, education, and location.
Employment is contingent upon your eligibility to access U.S. export-controlled technology. You may be required to provide proof of citizenship or permanent residency, or obtain a license from the U.S. Commerce Department, to comply with U.S. Export Administration Regulations. Offers may be rescinded if export control requirements cannot be met.