Software Engineer – Inference Serving (Toronto)

Software Engineer – Inference Serving (Toronto)

08 Sep
|
Taalas
|
Toronto

08 Sep

Taalas

Toronto

At Taalas we believe that fundamental progress is achieved by those who are willing to understand and assail a problem end-to-end, without regard for commonly accepted abstractions and boundaries.

We are building a team of hands-on technologists who dislike overspecialization and seek to excel in both depth and breadth.

In this position the successful candidate will build software infrastructure for an inference serving cluster built around Taalas hardcore AI model chips.

Job Responsibilities

- Adapt open-source inference servers like vLLM and Punica to interface with Taalas’ hardcore AI models
- Implement a highly productive LoRA swapping solution for multi-{tenant,LoRA} environments




- Build and test a scalable inference serving cluster using K8 and Traefik or similar

Qualifications

- Bachelor’s or higher degree in Computer Science, or Electrical/Computer engineering
- Experience with K8, HTTP load balancers, web-servers
- Good knowledge of computer architecture and low-level programming: Linux virtual memory and page table management, direct memory access, CUDA
- Familiarity with ML, Python and Pytorch

Interested in joining our team? Submit your resume to [email protected] to be considered for the exciting opportunity!

📌 Software Engineer – Inference Serving (Toronto)
🏢 Taalas
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: software engineer – inference serving (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: software engineer – inference serving (toronto) / toronto