03 Oct
|
Axelon Services
|
Montreal
03 Oct
Axelon Services
Montreal
Experience Level: Level 3 (senior): minimum 5 years
Duration: 12 Months contract position
Location: Montreal
Work Mode: Onsite (Day 1 onboarding onsite/in office presence 3x/week)
Responsibilities:
- Design, build, and operate the AI Gateway's Azure and AWS deployments, transitioning from proof of concept to production.
- Develop and extend Python services (FastAPI / Flask) for the Gateway's inference, onboarding, and administrative APIs.
- Integrate recent model providers and model families, including Azure AI Foundry / Azure OpenAI and AWS Bedrock, covering request signing, streaming responses, failover, and quota handling.
- Implement cloud-native authentication and secrets handling — Entra ID with Managed Identity and workload federation, AWS IAM roles, and STS — aiming to eliminate stored credentials.
- Build and evolve the entitlement and authorization data layer across SQL Server and PostgreSQL, including schema changes, migrations, and data-correctness controls.
- Manage platform controls for governance: rate limiting, token accounting, content guardrails, audit logging, and chargeback reporting.
- Deploy and run services on Kubernetes (on-premises, AKS, and EKS) using Helm, GitOps, and Terraform, and maintain CI/CD pipelines (Jenkins, GitHub Actions).
- Develop observability tools to track requests post-factum — metrics, logs, and dashboards across Prometheus, Grafana, Loki, and Snowflake.
- Collaborate with cloud platform, network, and security teams on connectivity, egress policy, network controls, and architecture review, and provide necessary evidence for reviews.
- Support production: participate in on-call, investigate incidents, and implement fixes and hardening into the code.
- Write tests and documentation as part of delivery and review peers' changes.
Requirements:
- Strong, production-grade Python, including a web framework — FastAPI or Flask — and a real testing discipline.
- Hands-on experience with Kubernetes: deploying, configuring, and troubleshooting workloads.
- Practical understanding of OIDC / OAuth 2.0: token validation, JWKS, client-credentials flows, claim, and audience handling.
- Microsoft Azure experience in at least three of the following: AKS, Entra ID, Azure OpenAI or Azure AI Foundry, Key Vault, Azure Database for PostgreSQL, Azure Cache for Redis, Azure Monitor.
- Amazon Web Services experience in at least three of the following: IAM and STS / assume-role, SigV4 request signing, Bedrock, EKS, VPC endpoints and private networking, Secrets Manager, CloudWatch.
- Experience with Infrastructure as code — Terraform, Bicep, or CDK — and CI/CD with Jenkins or GitHub Actions.
- Proficiency in SQL and relational data modeling, including schema migrations.
- Clear written and verbal communication skills, and the ability to work directly with security, network, and platform teams.
Preferred Skills:
- Experience building or operating an API gateway, reverse proxy, or multi-tenant platform.
- LLM platform engineering specifics: streaming and server-sent events, token accounting, prompt and response guardrails, model evaluation.
- Experience with Kafka and Snowflake for audit and consumption data pipelines.
- Observability depth: Prometheus and PromQL, Grafana, Loki, OpenTelemetry.
- Advanced Redis or Valkey use beyond basic caching — counters, TTLs, distributed rate-limiter semantics.
- Experience delivering in a regulated enterprise environment with corporate proxies, private networking, and strict change control.
This role is for an existing vacancy.
📌 AI Platform Engineering Specialist (Montreal)
🏢 Axelon Services
📍 Montreal