11 Sep
|
GuestTek Interactive Entertainment
|
Calgary
11 Sep
GuestTek Interactive Entertainment
Calgary
This is not a greenfield project.
We already run a working platform: the TIP OpenWiFi (uCentral) cloud SDK plus an in-house multi-tenant layer above it (Organisation -> Site -> Device model, canonical device catalogue, Keycloak/OIDC role-based access control, audit trail) and a React operator console. Real access points, switches and site gateways are provisioned end to end in our lab today. You are joining to take that from a working prototype to a supportable production service, and to extend it onto hardware and customers it does not yet support.
What the platform does not yet have is an architecture that survives being handed to a customer.
Three problems are open and they are yours: device identity has no working renewal path, so certificates expire and devices cannot recover without someone physically reaching them; there is no proven failover story for the equipment we put on a customer site; and the scale ceiling of the control plane is estimated rather than measured. This is a hands-on architecture role. You are expected to write the decision down, then build the thing you decided.
Key Requirements:
- Identity and certificate lifecycle. Device identity is mutual TLS against our own certificate authority. Certificates are currently issued with a short lifetime, the device agent has no in-band renewal command, and a device that expires while offline cannot re-enrol on its own. Own the design and the implementation of a renewal and recovery path that does not depend on physical access, plus expiry as a monitored fleet-wide signal.
- Runtime architecture. Device connections are raw TLS over TCP, not HTTP. They terminate on a Kubernetes LoadBalancer service with local external traffic policy to preserve source addresses. An HTTP ingress controller in this path is incorrect and breaks mutual TLS. Own the service topology,
storage placement across external Ceph and node-local NVMe, and the failure domains.
- Site gateway architecture. Customer sites terminate on a small x86 appliance running our own OpenWrt-based image: wide area network, guest network, RADIUS and policy enforcement. Own the redundancy model - a failover pair is the target - and the remote recovery path for an appliance nobody can reach.
- Delivery and environment promotion. Helm chart management for the upstream deployment charts plus our own, delivered through GitOps with a real promotion path between environments, and a documented delta against upstream that stays rebaseable.
- Observability that pages. Service level objectives, and alerting that reaches a human before a customer does. Dashboards nobody is paged from do not count.
- Scale envelope. Establish a defensible concurrent-device ceiling by measurement, name the first bottleneck, and re-measure it as the platform changes.
Required experience:
- Production on-premises Kubernetes built and operated with kubeadm or equivalent - not managed cloud only.
- Certificate authority operations at a practical level: issuing policy and lifetimes, renewal, revocation, and what happens to a device that misses its window. HashiCorp Vault or comparable.
- Bare-metal load balancing with MetalLB or Cilium, including BGP or layer 2 modes, and source address preservation for non-HTTP services.
- Ceph as a consumer: block storage CSI configuration, storage class tuning,
and diagnosing storage-induced application timeouts.
- Helm at an authoring level, and GitOps-based delivery with environment promotion.
- Prometheus, Grafana and Loki, with alerting tied to service level objectives rather than to raw thresholds.
- The ability to write an architecture decision down and defend it in review. We will ask to read something you wrote.
Valuable but not required
- TIP OpenWiFi deployment charts, or Kafka operations.
- High availability at the network edge: VRRP or keepalived, multi-WAN, and remote recovery of unattended appliances.
- Capacity planning for telemetry-heavy workloads.
- Experience operating a service that other people's hardware depends on.
What success looks like in the first 90 days
- A device certificate renewal path that is implemented, not just documented, including a stated recovery route for a device that expires while offline.
- The southbound service path running in production with source address preservation verified end to end.
- A measured storage latency baseline for the relational tier, with a go or no-go call on placement.
- Alerting that has caught at least one real incident before a user reported it.
What We Offer
At GuestTek, you won’t just have a job, you will have the prospect to build your career while working with innovative technology and a global team.
- Competitive compensation and comprehensive benefits
- Opportunities for career growth and professional development
- Exposure to innovative technology, AI, cybersecurity, and global projects
- Collaborative and supportive work environment
- Opportunities to work with teams and customers around the world
- Challenging projects that make a real impact
- A culture that values innovation, teamwork, and employee contributions
📌 Senior Cloud / Platform Architect (Calgary)
🏢 GuestTek Interactive Entertainment
📍 Calgary