23 Aug
|
CloudFactory
|
Canada
23 Aug
CloudFactory
Canada
- As a Senior SRE, you will design and build scalable infrastructure, working closely with cross-functional teams to develop systems and pipelines that support the automation, reliability, and scalability of our production environments
- You will bring a high degree of autonomy to designing recent infrastructure components and applying site-reliability practices across our systems, while communicating complex technical issues clearly to stakeholders across the business
- This is an exciting opportunity to make a real impact while working alongside talented people from developing nations
- Please note: This is a full-time, fixed-term employee position with an expected duration of 6 months
- Infrastructure design and automation:
- Design and implement new core infrastructure components with a high degree of autonomy
- Optimize and improve existing systems and operations, such as deployment pipelines, environment provisioning, and high-throughput batch jobs
- Use Infrastructure as Code (IaC) tools, such as Terraform, to manage and scale complex infrastructure
- CI/CD automation:
- Develop CI/CD pipelines to automate build, test, deployment, and monitoring processes
- Create and manage multi-step CI/CD pipelines, including environment setup and artifact handling
- Reliability and availability:
- Support the reliability, availability, and performance of production systems,
applying site-reliability practices across the infrastructure
- Set up monitoring, alerting, and observability tooling to maintain visibility into system health
- Collaboration and communication:
- Collaborate closely with software engineers, product, and business stakeholders on the design and delivery of infrastructure and deployment systems
- Communicate complex technical issues clearly to stakeholders from technical and non-technical backgrounds alike- Knowledgeable about cloud platforms such as GCP or AWS
- Experience using CI/CD platforms to automate build, test, and deployment pipelines
- Experience with Docker and Kubernetes
- Fluent in Python, with strong experience writing production-ready code
- Comfortable applying site-reliability principles, such as availability, observability, and automation, across production systems
- 5+ years of experience building and operating infrastructure in production environments
- Experience using Infrastructure as Code (IaC) tools such as Terraform
- Degree in Computer Science, Engineering, or another quantitative or computational field, or equivalent practical experience
- Familiarity with monitoring tools such as Prometheus or Grafana
- Exposure to multi-cloud or hybrid-cloud environments
- Experience with configuration management tools (e.g., Ansible, Chef, Puppet)
📌 Senior Site Reliability Engineer (Canada)
🏢 CloudFactory
📍 Canada