Employment Status: Permanent
Schedule: 40 hours/week – 100% remote work
Job Description
We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms.
Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infrastructure supporting critical cloud-based services. This role combines hands-on engineering with operational leadership, giving you direct ownership of system availability, scalability, and incident response.
You will be involved throughout the service lifecycle, from architecture and launch preparation to production monitoring and continuous improvement.
Responsibilities
- Partner with engineering teams during system design, capacity planning, launch readiness, and production deployment.
- Monitor and improve service availability, latency, performance, and overall system health.
- Identify recurring operational issues and implement sustainable solutions that improve scalability and resilience.
- Define and evolve observability practices, including dashboards, alerts, SLOs, and SLIs.
- Build and maintain automated infrastructure using Terraform and CI/CD pipelines.
- Develop automation and operational tooling to reduce manual intervention and support self-healing systems.
- Coordinate incident response and act as Incident Commander during critical production events.
- Facilitate blameless post-incident reviews and ensure that corrective actions are completed.
- Use AI-assisted engineering tools responsibly to accelerate development and operational workflows.
- Maintain clear technical documentation and contribute to the continuous improvement of SRE practices.
Required profile
- Significant experience in Site Reliability Engineering, Cloud Engineering, DevOps, or a similar infrastructure-focused role.
- Experience supporting complex or large-scale SaaS environments with high availability requirements.
- Strong hands-on knowledge of AWS services and architecture, including multi-account environments, VPC, EC2, and EKS.
- Proven experience operating and troubleshooting Kubernetes environments at scale.
- Strong knowledge of Infrastructure as Code, particularly Terraform.
- Experience building or maintaining CI/CD pipelines using GitLab, Jenkins, or comparable tools.
- Experience with enterprise observability platforms such as Datadog, Prometheus, Grafana, or equivalent solutions.
- Strong scripting skills using Python, Bash, or a similar language.
- Direct experience participating in on-call rotations, coordinating incident response, and conducting post-incident reviews.
- Familiarity with Java or .NET application environments is considered an asset.
- Ability to communicate clearly and collaborate with development, infrastructure, security, and operations teams.
- Must be legally authorized to work in Canada.
What to Expect
- Remote-first work environment within Canada.
- Occasional visits to a local office or participation in in-person meetings may be required, representing less than 10% of the role.
- Participation in a scheduled on-call rotation is required.
Does this opportunity sound like a good fit for you? Apply now through our website or by sending your resume to e.henry@totemtalent.ca.
Thank you for your interest in this position; only candidates who meet our client’s requirements will be contacted.
The masculine gender is used as a neutral form.
#totemtech


