cover image
GitLab

Senior Site Reliability Engineer, Environment Automation

Remote

Canada

Senior

Full Time

08-10-2025

Share this job:

Skills

Incident Response Cloud Security GitLab Kubernetes Monitoring Ansible Networking Programming AWS Software Development SDLC GCP Terraform Prometheus Grafana Microservices

Job Specifications

GitLab is an open-core software company that develops the most comprehensive AI-powered DevSecOps Platform, used by more than 100,000 organizations. Our mission is to enable everyone to contribute to and co-create the software that powers our world. When everyone can contribute, consumers become contributors, significantly accelerating human progress. Our platform unites teams and organizations, breaking down barriers and redefining what's possible in software development. Thanks to products like Duo Enterprise and Duo Agent Platform, customers get AI benefits at every stage of the SDLC.

The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software.

An Overview Of This Role

As a Site Reliability Engineer (SRE) at GitLab, you'll help keep all user-facing services and production systems reliable, scalable, and efficient. Our SREs combine a pragmatic operations mindset with strong software engineering practices to drive automation, reduce toil, and improve resilience across our platform.

In the Environment Automation specialization, your focus is on operating and automating hundreds of GitLab environments--from initial provisioning to day-to-day maintenance tasks.

Unlike other SRE roles, this position centers on automating the lifecycle of many tenant environments, ensuring they remain secure, consistent, and reliable at scale.

Some examples of the projects you could work on:
Designing infrastructure automation that provisions and operates GitLab environments using Terraform, Ansible, and Kubernetes
Creating and maintaining deployment packages for GitLab, such as Helm Charts and omnibus-gitlab
Building and operating Dedicated GitLab instances integrated with cloud-native services (e.g., GCP, AWS)
Developing tools to orchestrate infrastructure-as-code workflows across multiple tenants
Deploying and managing microservices on Kubernetes clusters at scale
Enhancing GitLab's observability stack (e.g., Prometheus, ELK) to support proactive monitoring and incident response
Integrating with and operating infrastructure in cloud provider ecosystems (e.g., IAM, networking, storage)
Championing and implementing cloud security best practices across automated infrastructure

What You'll Do

Build & Scale Multi-Tenant Infrastructure: Design and implement automation that provisions and manages hundreds of isolated GitLab environments using Terraform, Ansible, and Kubernetes. Manage complex state strategies and workspace configurations to support scale and maintainability.
Debug & Resolve Production Issues: Troubleshoot issues across Kubernetes clusters, cloud services, and GitLab apps--identifying root causes of failed deployments, crash loops, and scheduling conflicts to ensure service continuity.
Automate Operations at Scale: Replace manual workflows with infrastructure-as-code solutions, including automated version upgrades, configuration rollouts, and provisioning pipelines that operate reliably across all tenants.
Monitor & Predict Capacity: Build observability systems that detect bottlenecks, predict usage trends, and optimize resource consumption using tools like Prometheus, ELK, and Grafana.
Respond & Lead During Incidents: Lead incident response and postmortem efforts, applying technical depth to resolve issues and establish operational standards that reduce future risk.
Architect & Collaborate: Influence architectural decisions around automation, scalability, and operational excellence. Partner with engineering teams to improve automation, platform resilience, and production-readiness.

What You'll Bring

Production-Scale Experience: Proven ability to operate and troubleshoot production workloads across multiple tenants or environments. Deep understanding of how distributed systems fail at scale and how to build in resilience.
Terraform & IaC Mastery: Strong hands-on experience with Terraform, including workspace strategies, state management, and automation patterns that scale. Comfortable solving state isolation issues and building reliable, reusable infrastructure code. Experience with Ansible and templating tools like Jsonnet is a plus.
Kubernetes in Production: Skilled at diagnosing deployment failures, interpreting pod logs, and debugging scheduling issues and rollback scenarios in live environments. Understands how pods, ReplicaSets, and controllers interact in production.
Programming & Code Analysis: Ability to read and debug

About the Company

GitLab is a complete DevOps platform, delivered as a single application, fundamentally changing the way Development, Security, and Ops teams collaborate and build software. From idea to production, GitLab helps teams improve cycle time from weeks to minutes, reduce development costs and time to market while increasing developer productivity. We're the world's largest all-remote company with team members located in more than 65 countries. As part of the GitLab team, you can work from anywhere with good internet. You'll have ... Know more