What is a Site Reliability Engineer at GitLab?
A Site Reliability Engineer (SRE) at GitLab sits at the intersection of software engineering and systems operations. You are not just maintaining infrastructure; you are building the automated platforms, CI/CD pipelines, and robust environments that allow millions of developers to ship code efficiently. Your work directly impacts the stability and performance of the GitLab product, ensuring that the platform remains a reliable backbone for global software development.
This role is defined by scale and complexity. Whether you are working on Dedicated Hosted Runners or Environment Automation, you are tasked with solving high-stakes problems in distributed systems, Kubernetes orchestration, and cloud-native architecture. You will be expected to treat operations as a software engineering problem, favoring automation and code-driven solutions over manual intervention.
Success in this role requires a balance of deep technical expertise and a "blameless" engineering mindset. You will collaborate with cross-functional teams to improve system observability, reduce toil, and enhance the developer experience. It is a high-visibility position that demands both the ability to deep-dive into complex production issues and the strategic thinking to prevent them from recurring.



