What is a Site Reliability Engineer at NVIDIA?
As a Site Reliability Engineer (SRE) at NVIDIA, you are at the intersection of cutting-edge hardware and massive-scale software infrastructure. You are not just maintaining servers; you are enabling the global acceleration of AI, autonomous vehicles, and high-performance computing (HPC). Your work directly impacts how NVIDIA delivers services like DGX Cloud and internal CI/CD pipelines that power the development of next-generation GPUs.
This role requires a unique blend of software engineering discipline and deep operational expertise. You will be responsible for the uptime, reliability, and scalability of complex, distributed systems. Because NVIDIA operates at such a high velocity, you will frequently bridge the gap between development teams and production environments, ensuring that new features are not only innovative but also stable and supportable at scale.




