What is a Site Reliability Engineer at NVIDIA?
As a Site Reliability Engineer (SRE) at NVIDIA, you are at the intersection of cutting-edge hardware, massive-scale AI infrastructure, and mission-critical software services. You are not just maintaining systems; you are architecting the reliability of the platforms that power the next era of computing, from DGX Cloud and Omniverse to autonomous driving and deep learning research. Your work ensures that NVIDIA’s global engineering teams have the high-availability environments they need to innovate at speed.
This role is highly strategic and technically demanding. You will manage complex on-premise and cloud-based infrastructure, implement sophisticated automation to eliminate toil, and define the observability standards that keep NVIDIA’s services resilient. Whether you are optimizing CI/CD pipelines for GPU development or scaling network operations for data centers, your impact is measured by your ability to turn infrastructure into a competitive advantage. You will thrive here if you are an engineer who enjoys solving deep, systemic problems under pressure and values a culture of "blameless" operations.




![[24]7.ai logo](https://storage.googleapis.com/company-logos-bucket/logos/247ai.png)