What is a Site Reliability Engineer at Lambda?
As a Site Reliability Engineer at Lambda, you are at the heart of building and scaling one of the world's most powerful AI cloud platforms. You are responsible for ensuring the high availability, performance, and scalability of infrastructure that powers cutting-edge machine learning research and production workloads. This is not a traditional maintenance role; it is a high-impact engineering position where you will design and implement the systems that allow Lambda to remain a leader in GPU-accelerated computing.
You will work on critical domains such as Fleet Management, Managed Kubernetes, Storage, Software Defined Networking (SDN), and Observability. Your work directly impacts how researchers and enterprises deploy massive AI models, meaning your technical decisions have immediate, large-scale consequences. You will collaborate with cross-functional engineering teams to automate operations, reduce toil, and build robust platforms that handle extreme computational demands.



