What is a Site Reliability Engineer at Meta?
At Meta, the Site Reliability Engineer role is known internally as Production Engineer (PE). Production Engineers operate at the critical intersection of software engineering and systems engineering. Rather than treating operations as a reactive duty, Meta embeds PEs directly into product and infrastructure teams to ensure that massive services—such as Facebook, Instagram, WhatsApp, Messenger, and Threads—remain performant, highly available, and scalable for over three billion global users.
As a Production Engineer, your core mission is to solve complex systems challenges through code. You will build and optimize backend software, architect reliable distributed systems, manage resource utilization across multi-megawatt data centers, and automate infrastructure at hyper-scale. Whether you are debugging low-level Linux kernel bottlenecks, designing fault-tolerant file distribution systems, or writing automation scripts to handle massive traffic spikes, your work directly safeguards the foundation of Meta's physical and cloud infrastructure.
The role offers exceptional technical breadth and impact. You will collaborate closely with Software Engineers (SWEs), hardware teams, and network architects to influence service design from the ground up. Succeeding in this role requires a deep understanding of Linux system internals, networking fundamentals, distributed systems architecture, and proficient hands-on coding capabilities.

