1. What is a Machine Learning Engineer at Cerebras?
As a Machine Learning Engineer at Cerebras, you will operate at the intersection of groundbreaking hardware and advanced artificial intelligence systems. Cerebras builds the world's largest AI chip—the Wafer-Scale Engine, which is significantly larger than traditional GPUs—and provides the extreme compute power necessary for ultra-high-speed generative AI training and inference. In this role, you contribute directly to building and scaling the software and hardware infrastructure that empowers top model labs, global enterprises, and AI-native startups to run massive machine learning workloads without the bottleneck of managing hundreds of disparate GPUs.
Your work directly impacts the core product capabilities that define Cerebras Inference and state-of-the-art training platforms. Whether you are optimizing low-level kernel performance, bringing up foundational open-source models like LLaMA and Qwen, or building automated observability platforms, your contributions ensure that users experience industry-leading throughput and low latency. You will collaborate closely with cross-functional teams spanning compiler development, hardware design, runtime engineering, and product teams to translate complex model architectures into high-performance execution on custom silicon.
The role demands a system-minded generalist who thrives in fast-paced bring-up environments and feels comfortable navigating the entire software stack. You will tackle unique challenges related to graph lowering, compiler optimizations, performance benchmarking, and end-to-end reliability. Expect a rigorous, high-impact environment where your engineering solutions directly redefine the operational boundaries of large-scale machine learning.



