Dataford
Interview QuestionsInterview GuidesExperiencesMock InterviewsPricing
Get started

Design Batch and Online Ad Serving

Medium
MediumSystem DesignFeature StoreModel ServingRecommendation Systems

Problem

Scenario

You are designing the ML serving stack for an ad ranking system in a large social app. The system must score ad candidates for feed requests while balancing relevance, revenue, and user experience. Some predictions can be precomputed in batch, while others must be generated in real time from fresh user context. The business goal is to improve ad performance without violating tight latency budgets on the main feed.

Scale

SignalValue
DAU180M
Peak feed request QPS900K
Ad catalog25M active ads
Candidates scored per request2,000 retrieved -> 150 ranked -> 10 returned
p99 latency budget120ms end-to-end
New/updated ads per day3M
User feature freshness target< 5 minutes

Question

How would you design the end-to-end ML system, including what to serve from batch versus online inference, and how those choices affect retrieval, ranking, feature computation, evaluation, monitoring, and failure handling at this scale?

Take this as a live interview session →

You are practicing as a guest. Sign up free to get your answer graded with AI feedback. Your draft stays right here.

Sign up freeI have an account
Sign up to unlock solutions
Next questions
ElevenLabsDesign Online and Batch ML ServingHardExperis NederlandOnline vs Batch Model ServingMediumDesign Sharded Ad Click PredictorHard