Problem
How would you design a model serving infrastructure capable of hosting a large language model like Gemini with strict sub-100ms latency requirements at Google?
Practicing as: ML Platform Engineer interview at GoogleHi, I'll play your Google interviewer for the ML Platform Engineer role. Candidates describe these interviews as mostly positive and moderately difficult, so expect me to be friendly and conversational. Take your time with the question above and answer like we're in the room.
You are practicing as a guest. Sign up free to get your answer graded with AI feedback. Your draft stays right here.


