Your question is Design a Real-Time TTS Backend. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are building the backend for a text to speech product that turns user text into audio in real time. Users expect fast playback, natural sounding voices, and consistent quality across many devices and surfaces.
Design a highly scalable backend architecture to support real-time text-to-speech synthesis for millions of active users.