Your question is Design an LLM Serving Platform. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are designing the serving stack for a product that uses large language models in production. Different requests have different latency, cost, and safety needs, and the system must handle traffic from multiple user segments.
How would you design an LLM serving system that balances latency, cost, scalability, and safety?