Your question is Serving Multiple Fine-Tuned LLMs. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
Design a scalable system for serving multiple fine-tuned LLMs with strict latency SLOs and varying traffic loads.
Clarify the workload, model sizes, request patterns, latency targets, availability requirements, and deployment environment before proposing an architecture. Address routing, batching, autoscaling, GPU utilization, online and batch inference, evaluation, observability, cost controls, and failure recovery.