Your question is Transformer vs RNN Architecture. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
Explain the architecture of a Transformer model and how it differs from an RNN.
In a practical interview answer, describe the main Transformer components, explain how a token is processed through the model, and compare training, inference, context handling, and parallelism with an RNN. Relate the explanation to modern LLM engineering, including latency, memory, and hallucination implications where relevant.