Your question is Compare Transformers and LSTMs. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are comparing two sequence models used in NLP, one based on recurrent state updates and one based on self-attention. Both can be used for text classification, tagging, and language modeling, but they learn and process context in different ways.
What are the basic machine learning concepts behind transformers and LSTMs?