Your question is Tokenize Text for NLP Pipelines. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are working on an NLP pipeline that starts with raw text from emails, chat logs, and support notes. Before any model can use the text, you need to split it into units that can be counted, embedded, or passed into a transformer.
What is tokenization, and why is it important in NLP pipelines?