Your question is Cross-Entropy vs MSE Gradients. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You're training a supervised model and comparing common loss functions for prediction tasks. You want to understand not just how the losses are defined, but how they change optimization behavior during training.
Explain the mathematical difference between Cross-Entropy and Mean Squared Error loss functions, and how do they affect gradient updates?