Your question is Diagnosing Vanishing and Exploding Gradients. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are training a deep neural network and notice unstable learning. Loss may plateau early, become NaN, or improve only in shallow layers.
How do you diagnose and mitigate vanishing or exploding gradients during the training of deep neural networks?