Your question is Vanishing Gradients in Deep Networks. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You are training a deep neural network and notice that early layers learn very slowly as depth increases. You want to understand why optimization becomes difficult and which architectural choices make training stable.
Describe the vanishing gradient problem in deep neural networks. How do residual connections, batch normalization, and specific activation functions help resolve it?