Welcome to your interview.
The question is on your right: Vanishing Gradients in Deep Networks. Take a moment with it first.
Talk your thinking through with me if you like - when you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes). Discussion and graded submissions share your five interviewer interactions, so spend them well.
You are training a deep neural network and notice that early layers learn very slowly as depth increases. You want to understand why optimization becomes difficult and which architectural choices make training stable.
Describe the vanishing gradient problem in deep neural networks. How do residual connections, batch normalization, and specific activation functions help resolve it?