Visual explainer
Vanishing and Exploding Gradients
Ten layers shrink your error to a whisper or blow it sky-high. See the multiply chain and the fixes that keep it near one.
You stack ten layers hoping for smarter features. Each backward pass multiplies your error by another small number, so layer one hears a whisper while layer ten shouts. No learning rate rescues that shrink.
Multiply Derivatives Backward
Backprop applies the chain rule: one derivative per layer, multiplied in a row. Sigmoid peaks at 0.25 and sits near zero most times. So your product is a fraction times a fraction, ten times over.
Shrink Ten Times Small
Watch the focal math: start at 1.0 and multiply by 0.25 ten times. You land near 0.000001. That whisper cannot move weights, so early layers sit frozen while late layers keep learning.
Grow Ten Times Large
Flip the numbers for the mirror failure. With weights near 1.5 the product grows fast: ten steps turn 1.0 into about 57.7. Updates explode, loss jumps to NaN, and your overnight run dies.
Skips Keep Scale Healthy
Same ten layers, new wiring. A skip path adds the input back, so gradient 1 flows straight through. ReLU passes 1 for positives, and clipping caps spikes at 1.0. Both scales now sit near 1.0.
Start Simple Then Skip
Start with ReLU plus clipping at norm 1.0 since both cost little. When depth passes ten layers and early loss stalls, add skip connections. Keep plain nets only for shallow models you fully control.
Where It Breaks
Skips and clips narrow the failure, they do not erase it. Dead ReLUs pass zero forever, and wild starts still spike before clipping bites. When loss parks high or jumps, check dead units and init first.
The Quick Version
- Chain rule multiplies one derivative per layer.
- Sigmoid peaks at slope 0.25, usually less.
- Ten small steps shrink 1.0 near 0.000001.
- Ten big steps grow 1.0 near 57.7.
- Skips carry gradient 1 past shrinking layers.
- ReLU plus clipping is your default guard.
- Dead units still freeze learning entirely.