Skip to content
AI360Xpert
Beta
Core ML

Visual explainer

Vanishing and Exploding Gradients

Ten layers shrink your error to a whisper or blow it sky-high. See the multiply chain and the fixes that keep it near one.

Ten shrinking multiplications freeze your early layers.
Ten shrinking multiplications freeze your early layers.

You stack ten layers hoping for smarter features. Each backward pass multiplies your error by another small number, so layer one hears a whisper while layer ten shouts. No learning rate rescues that shrink.

Multiply Derivatives Backward

The chain rule multiplies one derivative per layer.
The chain rule multiplies one derivative per layer.

Backprop applies the chain rule: one derivative per layer, multiplied in a row. Sigmoid peaks at 0.25 and sits near zero most times. So your product is a fraction times a fraction, ten times over.

Shrink Ten Times Small

Ten sigmoid steps shrink a gradient of 1.0 to about 0.000001.
Ten sigmoid steps shrink a gradient of 1.0 to about 0.000001.

Watch the focal math: start at 1.0 and multiply by 0.25 ten times. You land near 0.000001. That whisper cannot move weights, so early layers sit frozen while late layers keep learning.

Grow Ten Times Large

Ten growth steps of 1.5 turn a gradient of 1.0 into about 57.7.
Ten growth steps of 1.5 turn a gradient of 1.0 into about 57.7.

Flip the numbers for the mirror failure. With weights near 1.5 the product grows fast: ten steps turn 1.0 into about 57.7. Updates explode, loss jumps to NaN, and your overnight run dies.

Skips Keep Scale Healthy

A skip path plus clipping keeps gradients near 1.0 on the same scale.
A skip path plus clipping keeps gradients near 1.0 on the same scale.

Same ten layers, new wiring. A skip path adds the input back, so gradient 1 flows straight through. ReLU passes 1 for positives, and clipping caps spikes at 1.0. Both scales now sit near 1.0.

Start Simple Then Skip

Default to ReLU with clipping, and add skips when depth passes ten layers.
Default to ReLU with clipping, and add skips when depth passes ten layers.

Start with ReLU plus clipping at norm 1.0 since both cost little. When depth passes ten layers and early loss stalls, add skip connections. Keep plain nets only for shallow models you fully control.

Where It Breaks

Dead units and wild starts still freeze or blow up your training.
Dead units and wild starts still freeze or blow up your training.

Skips and clips narrow the failure, they do not erase it. Dead ReLUs pass zero forever, and wild starts still spike before clipping bites. When loss parks high or jumps, check dead units and init first.

The Quick Version

  • Chain rule multiplies one derivative per layer.
  • Sigmoid peaks at slope 0.25, usually less.
  • Ten small steps shrink 1.0 near 0.000001.
  • Ten big steps grow 1.0 near 57.7.
  • Skips carry gradient 1 past shrinking layers.
  • ReLU plus clipping is your default guard.
  • Dead units still freeze learning entirely.