Skip to content
AI360Xpert
Glossary
Definition

Vanishing Gradients

Gradients shrinking toward zero as they propagate backward through many layers, leaving early layers with almost no signal to learn from during training.

Each layer's backward pass multiplies the incoming gradient by a local derivative, and when those derivatives are consistently smaller than one, the product shrinks exponentially with depth. Twenty layers with a typical factor of 0.5 leaves the earliest layer with roughly a millionth of the gradient the output layer sees.

The usual causes are saturating activations like sigmoid and tanh, poor weight initialization, and plain depth. Residual connections, normalization layers, and ReLU-family activations are the three standard fixes, each attacking a different piece of the multiplication.