Mel-Spectrogram
A picture of a sound: time along one axis, perceptually-spaced frequency bands up the other, brightness for energy.
Built in three steps. Cut the waveform into overlapping windows — 25ms with a 10ms hop is the speech convention — and take a Fourier transform of each, giving a spectrogram whose frequency bins are evenly spaced in hertz. Then sum those bins into a smaller number of mel bands, narrow at the bottom and wide at the top, because human pitch perception is roughly logarithmic and 100 to 200Hz is an octave while 7,000 to 7,100Hz is inaudible. Then take a log, since loudness is perceived logarithmically too.
The window and hop lengths are a trade you cannot escape: a longer window resolves frequency better and time worse.
Because the result is a 2-D array of intensities, image models and image augmentation apply directly — SpecAugment simply masks bands of time and frequency.