Benchmark Contamination
A catastrophic evaluation failure where the exact test questions used to evaluate a language model were accidentally included in its massive training corpus.
Think of It Like This
Like a student stealing the final exam answer key a week before the test; their perfect score proves they can memorize, not that they actually learned the material.
When a model achieves superhuman scores on a reasoning test, it is often because it simply memorized the answers during pre-training rather than actually learning to reason. This forces researchers to constantly invent new, secret benchmarks or use dynamic evaluation methods to get a true measure of a model's zero-shot generalization capabilities.