Benchmark
Noun · Verb · AI & Machine Learning
Definitions
A standardized test or dataset used to evaluate and compare AI model performance. MMLU, HellaSwag, HumanEval — the SATs of the AI world. Models are optimized for benchmarks so aggressively that benchmark scores increasingly diverge from real-world usefulness.
In plain English: A standardized test for AI models, used to compare how well different models perform on the same tasks.
Goodhart's Law applied to AI: once a benchmark becomes a target, it ceases to be a good measure. Models are increasingly 'teaching to the test,' performing well on benchmarks while failing at real-world tasks the benchmarks were supposed to predict.
Example: 'Their model tops the MMLU leaderboard but can't follow basic multi-step instructions. Classic benchmark overfitting.'
Source: critique / Goodhart's Law
Etymology
- 1842
- Originally a surveyor's term — a mark cut into stone as a reference point for measuring elevation
- 1970s
- Computing adopts benchmarks for standardized performance testing; Whetstone (1972) and Dhrystone (1984) become classics
- 2020s
- AI benchmarks (MMLU, HumanEval, GPQA) drive the 'benchmark wars' between LLM providers, though critics question their validity