Benchmark

Noun · Verb · AI & Machine Learning

Definitions

  1. A standardized test or dataset used to evaluate and compare AI model performance. MMLU, HellaSwag, HumanEval — the SATs of the AI world. Models are optimized for benchmarks so aggressively that benchmark scores increasingly diverge from real-world usefulness.

    In plain English: A standardized test for AI models, used to compare how well different models perform on the same tasks.

  2. Goodhart's Law applied to AI: once a benchmark becomes a target, it ceases to be a good measure. Models are increasingly 'teaching to the test,' performing well on benchmarks while failing at real-world tasks the benchmarks were supposed to predict.

    Example: 'Their model tops the MMLU leaderboard but can't follow basic multi-step instructions. Classic benchmark overfitting.'

    Source: critique / Goodhart's Law

Etymology

1842
Originally a surveyor's term — a mark cut into stone as a reference point for measuring elevation
1970s
Computing adopts benchmarks for standardized performance testing; Whetstone (1972) and Dhrystone (1984) become classics
2020s
AI benchmarks (MMLU, HumanEval, GPQA) drive the 'benchmark wars' between LLM providers, though critics question their validity

Related Terms

Collections Including This Term

Deep Dives Covering This Term