Benchmark

Noun · Verb · AI & Machine Learning

Definitions

  1. A standardized test or dataset used to evaluate and compare AI model performance. MMLU, HellaSwag, HumanEval — the SATs of the AI world. Models are optimized for benchmarks so aggressively that benchmark scores increasingly diverge from real-world usefulness.

    In plain English: A standardized test for AI models, used to compare how well different models perform on the same tasks.

  2. Goodhart's Law applied to AI: once a benchmark becomes a target, it ceases to be a good measure. Models are increasingly 'teaching to the test,' performing well on benchmarks while failing at real-world tasks the benchmarks were supposed to predict.

    Example: 'Their model tops the MMLU leaderboard but can't follow basic multi-step instructions. Classic benchmark overfitting.'

    Source: critique / Goodhart's Law

Etymology

1842
Originally a surveyor's term — a mark cut into stone as a reference point for measuring elevation
1970s
Computing adopts benchmarks for standardized performance testing; Whetstone (1972) and Dhrystone (1984) become classics
2020s
AI benchmarks (MMLU, HumanEval, GPQA) drive the 'benchmark wars' between LLM providers, though critics question their validity

Related Terms