Benchmarks Glossary
Browse 8 benchmarks terms defined in plain English, from the cultural dictionary of computing.
8 Benchmarks Terms
- AI Safety Benchmark
- A benchmark designed to measure model behavior on harmful, risky, deceptive, or policy-sensitive tasks.
- benchmark contamination
- The phenomenon where a model's training data inadvertently (or deliberately) includes examples from evaluation benchmarks, inflating its apparent performance...
- Benchmarketing
- The practice of using benchmarks more as marketing theater than as honest performance evaluation. In tech slang, benchmarketing implies selective tests,...
- Benchmark Saturation
- A situation in which benchmark scores become so high that the benchmark no longer meaningfully distinguishes between systems or predicts real-world usefulness....
- Eval Suite
- A collection of automated tests, tasks, or scoring scripts used to evaluate model quality across multiple behaviors.
- GLUE Benchmark
- A benchmark suite of language understanding tasks used to compare NLP model performance. It influences how models are trained, evaluated, or served, and it can...
- MMLU Benchmark
- A benchmark that measures broad multitask knowledge and reasoning across many academic and professional domains.
- Model Benchmark
- A benchmark used to compare models on one or more tasks, capabilities, or operational metrics. Model benchmarks help guide selection, but they should be...