LLM Benchmark
Noun · AI & Machine Learning
Definitions
A benchmark used specifically to evaluate large language models on tasks such as reasoning, coding, factuality, tool use, or safety. LLM benchmarks are useful for comparison, but they can mislead if they stop reflecting real product needs.
In plain English: A benchmark used to measure how well a large language model performs.
Example: "They added an internal LLM benchmark because public leaderboards did not capture the tone and citation requirements of their support workflow."
Related Terms
- LLM-as-Judge
- AI Metric
- AI Performance
- AI Testing
- Human Baseline
- LLM Cache
- LLM Calling
- LLM Completion
- LLM Config
- LLM Gateway
- LLM Judge
- LLM Memory
- LLM Observability
- LLM Optimization
- LLM Pipeline
- LLM Platform
- LLM Plugin
- LLM Proxy
- LLM Reliability
- LLM Router
- LLM Scale
- LLM SDK
- LLM Security
- LLM Streaming
- LLM Throughput
- LLM Token
- LLM Trace
- LLM Wrapper
- Model Eval
- Prompt Evaluation
- Safety Evaluation
- Truthfulness