LLM Benchmark

Noun · AI & Machine Learning

Definitions

  1. A benchmark used specifically to evaluate large language models on tasks such as reasoning, coding, factuality, tool use, or safety. LLM benchmarks are useful for comparison, but they can mislead if they stop reflecting real product needs.

    In plain English: A benchmark used to measure how well a large language model performs.

    Example: "They added an internal LLM benchmark because public leaderboards did not capture the tone and citation requirements of their support workflow."

Related Terms