LLM Judge

Noun · AI & Machine Learning

Definitions

  1. A large language model used to evaluate, rank, or critique the outputs of another model or system according to some rubric. LLM-as-judge approaches can scale evaluation, but they need calibration because judges can inherit bias and inconsistency too.

    In plain English: A large language model used to assess other AI outputs.

    Example: "They used an LLM judge for early screening, then spot-checked disagreements with human reviewers before accepting the metric."

Related Terms