LLM Judge
Noun · AI & Machine Learning
Definitions
A large language model used to evaluate, rank, or critique the outputs of another model or system according to some rubric. LLM-as-judge approaches can scale evaluation, but they need calibration because judges can inherit bias and inconsistency too.
In plain English: A large language model used to assess other AI outputs.
Example: "They used an LLM judge for early screening, then spot-checked disagreements with human reviewers before accepting the metric."