AI judge
Human judgment
An AI system configured to evaluate an item against specified criteria.
A judge can assess labels, outputs or other evidence using a rubric. Its capability should be tested against relevant expert judgments. Multiple judges provide additional observations, but shared models or prompts can produce correlated mistakes. A judge's own confidence is not a substitute for measured performance.
In practice
A judge checks whether a task success label is supported by the full episode rather than generating a new robot action.