Skip to content
Glossary

Calibration set for AI judges

Human judgment

Expert reviewed examples used to compare judges before selecting them for a review task.

In this context, calibration means evaluating alignment with expert judgments. It is distinct from calibrating predicted probabilities. Keep the reference answers hidden from the judges being tested and use the same evaluation rules for each candidate. Reusing the set for extensive tuning can weaken its value as independent evidence.

In practice

Candidate judges answer the same examples, and their decisions are scored against the expert reference.