How Proof of Quality works.
You define the rubric. Human experts establish the reference judgments. AI judges compete to match them, with consensus requirements tailored to your dataset.
No single reviewer determines the outcome.
Watch how it works ↓See the review process.
Read the film transcript
This is a silent illustrative film. The judgments assess supplied label quality against the customer’s rubric. Robot task success, label quality and reviewer agreement are different things. Votes, thresholds and alignment scores are illustrative, not measured customer results. Accurately labelled unsuccessful episodes can still be useful.
0:00 to 0:12 · Opening
Every episode needs a judgment. One episode. The complete sequence. Your data grows. Review becomes the bottleneck.
Visual action. The camera pulls back from a robot workcell to an episode represented by its complete temporal sequence, then to a growing dataset. This introduces the human review burden as data volume grows.
0:12 to 0:24 · Expert reference
You define the rubric. EXAMPLE: INSERTION SUCCESS Success means it stays fully seated after the gripper releases. REFERENCE R011 · LABEL: SUCCESS Experts review independently. Their consensus sets the reference. You set the consensus rule. Expert consensus: 3 of 3 Shared examples and expert reasoning guide judges.
Visual action. Reference episode R011 carries the supplied label Success. One large evidence view shows approach, insertion, release and the period after release. The part remains fully seated. Three experts independently apply the criterion and each gives Pass. Their consensus forms the reference. The customer sets the human consensus rule. Shared examples and expert reasoning guide the judges.
0:24 to 0:36 · Judge alignment
Unseen examples test alignment. Expert answer withheld Judges compete to match your experts. TESTED AGAINST UNSEEN EXPERT ANSWERS Example threshold: 99% alignment Illustrative scores Top qualified judges are elected.
Visual action. Distinct withheld episode Q028 is used for alignment testing. Its expert answer is hidden from the AI judges. Six illustrative alignment scores are shown: Judge 1, 99.4%; Judge 2, 87.2%; Judge 3, 99.2%; Judge 4, 99.6%; Judge 5, 96.8%; Judge 6, 93.1%. Judges 1, 3 and 4 qualify and are elected. These are example scores, not measured product performance.
0:36 to 0:45 · AI review
Every elected judge reviews every item. P041 · SUPPLIED LABEL: SUCCESS Required consensus: 3 of 3 You set the consensus threshold. AI agreement: 2 of 3
Visual action. Production episode P041 has the supplied label Success. Every elected judge reviews this same complete temporal episode. After the gripper has withdrawn, the part backs out and is no longer fully seated. Required consensus of 3 of 3 is visible before the votes. Judge 1 gives Pass, Judge 3 gives Pass and Judge 4 gives Fail. Their agreement is 2 of 3, below the configured threshold, so the item remains unresolved.
0:45 to 0:56 · Human escalation
Insufficient consensus returns to experts. PENDING HUMAN CONSENSUS No verdict until experts agree. SAME P041 · SUPPLIED LABEL: SUCCESS Experts review independently. Human consensus resolves the item. Success label rejected: not fully seated. LABEL QUALITY Fail HUMAN CONSENSUS 3 of 3
Visual action. The exact same production episode P041 and supplied Success label return to human experts. They replay the complete evidence and each gives Fail. Their individual votes and roles then fade away. The final display keeps one LABEL QUALITY verdict, Fail, separate from HUMAN CONSENSUS, 3 of 3. The visible reason is "Success label rejected: not fully seated." The quality verdict concerns that supplied label. An unsuccessful robot episode can still be useful when accurately labelled.
0:56 to 1:04 · Blind checks
Random expert checks continue. B063 · SUPPLIED LABEL: SUCCESS Expert answer kept hidden Hidden from the AI judges. Expert consensus: Fail · 3 of 3
Visual action. A distinct later random episode, B063, carries the supplied label Success. Its part backs out after release. Judges 1, 3 and 4 independently give Fail. Only after every AI vote does the film reveal the human reference to the audience: Fail with consensus 3 of 3. The expert answer remains hidden from the judges. This sampled check does not certify neighbouring items.
1:04 to 1:18 · Continuing competition
Judges that drift are replaced. Later human checks detect drift. Qualified challengers can take a place. J8 · Qualified and faster Same standard. Faster review. Quality first. Speed before cost. Your dataset sets the priorities.
Visual action. Later human checks show Judge 3 drifting below the required alignment standard. Qualified Judge 7 takes its place. Qualified challenger Judge 8 then replaces Judge 1 through faster review while meeting the same standard. Judge 4 remains. The dataset sets the performance priorities. The customer priority statement does not establish a fixed universal ranking algorithm.
1:18 to 1:22 · Close
Scale review. Keep expert judgment. Illustrative workflow
Visual action. The closing statement sits beside the dataset and holds at full brightness through the final frame.
Judges keep earning their place.
Elected judges are not permanent. Competition continues in the background, and every judge must keep meeting the agreed standard.
Alignment slips.
Blind spot checks show a judge falling below the agreed standard. It is removed and replaced by a qualified judge.
A challenger performs better.
A new candidate meets the alignment requirement and performs better on objectives you agree, such as speed or cost. It can take an incumbent’s seat. Cheaper or faster never bypasses the quality bar.