An instrument you cannot calibrate is a rumor. Before any version of the examiner reaches a candidate, it runs against a bench of transcripts with agreed-upon scores, interviews scored independently by people who argued until they converged. The examiner's job is to land inside that agreement, and to show its reasoning when it does not.
The bench is built to be hostile
Brilliant answers delivered badly. Confident nonsense. Candidates who fish for hints, and candidates who go quiet. Each release is measured on the full bench, and a regression on any slice blocks it. A model that got better on average and worse on quiet candidates does not ship.
An average is a place for a regression to hide. We grade the slices.
Drift gets the same treatment on a slower clock
Rubric language ages, roles change shape, and an examiner tuned last quarter meets this quarter's candidates. We re-run old transcripts under new versions and diff the reasoning, not just the scores. A stable number reached by a different argument is drift wearing a disguise.
None of this makes the examiner infallible. It makes it inspectable, which is the property we actually promise. The lab's oldest rule applies to its own instruments first: show the working.
watchglass (2026). Grading the grader: how we evaluate the examiner. The notebook, volume № 01, § 02. watchglass.app.