The first draft of every rubric is a wish list. Strong fundamentals, clear communication, good judgment under ambiguity: words that feel precise in a document and dissolve on contact with an actual answer. The candidate does something you did not anticipate, and the rubric has no opinion.
The rubrics that hold up describe evidence, not virtues
Name the observable, not the quality
Not communicates clearly but states their assumption before committing to it. The second one can be found in a transcript; the first one can only be felt.
Say what separates adjacent scores
The hard call is never between a two and a five. It is between a three and a four, and a rubric that does not describe that boundary has delegated it back to whoever is reading.
Name what does not count
This is where most private bars leak in. Pace, accent, confidence, familiarity with your stack: if they are not part of the job, write down that they are not part of the score.
Stress-test the rubric the way you test code
We have started running rubrics against answers designed to embarrass them. The brilliant solution delivered with hostile communication. The wrong answer reached by excellent reasoning. The candidate who says almost nothing and is right. If two readers of the rubric score those differently, it is the rubric that needs the work, not the readers.
A criterion that every reader applies the same way is finished. Everything else is still a draft.
None of this requires an AI. It is simply what writing the bar down forces you to confront. The examiner only makes the forcing function impossible to skip, because it will apply exactly what you wrote, to everyone, without the silent repairs a human reader performs on a vague criterion.
watchglass (2026). Writing rubrics that survive contact with real candidates. The notebook, volume № 01, § 05. watchglass.app.