What the lab is learning while it builds.
Positions we hold, methods we use, records we keep. A new note when the work earns one, not on a schedule, and any figure drawn from an argument rather than a measurement says so in its caption.
- 07
- notes filed
- 05
- registers
- 27
- sources verified
Nobody owned the design
What two years of evidence says about the engineer's job in AI-native development, and about the part of it we still cannot hand over.
slower than gcc -O0 on SQLite, from a compiler that passes 99% of the GCC torture suite.
- § 0602 Jul 2026
Every interviewer holds a different bar
Hiring rarely breaks at the résumé screen. It breaks the moment the standard depends on who happens to be in the room.
standardscalibrationposition6 min - § 0518 Jun 2026
Writing rubrics that survive contact with real candidates
A rubric drafted in an afternoon meets its first surprising answer within the hour. Notes on criteria that bend without breaking.
rubricsauthoringmethod8 min - § 0429 May 2026
What the examiner must never decide
An AI can hold a standard steady all night. It still should not have the last word. Where the line sits, and why it stays there.
judgmentlimitsposition5 min - § 0307 May 2026
Reading the transcript: what reasoning shows that scores cannot
Two candidates, one score, different work entirely. The number says they tied; the reasoning says which one you want.
reasoningreviewmethod7 min - § 0210 Apr 2026
Grading the grader: how we evaluate the examiner
Before Evals judges anyone, it sits its own exam: transcripts with known answers, counter-examples, and drift checks on every release.
evaluationdriftengineering9 min - § 0113 Mar 2026
Notes on instrument № 01
watchglass studies where human judgment belongs in a world that works with AI. Why the first instrument out of the lab is a hiring tool.
serieslabseries6 min