Evals, instrument № 01

Know how your next hire actually works.

Evals is an AI examiner. It interviews every candidate on the real work of the job, the debugging, the trade-offs and the judgment, then files the reasoning behind every score. You write the bar once and share a link.

the brief

You set the bar. The AI holds it.

Write the evaluation once: the rubric it scores against, how long the interview runs, and what the role actually requires. Plain language is the whole specification. From there the standard stops depending on who is in the room, or on what time it is.

Isometric technical plate: a height gauge clamped at zero drift holds the bar constant, met by stacked blocks for role, duration, and rubric, with the brief taped to the surface plate
Fig. 1Hiring breaks when every interviewer holds a different bar. Write the standard once and hold it fixed for everyone.
the run

Every candidate, the same exam.

Each candidate opens a link and meets a live examiner: the same questions, the same standard, whatever the hour. Scored results land as they finish, not weeks later.

Isometric technical plate: a footbridge spans from solid ground to new ground on anchored supports, its deck built from repeatable modules, with criteria — safe, sound, true to the path — on a note taped beside it
Fig. 2Interviews go uneven: different questions, different moods, different hours. Holding them level is the point.
the record

Scores that show their working.

Every score is filed with the reasoning that produced it, criterion by criterion, next to the full transcript. Read why before you trust it, and overrule it when you know better. The instrument measures; you decide.

Isometric technical plate: a rubric drawn as an exploded stack of layers, one per criterion, each layer's thickness set by its weight, descending onto a base plate of colored probe cubes marked role and duration
Fig. 3A score without its reasoning is a black box. Read the working, then keep the last word.
the specification

What the instrument handles.

Table 1The seven functions of the instrument, in running order: create, run, judge, watch.
  1. 01examiner

    It runs the whole interview

    The questions, the follow-ups, the scoring, end to end, with no engineer in the room.

  2. 02authoring

    You write it in plain language

    Describe the role and the bar in a paragraph. That paragraph is the whole specification.

  3. 03fit

    It is cut to the role

    Every evaluation is built to the seat you are filling, not drawn from a question bank.

  4. 04range

    Difficulty follows the candidate

    It presses harder when someone cruises and offers a hint when they stall, so a score is a reading rather than a ceiling.

  5. 05integrity

    It notices what does not add up

    Pasted answers, coached responses, work that does not match the conversation: flagged on every run.

  6. 06judgment

    It scores criterion by criterion

    Each answer is graded against your rubric, with the reasoning filed beside the score.

  7. 07record

    Nothing happens off the record

    Follow every candidate from the moment the link opens to the moment the score is filed.

Herbarium plate 27A: a pressed Asteraceae specimen taped to a field-note sheet, with capitulum, involucre, and leaf studies, measured dimensions, and handwritten collection notes
the series

One instrument at a time.

watchglass studies where human judgment belongs in a world that works with AI. It publishes an instrument when the instrument becomes dependable, because every one added before the last is trusted subtracts from all of them.

  1. № 01
    EvalsAI-run technical evaluations that file their reasoning with every score.
    In use
  2. № 02
    UntitledIn the workshop. It ships when it stops surprising us, and not before.
    In preparation
  3. № 03

A standard nobody can state is not a standard. Write it once, hold it for everyone, and keep the last word.

Open evals