Reference for your agent. The precise shape of things — signatures, fields,
rules — written so an agent can read it and act. You do not need to: ask your
agent for what you want in words, and if it needs this page, hand it the URL or
use Copy page. What to ask for, and how to check it, is in Ask;
how to see what changed is in Look inside.
scenarios and ends up in the same runner, which is how a
case is tried before it becomes a file.
Cases are independent of one another, a chat and a run each, so they run in
parallel. The scenarios inside one case run in order, because they share a
conversation.
A case
user is one turn or a list of them; a list is sequential turns on the same chat,
each starting after the previous run reached a terminal state. trials and
timeout default from defaults and are overridable per scenario.
A case in the tree carries nothing else. suite:, agent:, tool:,
harness_id:, machine_id: and files: all describe a throwaway harness to test
and the computer to test it on, and a case in the tree tests the harness the tree
is, wherever that tree is already running. Each of them is refused by name, with
the reason.
What a scenario may assert
A trial passes when every assertion of its scenario passes. There is no partial
credit inside a trial.
hook_called is the one that pays for itself fastest. A hook that agrees with the
platform and a hook that has been timing out all week produce the same conversation;
hook_called: {name: tools.before, decision: deny} is a check that the guard is
actually guarding.
Two numbers, three verdicts
A scenario is runtrials times, and the two numbers that come out of it are the
ones worth knowing:
The verdict is those two together: pass when both hold, flaky when it passed
at least once but not always, fail when no trial passed.
Flaky is its own word on purpose. A check that passes four times in five is not a
check that passes, and one trial cannot tell you which of the three you have — so
trials: 1 can only ever report pass or fail, and a case that matters is worth
running three or five times.
Scoring is a pure function of the artifact, so evals/results/ from last month can
be scored again by today’s rules without running anything.
A grader
Assertions match text and calls. Anything that needs judgment is a grader: Python in the tree, called by name from a case, running in the sandbox that holds the tree — which is the whole point, because it is handedt, and so t.llm: a judge model
called through the platform, on the platform’s key, against the platform’s
accounting.
<file>.<function> — quality.answers above — the way a tool is
<file>__<function>, and the name is read back to the first dot, so neither the
file nor the function may carry one. A case calling a judge the tree does not hold
is refused at publish time, not discovered on a run.
It is handed one scenario and one trial of it, not the file:
True/False, or {"pass": bool, "score": float, "why": str}. Anything
else is malformed and is recorded as an error, not as a decision: counting an
unreadable answer as a pass is worse than not grading at all. Same for a grader that
raised or ran over its 30 seconds — longer than a hook gets, because a judge is
allowed to call a model.
pass_below in the case is the score the case passes at: the trial fails when the
grader answered below it, and fails too when it answered no score at all. The
grader’s own pass still has to hold either way.
A judge is a file in a checkout, and a checkout is on a machine. Any case whose
expect names a grader needs machine_id — one of the harness owner’s machines
that runs this harness. harness_eval without one runs the assertion-only cases
fine.Reading the result
harness_eval answers one line per scenario — case, scenario, PASS or FAIL, the
first failure in the words the assertion used, how long it took — and the path of
the full artifact.
The artifact holds every trial: what was said, every tool call, every hook decision,
every verdict. When a case fails, the first failure tells you what broke; the
artifact tells you why.
Two follow-ups are usually the fastest:
harness_hook_tracefor the same run, if the failure smells like a hook. It says whether your file answered, deferred, failed or was never asked.- The artifact’s tool calls, if the failure is a
tool_calledthat did not happen. An agent that never called the tool usually was not told it had one — check thetools=entry and whether it was narrowed after a#.
Checking a harness that is not this one
harness_eval with no harness_id is about the harness being run, the checkout at
~/harness. Passing harness_id points it at another harness of the same owner —
one this machine does not carry, which is how you evaluate a harness you just built.
A harness this account does not own is refused.

