Lab · PLANNED
Agent Field Test
Can an agent execute, verify, and document real work?
Layers under test
The test
Give an agent a real task with a real deadline and a real definition of done, then judge it on the finished work rather than on the transcript.
What gets measured
-
Task completion against a specification written before the run.
-
Whether the agent verified its own output, and how it detected failure.
-
The quality of the record it left behind for a human to audit.
-
Where a human had to intervene, and why.
Status
[PLACEHOLDER] Method is being finalised. The specification is published before the run so the result cannot be graded retroactively.