Stop Pressing 1: measuring oversight of AI agents
2025 to present. Independent research.
When a person approves an AI agent’s action, did they check what the agent will actually do? I call the failure “the reading problem” and treat it as a measurement question. Oversight fidelity is the match between what a person predicts an agent will do and what it actually does.
- Instrument
- A browser-only, fully scripted lab with no live model. Participants review six agent action requests across a code repository, an inbox, and an operations console. Three are seeded with real incident classes: a hidden
.envedit that sends logs to an outside endpoint, a reply that quietly copies hundreds of recipients, and a typosquatted dependency. Evidence of each problem stays on screen the whole time. - Design
- Three between-subjects conditions: passive confirmation (press 1), a generative gate where participants predict the agent’s action before approving, and a quiz written by the agent itself. Responses are scored with signal detection measures, d′ and criterion c.
- Hypotheses
-
- Under the generative gate, accuracy on ground-truth questions reaches 70% or higher.
- Students who grew up with AI assistants and recent adopters both improve under the gate, with a larger gain for the first group.
- The agent-written quiz performs like passive confirmation: verification theater.
- Outputs
- Poster at the DEF CON 34 AI Village (August 2026) and a talk at DC State of the Stack (September 2026). The current Stop Pressing 1 Lab puts the gate into practice: students build a small Python tool with a simulated coding agent and answer questions about each change before they can approve it (source).
- Next
- A classroom pilot, and a comprehension-gate service that coding agents can call before a person approves an action.