AI Workshop
Episode 7 of 10Upcoming
Harness: the agent under control
What a harness is: evals, tests, hooks and a sandbox — building the frame inside which an agent can be trusted.
The link arrives by email an hour before the stream
Sign up — a day before we'll send a reminder with the episode plan, and an hour before the start the stream link. The player appears on this page 15 minutes before the start.
What we'll show
- What a harness is in plain words: everything around the model — tools, permissions, checks, memory
- Hooks: automatic lint and typecheck after every agent edit
- Tests as the definition of done: the agent can't say «done» while they're red
- Sandbox and dev container: where the agent may do anything, and where nothing
- An eval set: 10 typical project tasks that show whether a model or prompt change made things better or worse
What you take away
Your own minimal harness: hooks, tests, a sandbox and a 10-task eval set.
Minimal harness checklist
- CLAUDE.md/AGENTS.md says how to verify the work (test and typecheck commands)
- A post-edit hook runs the linter/formatter; a pre-commit hook runs tests
- The agent works in a sandbox or dev container with no production access
- Dangerous commands need confirmation; the allowlist lives in project settings
- There's an eval set: 10 tasks with expected results, run after any model or prompt change
- Run results are recorded — you see a trend, not a single impression
- Long tasks keep a progress file: the agent writes down what's done and what's next
Harness in this episode
This episode is about the harness itself. Every other episode has a «Harness in this episode» block — the layer we added.
What to have ready
- A project with at least one test (if not, we'll write one)
- Docker, if you want to repeat the sandbox
