AI Workshop
Episode 7 of 10Upcoming

Harness: the agent under control

What a harness is: evals, tests, hooks and a sandbox — building the frame inside which an agent can be trusted.

Wednesday 18 November at 19:00 Moscow time· 90 min

The link arrives by email an hour before the stream

Sign up — a day before we'll send a reminder with the episode plan, and an hour before the start the stream link. The player appears on this page 15 minutes before the start.

What we'll show

  1. What a harness is in plain words: everything around the model — tools, permissions, checks, memory
  2. Hooks: automatic lint and typecheck after every agent edit
  3. Tests as the definition of done: the agent can't say «done» while they're red
  4. Sandbox and dev container: where the agent may do anything, and where nothing
  5. An eval set: 10 typical project tasks that show whether a model or prompt change made things better or worse

What you take away

Your own minimal harness: hooks, tests, a sandbox and a 10-task eval set.

Minimal harness checklist

  • Да: CLAUDE.md/AGENTS.md says how to verify the work (test and typecheck commands)
  • Да: A post-edit hook runs the linter/formatter; a pre-commit hook runs tests
  • Да: The agent works in a sandbox or dev container with no production access
  • Да: Dangerous commands need confirmation; the allowlist lives in project settings
  • Да: There's an eval set: 10 tasks with expected results, run after any model or prompt change
  • Да: Run results are recorded — you see a trend, not a single impression
  • Да: Long tasks keep a progress file: the agent writes down what's done and what's next

Harness in this episode

This episode is about the harness itself. Every other episode has a «Harness in this episode» block — the layer we added.