# Minimal harness checklist

AI Workshop · episode 7: Harness: the agent under control
https://mihalkevich.com/en/workshop/harness

- [ ] CLAUDE.md/AGENTS.md says how to verify the work (test and typecheck commands)
- [ ] A post-edit hook runs the linter/formatter; a pre-commit hook runs tests
- [ ] The agent works in a sandbox or dev container with no production access
- [ ] Dangerous commands need confirmation; the allowlist lives in project settings
- [ ] There's an eval set: 10 tasks with expected results, run after any model or prompt change
- [ ] Run results are recorded — you see a trend, not a single impression
- [ ] Long tasks keep a progress file: the agent writes down what's done and what's next

## Harness in this episode

This episode is about the harness itself. Every other episode has a «Harness in this episode» block — the layer we added.

## Official documentation

- [Anthropic — Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents)
- [Claude Code — hooks guide](https://code.claude.com/docs/en/hooks-guide)
- [Claude Code — sandboxing](https://code.claude.com/docs/en/sandboxing)
- [Anthropic — Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)
- [Define success and build evals](https://platform.claude.com/docs/en/test-and-evaluate/develop-tests)
- [Claude Code — dev containers](https://code.claude.com/docs/en/devcontainer)

— mihalkevich school
