Specification for AI
Write specs that an AI can't misread. PRD-as-tests, acceptance criteria, deterministic checks.
Try Lite freeSimulator
Drop a spec → see exactly what an AI would do with it. 3-layer validator: format, autotest, rubric.
Free · 12 runs/hour · no email needed
Your spec reads clean to a human. Then the agent ignores half of it — picks the wrong output shape, leaks a PII field, fills an ambiguous acceptance criterion with whatever it felt like. You only find out after it ships, when someone asks why the agent did the opposite of what you 'clearly wrote.' The gap isn't in the model; it's between what you wrote and what it actually heard, and you have no way to see that gap before it bites.
After completing this
You stop trusting the human read and start seeing your spec through the model's eyes. Drop a spec, get three layers back — format, a live autotest on a real model, and a rubric grade — and the gap shows up in 30 seconds, line-quoted. After a few specs you write the next one already passing the linter, because you've internalized what makes an agent obey.
You don't read about spec quality — you drop a real spec and watch a real model run it back, with the failures quoted line by line.
You commit to your spec, then see exactly where it broke. The lesson lands because you already had skin in it.
Once you can read a spec like a model, the paid 'Specification for AI' track turns this into PRD-as-tests and acceptance suites your team and your AI both trust.
Catch the ambiguous acceptance criterion before an agent ships against it — and before a contractor bills you to discover it.
Triage a stack of specs from worst to best in minutes. The rubric score does the first pass so you only deep-read the salvageable ones.
See the autotest run your minimal prompt on a real model, so you fix the SOP-to-spec gap before it reaches production traffic.
Not for you if
The autotest runs the minimal prompt against a real model — you see the gap between what you wrote and what it heard, in 30 seconds.
Format + rubric flag the parts a developer would push back on — missing acceptance criteria, ambiguous output schema, leaked PII fields.
Lint each one, rank by rubric score, kill the bottom third before they reach an agent.
Used internally on every spec before it ships to a paying customer.
Write specs that an AI can't misread. PRD-as-tests, acceptance criteria, deterministic checks.
Try Lite freeUpload an existing AI folder — get the 3 most critical issues, ranked, with fix suggestions.
Sniff a folderMake the numbers honest. Interrogate metrics, audit triage agents, write reconciliation tests, brief AI that surfaces bad news instead of burying it.
Start Ch1 freeYes, fully. Lint as many specs as you want — no card, no signup.
No. Paste a spec, get the three-layer report. There's no payment step.
One spec lints in about 30 seconds. A real habit forms after you run three or four of your own.
Once you can read a spec like a model does, the paid 'Specification for AI' track turns this reflex into PRD-as-tests with weighted rubrics and acceptance suites.
It runs the minimal prompt your spec implies against a real target model, so you see the gap between what you wrote and what the model heard — not a static checklist.
Made by mihalkevich.com — runs on the FolderAIThe engine that powers the courses and simulators engine.