Spec Linter Lite
Drop a spec → see exactly what an AI would do with it. 3-layer validator: format, autotest, rubric.
Lint a specSimulator
Write specs that an AI can't misread. PRD-as-tests, acceptance criteria, deterministic checks.
Free · 12 runs/hour · streams in 5-10 sec
Your spec said 'helpful tone' and the agent shipped sycophancy. It said 'done' and prod broke. Two agents passed the same spec and produced wildly different work. You're grading agent output by feel, your eval rubric drifts every week, and nobody trusts the scores — so 'done' is a vibe, not a number. The spec reads fine to a human and means nothing to a model, and you find out only after the rollback.
After completing this
You rewrite vague adjectives into rubric-scored criteria and turn a PRD into failing tests plus sample data. 'Helpful' becomes three weighted checks an evaluator grades in seconds. Two agents that pass now pass for the same reason. You grade an agent's PR before reading a line of code, defend a 'this isn't done' verdict with numbers, and ship specs your contractor and your model both trust.
You don't read about spec quality — you rewrite a real vague PRD into rubric-scored tests and watch the evaluator grade it live.
You commit to your acceptance criteria, then see where a model slips through them. The discipline sticks because you got burned first.
The free Lite playground flows into the full track, where PRD-as-tests becomes a cert-worthy spec your contractor and model both trust.
Replace adjectives with weighted rubrics, so 'done' is verifiable on the spot instead of relitigated after the rollback.
Pin behavior with PRD-as-tests and frozen sample data — if two agents pass, they pass for the same reason.
Hand over an acceptance suite that runs before you read a line of code, so 'done' stops being a vibe and becomes a score.
Not for you if
You'll rewrite vague adjectives into rubric-scored criteria — so 'helpful' becomes three weighted checks the evaluator can grade in seconds.
PRD-as-tests pins behavior with failing tests + sample data. If both agents pass, they pass for the same reason.
The acceptance suite runs before you read a line of code. 'Done' stops being a vibe and becomes a number.
Weighted rubrics + frozen sample data make eval reproducible — same input, same score, week over week.
Built with the same spec-as-tests discipline it teaches.
Drop a spec → see exactly what an AI would do with it. 3-layer validator: format, autotest, rubric.
Lint a specShip like a team of one. Ralph-loop, TDD prompts, parallel worktrees, self-healing pipelines.
Start Ch1 freeYes. The Lite playground is free and needs no card. The full track is $39, or reserve a seat now to lock launch pricing.
No card for Lite. You only pay if you take the full track after the demo.
The Lite demo is one sitting. The full track is self-paced — you'll convert your own PRD into a failing-test checklist in a sitting or two.
The Lite demo bridges into the full 'Specification for AI' track — PRD-as-tests, weighted rubrics, and a final cert: a spec your team and your AI both trust.
You'll write checklists, acceptance criteria, and lightweight tests — closer to structured markdown than software. No framework required.
No card, no commitment — try it first, decide later.
Try Lite free — no signupMade by mihalkevich.com — runs on the FolderAIThe engine that powers the courses and simulators engine.