Back to all simulators

Simulator

Specification for AI

$39Lite playable · full track on waitlist

Write specs that an AI can't misread. PRD-as-tests, acceptance criteria, deterministic checks.

Free · 12 runs/hour · streams in 5-10 sec

specification-for-ai · demoDemo

Your fuzzy idea

I want an AI that watches my Gmail and drafts replies to emails that look like sales leads. I don't want it to send anything on its own.

Result

  1. ONE-LINE INTENTa draft reply for every lead email, never sent
  2. BEHAVIOR CONTRACTMUST NEVER: send without your click
  3. EDGE CASESemail from an existing client → not a lead
  4. ONE EVAL CASEpricing request → draft asks about volume
  5. GAPSneeds decision: lead vs. newsletter
Spec skeleton · 2 decisions left to you

A sample run on demo data — this is what the result looks like in the simulator.

Want the full track ($39)? Get notified when it ships.

Coming soon · join the waitlist

Get notified when Specification for AI opens

Single email, no spam. Early signups get a launch discount.

The problem this solves

Your spec said 'helpful tone' and the agent shipped sycophancy. It said 'done' and prod broke. Two agents passed the same spec and produced wildly different work. You're grading agent output by feel, your eval rubric drifts every week, and nobody trusts the scores — so 'done' is a vibe, not a number. The spec reads fine to a human and means nothing to a model, and you find out only after the rollback.

After completing this

You rewrite vague adjectives into rubric-scored criteria and turn a PRD into failing tests plus sample data. 'Helpful' becomes three weighted checks an evaluator grades in seconds. Two agents that pass now pass for the same reason. You grade an agent's PR before reading a line of code, defend a 'this isn't done' verdict with numbers, and ship specs your contractor and your model both trust.

How it works

  1. It's a simulator, not an article

    You don't read about spec quality — you rewrite a real vague PRD into rubric-scored tests and watch the evaluator grade it live.

  2. You answer first, theory after

    You commit to your acceptance criteria, then see where a model slips through them. The discipline sticks because you got burned first.

  3. From the demo, a direct bridge to the paid track

    The free Lite playground flows into the full track, where PRD-as-tests becomes a cert-worthy spec your contractor and model both trust.

Who this is for

PMs whose AI feature shipped and immediately got rolled back

Replace adjectives with weighted rubrics, so 'done' is verifiable on the spot instead of relitigated after the rollback.

Senior engineers tired of grading agent output by feel

Pin behavior with PRD-as-tests and frozen sample data — if two agents pass, they pass for the same reason.

Founders writing the first spec a contractor will build against

Hand over an acceptance suite that runs before you read a line of code, so 'done' stops being a vibe and becomes a score.

Not for you if

  • Нет: People who want a spec template to fill in — you'll build the eval discipline, not copy a form.
  • Нет: Anyone who refuses to write a single failing test; the whole method is tests-before-prose.
  • Нет: Teams with a mature, reproducible eval harness already gating every release.

When you'll reach for this

01

Your spec says 'helpful tone' and the agent ships sycophancy

You'll rewrite vague adjectives into rubric-scored criteria — so 'helpful' becomes three weighted checks the evaluator can grade in seconds.

02

Two agents pass the same spec but produce very different work

PRD-as-tests pins behavior with failing tests + sample data. If both agents pass, they pass for the same reason.

03

You're handing a spec to a contractor and want to grade their PR honestly

The acceptance suite runs before you read a line of code. 'Done' stops being a vibe and becomes a number.

04

Your eval rubric drifts every week and nobody trusts the scores

Weighted rubrics + frozen sample data make eval reproducible — same input, same score, week over week.

Powers you walk away with

  • Да: Convert any vague PRD into a failing-test checklist in one sitting
  • Да: Write acceptance criteria that hold up under three different models
  • Да: Grade an agent's output without reading every word it produced
  • Да: Defend a 'this isn't done' verdict with numbers, not vibes
  • Да: Ship specs your contractor and your model both trust

What's inside

  1. 01PRD = markdown checklist + failing tests + sample data
  2. 02Acceptance criteria that survive non-deterministic agents
  3. 03Live evaluator with weighted rubrics
  4. 04Final cert: a spec your team and your AI both trust

Built with the same spec-as-tests discipline it teaches.

What stays in your portfolio

  • Да: The skill to convert any vague PRD into a failing-test checklist in one sitting.
  • Да: Acceptance criteria that hold up under three different models.
  • Да: The ability to grade an agent's output without reading every word it produced.
  • Да: A 'this isn't done' verdict you can defend with numbers, not vibes.

Pairs well with

Free

Spec Linter Lite

Drop a spec → see exactly what an AI would do with it. 3-layer validator: format, autotest, rubric.

Lint a spec
$99Ch1 free

Automation AI Engineer

Ship like a team of one. Ralph-loop, TDD prompts, parallel worktrees, self-healing pipelines.

Start Ch1 free

FAQ

Is the Lite demo free?

Yes. The Lite playground is free and needs no card. The full track is $39, or reserve a seat now to lock launch pricing.

Do I need a card to try it?

No card for Lite. You only pay if you take the full track after the demo.

How long does it take?

The Lite demo is one sitting. The full track is self-paced — you'll convert your own PRD into a failing-test checklist in a sitting or two.

What's next after the demo?

The Lite demo bridges into the full 'Specification for AI' track — PRD-as-tests, weighted rubrics, and a final cert: a spec your team and your AI both trust.

Do I need to write code?

You'll write checklists, acceptance criteria, and lightweight tests — closer to structured markdown than software. No framework required.

Start with the free part

No card, no commitment — try it first, decide later.

Try Lite free — no signup

Made by mihalkevich.com — runs on the FolderAIThe engine that powers the courses and simulators engine.