Back to all simulators

Simulator

Spec Linter Lite

Free

Drop a spec → see exactly what an AI would do with it. 3-layer validator: format, autotest, rubric.

Free · 12 runs/hour · no email needed

spec-linter-lite · demoDemo

Your spec

# Sales agent Greets customers, answers questions about products, closes sales. Should be friendly.

Result

  1. LAYER 1 — FORMATschema_present: FAIL — inputs are not described
  2. LAYER 1 — FORMATpii_safe: PASS — no secrets or emails
  3. LAYER 2 — AUTOTEST“Customer asks about a competitor” → expect a neutral answer
  4. LAYER 3 — RUBRICclarity 2/5 — “friendly” can't be tested
VERDICT · NEEDS_REVIEW

A sample run on demo data — this is what the result looks like in the simulator.

The problem this solves

Your spec reads clean to a human. Then the agent ignores half of it — picks the wrong output shape, leaks a PII field, fills an ambiguous acceptance criterion with whatever it felt like. You only find out after it ships, when someone asks why the agent did the opposite of what you 'clearly wrote.' The gap isn't in the model; it's between what you wrote and what it actually heard, and you have no way to see that gap before it bites.

After completing this

You stop trusting the human read and start seeing your spec through the model's eyes. Drop a spec, get three layers back — format, a live autotest on a real model, and a rubric grade — and the gap shows up in 30 seconds, line-quoted. After a few specs you write the next one already passing the linter, because you've internalized what makes an agent obey.

How it works

  1. It's a simulator, not an article

    You don't read about spec quality — you drop a real spec and watch a real model run it back, with the failures quoted line by line.

  2. You answer first, theory after

    You commit to your spec, then see exactly where it broke. The lesson lands because you already had skin in it.

  3. From the demo, a direct bridge to the paid track

    Once you can read a spec like a model, the paid 'Specification for AI' track turns this into PRD-as-tests and acceptance suites your team and your AI both trust.

Who this is for

Founders writing their first PRD for an AI feature

Catch the ambiguous acceptance criterion before an agent ships against it — and before a contractor bills you to discover it.

Tech leads reviewing specs from a junior who 'used ChatGPT'

Triage a stack of specs from worst to best in minutes. The rubric score does the first pass so you only deep-read the salvageable ones.

Ops managers turning SOPs into agent instructions

See the autotest run your minimal prompt on a real model, so you fix the SOP-to-spec gap before it reaches production traffic.

Not for you if

  • Нет: People who want the spec written for them — this grades what you wrote, it doesn't author it.
  • Нет: Anyone who treats a failing lint as an insult; the point is to fail here instead of in prod.
  • Нет: Teams with a mature eval harness already gating every spec — you've built this in-house.

When you'll reach for this

01

Your spec passes a human read but the agent does whatever it wants

The autotest runs the minimal prompt against a real model — you see the gap between what you wrote and what it heard, in 30 seconds.

02

You're about to send a spec to a contractor and want a second pair of eyes

Format + rubric flag the parts a developer would push back on — missing acceptance criteria, ambiguous output schema, leaked PII fields.

03

You inherited a folder of specs and don't know which ones are rotten

Lint each one, rank by rubric score, kill the bottom third before they reach an agent.

Powers you walk away with

  • Да: Catch a vague acceptance criterion before an agent ships against it
  • Да: See your spec through the model's eyes, not your own
  • Да: Triage a stack of specs from worst to best in under ten minutes
  • Да: Write the next spec already passing the linter — habits, not rules

What's inside

  1. 01Format check: schema, PII, source attribution
  2. 02Auto-test: minimal prompt run on a target model
  3. 03Rubric grade: clarity, modularity, model-fit

Used internally on every spec before it ships to a paying customer.

What stays in your portfolio

  • Да: A reflex for spotting a vague acceptance criterion before it reaches an agent.
  • Да: The habit of reading your own spec through the model's eyes, not your own.
  • Да: A repeatable triage move: rank a folder of specs by rubric score, kill the bottom third.
  • Да: Specs that pass the linter on the first write because the rules became instinct.

Pairs well with

$39

Specification for AI

Write specs that an AI can't misread. PRD-as-tests, acceptance criteria, deterministic checks.

Try Lite free
Free

Folder Sniff

Upload an existing AI folder — get the 3 most critical issues, ranked, with fix suggestions.

Sniff a folder
$89Ch1 free

AI Ops Analyst

Make the numbers honest. Interrogate metrics, audit triage agents, write reconciliation tests, brief AI that surfaces bad news instead of burying it.

Start Ch1 free

FAQ

Is it free?

Yes, fully. Lint as many specs as you want — no card, no signup.

Do I need a card?

No. Paste a spec, get the three-layer report. There's no payment step.

How long does it take?

One spec lints in about 30 seconds. A real habit forms after you run three or four of your own.

What's next after the demo?

Once you can read a spec like a model does, the paid 'Specification for AI' track turns this reflex into PRD-as-tests with weighted rubrics and acceptance suites.

What does the autotest actually do?

It runs the minimal prompt your spec implies against a real target model, so you see the gap between what you wrote and what the model heard — not a static checklist.

Start with the free part

No card, no commitment — try it first, decide later.

Try free — no signup

Made by mihalkevich.com — runs on the FolderAIThe engine that powers the courses and simulators engine.