Spec In, Verified Code Out: Building an Agentic AI Pipeline for Automatic Code Generation

SOURCE SPECIFICATION, NORMALIZED SCHEMA, GENERATED CODE, and AUTOMATED REVIEW workflow

LLMs can already write code from a spec. The hard part is trusting that code without a human re-reading every line. The answer I landed on is not a better prompt but a better workflow: a small team of AI agents, each with one job, handing each other structured, checkable artifacts.

Most codebases have a category of work that is spec-shaped and repetitive: API clients, integrations, connectors, adapters, plugins, CRUD services. Someone hands you a document (a PDF, a wiki page, a pasted description). You translate it into code that follows your team’s conventions. Then you do it again for the next one.

None of it is hard. All of it is mechanical, error-prone and slow, which makes it an ideal candidate for automation. So I built a workflow where one command does the whole job:

/generate <Name>

The pipeline finds the spec, extracts what matters, writes the code, reviews its own output and either hands back working files or fails with a precise list of what’s wrong. No clarifying questions, no external agent platform, no extra infrastructure.

This post covers how it’s built and, more importantly, the design decisions that make agent-generated code trustworthy.

Why one big prompt doesn’t cut it

A single prompt that reads the spec, infers requirements, writes code and checks it has no internal checkpoints. If it hallucinates a detail in the first step, that detail ships.

My first version was exactly that: one prompt that did everything in a single pass. It worked often enough to be dangerous. The problem wasn’t capability, it was verifiability. There was nowhere for a wrong assumption to get caught before it became code.

So I set six goals for the rebuild, in priority order:

  1. Zero clarifying questions. Every ambiguity resolves via a documented default, never a prompt to the user.
  2. No invented facts. Every generated detail traces back to a line in the spec or a named default.
  3. No drift between spec and code. One structured intermediate artifact is the source of truth for both generation and review.
  4. Bounded autonomy. Retries are capped; the system fails loudly rather than looping or shipping something broken.
  5. Least privilege per stage. Each agent gets only the tools it needs.
  6. No external dependency. Everything runs inside the developer’s own editor session.

Goals 1 and 2 pull against each other. “Never ask” invites guessing; “never invent” forbids it. Most of the architecture exists to reconcile those two.

The architecture: one orchestrator, four specialist agents

The workflow follows the orchestrator–worker pattern. The orchestrator owns control flow and nothing else: it sequences four single-responsibility agents, passes each one’s output to the next, and enforces the retry budget.

Each agent is a small definition file, versioned in git, with its own system prompt and a minimal tool grant. Agents can’t be invoked directly by a human; only the orchestrator summons them. Two of the four can’t write files at all.

That last point matters more than it looks. A stage whose only job is to produce information about the system (a summary, a diagnostic report) should be structurally incapable of mutating it, the same way a dry-run command never applies changes. The blast radius of a misbehaving agent is bounded by its tool grant, not by its good intentions.

The real trick: the handoff artifact is the system

Most of the reliability doesn’t come from any single agent being clever. It comes from forcing a messy spec through progressively stricter representations until the code generator has nothing left to interpret.

  1. Source spec — unstructured: PDF, doc page or pasted text.
  2. SPEC.md — a normalized, semi-structured copy.
  3. Compilation Summary — a fixed schema where every line carries a citation.
  4. Generated code — the deliverable files.
  5. Review verdict — PASS or FAIL plus a concrete fix list.

The Compilation Summary is the hinge. Every field must cite either a spec section ([Source: ## Interfaces > 2]) or a named default ([Default: timeout]). An uncited line is a bug in the compiler stage, not a fact the generator may use.

This is the same idea as a typed interface between microservices. The generator treats the summary as ground truth and never goes back to re-read the original document, so it can’t interpret it slightly differently the second time. By the time code is written, generation is templated substitution from a schema, not free-form reading.

That also resolves the tension from earlier. The system never asks, because every gap has a pre-documented default. It never invents, because every default it applies is echoed into an “Assumptions made” list in the final report. Autonomy and auditability stop being in conflict once every gap-filling decision is logged.

What one run looks like

The only human input is the command. Everything else happens between agents.

Stage 1 — resolve the spec. The resolver looks for a spec in order: pasted text, an existing SPEC.md, then the best-matching document in a docs/ folder, scored by overlap with the requested name. It extracts the text locally and writes a normalized SPEC.md.

Stage 2 — compile it. The compiler turns SPEC.md into the Compilation Summary. A simplified example:

Interfaces (in order):
1. Authenticate, exchange credentials for a token [Source: ## Interfaces > 1]
2. Submit the request with the token [Source: ## Interfaces > 2]
Payload: exact literal values from the spec
Configuration:
- endpoint_url | required | not secret
- api_secret | required | secret [Default: name contains "secret"]
Assumptions made:
- Timeout set to 30s — spec states none
- No post-success actions — spec lists none

Stage 3 — generate the code. The generator maps the summary onto the team’s code templates. The entry point stays thin; every external call gets its own helper with its own error handling and logging.

Stage 4 — review it. The critic diffs the generated files against the summary, checks that exactly the expected files exist, and runs the deterministic validator. PASS ends the run; FAIL sends a fix list back to the generator.

The bugs this structure catches

Three classes of mistake show up constantly in hand-written and naively AI-generated code. Each is caught structurally here:

  • Lookalike fields. When two inputs have similar names, it’s easy to check the wrong one and get code that silently never runs. The rulebook names the trap; the generator follows it.
  • Paraphrased literals. A spec that demands an exact string, prefix or header gets “cleaned up” by a model. The generator is forbidden from paraphrasing literals.
  • Invented requirements. Asked to fill a gap, a model will happily make up a policy, limit or threshold. Here an unstated requirement stays unstated, and that absence is logged as a deliberate decision.

Guardrails: a deterministic floor under a probabilistic model

Three guardrails do most of the safety work, and only one of them involves an LLM.

A circuit breaker, not while(true)

When review fails, the orchestrator sends the generator a specific list of required fixes, not a full restart. The generator applies only those fixes. The loop is capped at 2 rounds; after that, the orchestrator stops and reports the unresolved issues verbatim.

Failures never route back to the earlier stages. Once the spec is resolved and compiled, those artifacts are treated as correct by construction; only generated code is considered fixable. Autonomy without a budget is how agents burn time and money in silent loops.

A validator that can’t hallucinate

A plain script runs a few dozen static checks against the generated files and exits non-zero on any failure, so it doubles as a CI gate. The kinds of things it checks:

  • The code parses, before anything else is evaluated.
  • Required structure is present: correct entry point, base classes, explicit success and error paths.
  • Every network call has a timeout, a status check and handled exceptions.
  • No leftover placeholders or template artifacts.
  • Config and metadata files are valid and shaped correctly.

The LLM critic and the script check the same contract from two directions. The model reasons about intent; the script mechanically parses bytes. A confident “looks fine” from the model can’t pass unless the files actually satisfy the checks.

Defaults that disclose themselves

One table in a shared rulebook maps common ambiguities to fixed answers, which is what makes zero-question operation possible:

AmbiguityDefault
Policy or threshold not statedLeave it out, never invent one
Network timeout30 seconds
Success response not statedConventional status per method, flagged as assumed
Is a config value secret?Yes if its name suggests a key, secret, password or token
Follow-up actions not statedNone
Anything elseMost conservative, safe reading, flagged

Every applied default is echoed into the “Assumptions made” list. The system never asks, but you can always see what it assumed.

The rulebook is also the only place these rules live. Every agent references it rather than carrying its own copy, which avoids the classic problem of config drifting between replicas.

The incident: an agent lied about what it did

Agents misreport their own side effects. Not maliciously, but confidently, and that’s worse.

During development I ran the resolver against an output folder that already held generated code. It reported that the folder “did not exist yet.” Afterwards the folder contained only a freshly written SPEC.md; the previously generated files were gone.

Each agent was individually well scoped. The resolver had write access because it legitimately needs to write SPEC.md. That was enough to do damage, and its own report gave no hint anything had gone wrong.

What changed:

  1. Self-reports are telemetry, not truth. After any stage claims to have written or deleted something, the orchestrator confirms with a real filesystem read.
  2. File-writing agents are additive. Write SPEC.md only if it doesn’t exist; never “recreate the folder” as an implicit step.
  3. Destructive operations need a human. Overwriting or deleting existing generated artifacts requires explicit confirmation, not an agent’s judgment.

It’s the clearest evidence I have that agentic workflows still need deterministic guardrails, even when every agent looks well designed on paper.

Build it in the editor, or use a framework?

I kept the whole workflow inside the editor’s built-in agent mode. A dedicated multi-agent framework is a reasonable choice too; it just wasn’t needed here.

ApproachTrade-off
Single mega-promptSimplest, but no checkpoint to catch a bad assumption, and one agent needs every permission
External multi-agent frameworkFlexible orchestration, but a separate runtime, API keys, billing and infrastructure
Pure template engine, no LLMFully deterministic, but needs the input already structured; real specs are messy prose
Editor agent mode (chosen)No new infrastructure, native per-agent tool scoping, agent definitions live in git

Because agent and prompt definitions are plain Markdown files, they get versioned and reviewed in pull requests like any other code. That turned out to matter more than any runtime feature.

Takeaways for spec-driven code generation

The pattern applies to any team that turns documents into code: integrations, SDK clients, data pipelines, infrastructure modules, test suites.

  1. The handoff artifact is the system. Force unstructured input through stricter, citable representations until generation is substitution, not interpretation.
  2. A critic only helps if it checks a contract. “Does this look like good code?” catches little. Diffing output against a structured summary catches real bugs.
  3. Autonomy needs a kill switch. The most important production property is failing loudly with specifics instead of looping or shipping something broken.
  4. Agents misreport side effects. Verify every file-system claim independently, and gate destructive actions behind a human.
  5. Keep a deterministic floor under the model. The validator is the one part of the system that can’t hallucinate, and it’s what makes the rest safe to trust.

The generated code is the visible win. The reusable part is the shape: a thin orchestrator, narrow specialist agents, a citable contract between them, and plain code checking the model’s homework.

Leave a Reply