Module 1

SDD Foundations

What spec-driven development actually is, how long a spec should live and how it changes, the spec-driven workflow from constitution to converge — with the human gates along the way — and the honest case against adopting it.

Time 50 min
Format Reading + knowledge check
Prerequisites None

Learning objectives

By the end of this module, you will be able to:

  1. Explain why AI agents move the bottleneck upstream, and what it means for a spec to be executable rather than disposable scaffolding.
  2. Define a “spec” and tell it apart from a PRD, a design doc, or a detailed prompt.
  3. Separate the two persistence questions — how long a spec matters, and how the artifact set changes — and name the main risk of each model.
  4. Walk the full and short Spec Kit workflow, from constitution to converge, and say what each step produces.
  5. Say where a person decides at each human gate, and what the implement–converge loop and “Converged” mean.
  6. Place the workflow in SPECTRA’s seven agentic SDLC phases.
  7. Recognize when SDD is overkill, and articulate its main critiques honestly enough to push back when adoption is oversold.

1.Why this is happening now

For decades, code has been the source of truth in software development. We write requirements docs and they drift. We draw architecture diagrams and they rot. We write tests, often after the fact. The code — whatever it actually does — becomes the de facto truth of the system, and “what should this function do?” gets answered with “read the code.”

That stayed manageable as long as humans were the bottleneck on writing code. AI coding agents have changed that. When an agent can produce a working endpoint in 30 seconds, the bottleneck shifts upstream. The new constraint is: how clearly can you state what you actually want?

Spec-driven development is the answer to that question. Rather than typing increasingly elaborate prompts into a chat window, you write a structured specification first, and the agent uses it as the authoritative description of what to build.

Spec Kit, the open-source toolkit this course is built on, frames it as reversing an old relationship. A spec used to be scaffolding, taken down once the “real” coding began. In SDD the spec stays up: it is the input the agent generates from and the yardstick the result is measured against. Spec Kit’s word for that is executable.

The simplest way to put the inversion (from Piskala’s 2026 paper):

In spec-driven development, code is the implementation detail of the specification — not the other way around.

That’s the whole idea. Everything below is detail.

Three kinds of work SDD is for

Spec Kit names three development phases the method serves. You’ll do two of them in this course.

2.What is a spec, exactly?

You’re going to hear “spec” used loosely in tutorials, vendor blog posts, and social media threads. Some people use it to mean “a slightly longer prompt.” Others mean “a regenerated copy of the codebase in markdown form.” Neither is what we mean.

For this course, a spec has three properties:

  1. Structured. It follows a known format, not just prose. Sections for user scenarios, requirements, success criteria, scope.
  2. Behavior-oriented. It describes what should happen and why, not how it’s implemented.
  3. Executable. It is consumed by tooling — AI agents, consistency checks, CI gates, contract tests — not just read by humans.

That third property is what separates a spec from a PRD or a design doc:

ArtifactRead byEnforced?Update cadence
PRDProduct, engineeringNo — humans interpret itOften stale within weeks
Design docEngineering peersNo — point-in-time reviewRarely updated
Detailed promptOne AI agentNo — vanishes after the chatPer-conversation
SDD specAI agent, reviewers, automated checksYes — the agent builds from it, and converge, tests, and CI flag divergenceDeliberately, under the team’s persistence model (section 3)

That’s the bar: when the code and the spec disagree, something notices — a converge pass, a failing test, a CI gate — and a person resolves it before merging.

What a spec is not

3.How long a spec lives — and how it changes

Not all SDD adoption looks the same, and most of the confusion online comes from people answering different questions without realizing it. There are two, and they are independent:

  1. How long does the spec matter? Only until the code exists, or for the life of the system?
  2. When requirements change, what happens to the artifacts you already have? Do you edit them, or leave them as history and start new ones?

3.1 How long the spec matters

LevelWhat it meansFitsRisk
Spec-firstWritten before coding, used during the work, then allowed to goTeams new to SDD; initial AI-assisted feature work, where the value is up-front clarityOnce the spec is dropped, the code quietly becomes the truth again — or a stale spec lingers that someone still trusts
Spec-anchoredKept after implementation and used for future changes; when behavior changes, both changeLong-lived production systems, regulated domains, contracts other teams depend onUpkeep: a kept spec nobody updates is worse than none, because people trust it
Spec-as-sourceThe only artifact humans edit; code is regenerated from it like a compiled binaryMature, certified code-generation pipelinesEverything rests on a generator you trust to overwrite code; any hand edit is lost or breaks the model

Spec-first is where almost every team should start. Spec-anchored is the sweet spot for most production systems: BDD with Cucumber and contract testing with Pact or Specmatic both predate the “SDD” label, and both are fundamentally spec-anchored. Spec-as-source is real in narrow domains — automotive control software generated from Simulink models for ISO 26262, for example — and tools like Tessl are exploring it more broadly, but it remains experimental for business software. If someone tries to sell you spec-as-source for your CRUD app, push back.

3.2 How the artifact set changes

The second question is about spec.md, plan.md, and tasks.md once they exist and a requirement moves. Spec Kit names three models:

ModelWhen a requirement changes…FitsRisk
Flow-backEdit whichever artifact (or the code) the change touches first, then reconcile the rest by handSmall teams, fast iteration, plans expected to shift as you learnSilent divergence — a lower-level artifact changes, the spec doesn’t, and nobody knows which to trust
Flow-forwardLeave finished feature directories untouched as a record; the new requirement gets a new feature directoryAudit trails, traceability, well-scoped features rarely revisitedScattered context — related decisions spread across directories unless you name and link them well
Living specEdit spec.md first — it is the contract — then regenerate or revise the plan and tasks from itA stable product contract; teams comfortable regenerating derived filesLost rationale — decisions buried in a regenerated plan disappear unless recorded elsewhere, such as an architecture decision record

3.3 A team convention, written in the constitution

Spec Kit doesn’t answer either question for you, and there is no CLI setting for it. The choice is a team convention: make it explicitly and record it in the constitution (.specify/memory/constitution.md), so every contributor — and every agent — reads the same rule. Two questions get you most of the way: are finished feature directories history or editable work areas? And is spec.md the single source of truth, or may plan.md and tasks.md become co-equal sources?

Where our team starts: spec-first for greenfield work, one small spec per change for brownfield work, spec-anchored where contracts are already stable (APIs you publish, modules other teams depend on). For the second question there is no house default — choose deliberately, per project.

Coming up.In Module 4 you’ll choose a model for the to-do app and write it into its constitution.

4.The spec-driven workflow

Every SDD tool you’ll encounter — GitHub Spec Kit, Kiro, OpenSpec, Tessl — shares one spine: specify what, plan how, break it into tasks, implement. Spec Kit adds a constitution up front, three optional quality gates, and a verification step at the end. We use Spec Kit’s vocabulary because the labs run on SPECTRA, which is built on it: the commands and artifacts are Spec Kit’s, and SPECTRA installs them and layers its own agents around them.

Command spelling.Every command has one identity, written with dots — speckit.specify, speckit.spectra.adr — and your coding agent decides how you type it. This course shows the trigger form Claude Code and GitHub Copilot use: a leading slash and dashes, as in /speckit-specify and /speckit-spectra-adr. Other agents differ: kiro-cli uses /speckit.specify, Codex uses $speckit-specify, and Kimi uses /skill:speckit-specify. Your agent’s own command list always shows the exact trigger.
The spec-driven workflow Constitution once per project, then specify, clarify (optional), plan, checklist (optional), tasks, analyze (optional), implement and converge. Converge either appends tasks, which sends you back to implement, or reports Converged, after which you review and open a pull request. Constitution Specify Clarify Plan Checklist Tasks Analyze Implement Converge Converged once per project spec.md optional gate plan.md optional gate tasks.md optional gate code + tests append-only review, then PR tasks appended → implement again core step optional quality gate A person reviews every artifact before the next step
The full path. The short path drops the three dashed gates.

4.1 Once per project: the constitution

/speckit-constitution writes .specify/memory/constitution.md: the principles — coding standards, testing rules, architectural constraints — that every later step is checked against. Run it before the first feature, and again whenever a principle changes. SPECTRA calls the agent behind it Guardrails: standards encoded once instead of repeated in every prompt. Module 2 looks at what belongs in it.

4.2 The full path, step by step

StepWhat it producesWhere you stand
/speckit-constitution
once per project
The constitutionYou approve every principle before it binds an agent
/speckit-specifyspec.md — the what and why: user scenarios, requirements, success criteria, no tech stack — plus the built-in checklists/requirements.md, which specify and clarify tick themselvesYou review behavior and scope, and strip out anything that is really “how”
/speckit-clarify
optional
Up to five targeted questions per pass; answers written into spec.mdYou answer the questions
/speckit-planplan.md — the how: stack, architecture, technical constraints — plus supporting files such as data-model.mdYou review the design and its trade-offs
/speckit-checklist
optional
A custom checklist under checklists/ — “unit tests for your requirements”: is the spec complete, clear, consistent?You tick each item, only when satisfied
/speckit-taskstasks.md in phases: Setup → Foundational → one per user story, in priority order → Polish; [P] marks parallel tasksYou check order, coverage, and scope creep
/speckit-analyze
optional
A read-only report of gaps and contradictions across spec, plan, and tasksYou fix each finding at its source, then re-run
/speckit-implementCode, and the tests the tasks call for; finished tasks marked [X]You review each stage before the next
/speckit-convergeConverged, or new tasks appended to tasks.mdYou decide it’s done, then review and open the pull request

Keeping specify and plan apart is the habit that matters most, and the one most engineers struggle with at first. Compare two prompts for the course’s to-do app:

/speckit-specify Let people filter the to-do list to show all, active, or completed items. The chosen filter should still be selected after a page reload.
/speckit-plan Next.js App Router, React, and TypeScript, as the app already uses. Keep the filter state in app/page.tsx next to the list, and persist it through lib/storage.ts.

The first says what a user gets. The second says how it will be built. When a specify prompt names a framework, a function, or a database column, it’s in the wrong step.

4.3 The short path

For a small feature, drop the three optional gates: once the constitution exists, run specify → plan → tasks → implement → converge. Only /speckit-specify is strictly required before /speckit-plan. SPECTRA’s roster lists clarify, checklist, and analyze as add-on agents you switch on as needed, next to the core agents that run the loop end to end. Switch them on whenever there is real ambiguity or the feature is headed for production — skipping them on a feature full of open questions just moves the questions into code review.

4.4 Implement and converge: the closing loop

/speckit-implement works through tasks.md phase by phase, in dependency order. Before it starts, it counts the items in every checklist; if any are unchecked, it asks whether to continue. It never ticks a box itself. On a large feature, scope each run and verify the result before starting the next:

/speckit-implement only execute tasks T001-T010, then stop and report progress

When implement finishes, run /speckit-converge. It compares the codebase against the spec, plan, and tasks, summarizes its findings by severity, and ends in one of two ways:

Converge is append-only: it leaves your code alone, never removing a line, and appending tasks is the only write it can make. So “Converged” is a statement about coverage, not a code review — the build matches what was written down. Whether that is right is still your call.

4.5 Where you stand at each gate

SPECTRA’s short version is agents draft, people decide — agentic, not autonomous. Four gates are easy to get wrong:

Because each step produces an artifact that constrains the next, each is also a review checkpoint. By the time you’re looking at a pull request, divergence has had several chances to be caught. When a team complains that “AI-generated code is hard to review,” it’s usually because they skipped straight to the code and tried to review only the output.

4.6 The workflow inside SPECTRA’s agentic SDLC

SPECTRA describes the whole lifecycle as an agentic SDLC in seven phases, each with a human gate. The Spec Kit workflow covers the first five. Every phase reads and writes the same durable artifacts — constitution, specs, plans, tasks, decision records — so no phase starts cold.

PhaseWhat runs thereHuman gate
00 Foundation/speckit-constitutionArchitects and engineering leads approve every standard
01 Plan/speckit-specify, /speckit-clarify, /speckit-checklistThe product owner approves intent, scope, and business alignment
02 Design/speckit-planArchitects approve the design and its decisions
03 Implement/speckit-tasks, /speckit-analyze, /speckit-implementEngineers review every change
04 Test/speckit-converge, plus the tests written alongside the code in each story’s phaseQE decides what a finding means
05 DeploySPECTRA’s pull-request agents, speckit.spectra.create-pr and speckit.spectra.review-prA maintainer merges
06 MaintainThe bug-fix workflow — Module 6The team decides the fix, and whether the lesson becomes a standard

Checklist runs after plan in the command sequence, but it tests the requirements, not the design — which is why it sits in Plan.

Two paths sit outside this sequence. Idea assessment (Module 5) happens before any of it: it decides whether Plan should start at all. Bug fixing (Module 6) is a separate path and doesn’t need SDD first — you can run it on a codebase that has no specs at all. Both are Spec Kit workflows, added to a project with specify extension add.

5.How SDD relates to things you already know

TDD

A unit test is a micro-spec. You write the test first, declaring what the function should do; then you write the function. SDD extends the same “specify before you build” thinking up to feature, system, and architectural scope.

Use both. TDD at the unit level, SDD at the feature/system level. They don’t compete.

BDD

Behavior-driven development is the most direct ancestor of modern SDD. Gherkin scenarios — Given / When / Then — are literally executable specs. Cucumber and SpecFlow have been turning them into automated tests for over a decade.

What changed: AI agents can now generate code from those scenarios, not just verify against them. This is why people like Bryan Finster (who has been doing BDD for years) call SDD “BDD with branding.” He’s not entirely wrong.

If your team already runs BDD, SDD is the smallest possible step from where you are. You’re already most of the way there.

Contract testing and contract-driven development

Pact and Specmatic have enforced agreements between services for years. Spec Kit folds the idea into SDD as contract-driven development: when a component exposes an interface that something else consumes — an API, a library, a CLI whose output gets parsed — agree on its observable obligations (inputs, outputs, errors, guarantees such as retry safety) before building either side. One side owns the contract, the agreement is settled during planning, and each side then builds and tests against it, whatever your architecture or repository layout.

PRDs and design docs

These are advisory. Humans read them and write code that hopefully matches. Drift is normal and expected; nobody is alarmed when a 6-month-old design doc no longer matches the code, because nothing in the system enforces alignment.

SDD specs are enforced: the agent builds from them, and a mismatch between spec and code is flagged and has to be resolved.

This is the cultural shift the team has to internalize: documents that nobody enforces become decoration. Documents that the workflow enforces become contracts.

Agile / Scrum

User stories with acceptance criteria are specs. The Definition of Done is a form of spec. You’re already specifying things. SDD asks you to do that with more discipline and to make the specs authoritative rather than advisory.

You don’t need to change your sprint structure to adopt SDD. You change what counts as “done” for a story and what artifacts get versioned.

6.When not to use SDD

There are tasks where spec ceremony costs more than it saves. The honest list, drawn from Augment’s “skip the spec” criteria:

Skip the spec when……because
The work is exploratory or experimentalYou don’t know what you want yet; specifying prematurely will lock in the wrong shape
A single prompt produces usable outputSpec overhead would exceed the work itself
The output can be reviewed in under five minutesYou can verify by reading; you don’t need a contract
You’re building a throwaway prototypeMaintenance benefits don’t apply
The change is mechanical or low-riskBumping a version number doesn’t need a spec

The trigger heuristic that’s worth memorizing:

If you’d be annoyed to have the agent interpret requirements differently than you meant, write the spec. If you could fix the output in a quick follow-up prompt, skip it.

This is calibrated by stakes, not by intrinsic complexity. A two-line auth token check might warrant a spec because misinterpretation has security consequences. A 200-line refactor of a private function might not, because if the AI misreads, you’ll catch it on diff review.

Skipping the spec doesn’t have to mean skipping structure. A defect with a clear symptom has its own lighter path — Spec Kit’s bug workflow, which you’ll run in Module 6 — and an idea nobody has vetted yet is better assessed (Module 5) than specified.

7.The honest case against (and why we’re doing it anyway)

Spec-driven development has serious critics, and you should hear them before adopting any practice. This is the part most vendor blog posts skip.

Birgitta Böckeler (Thoughtworks): “Verschlimmbesserung”

Böckeler used Spec Kit, Kiro, and Tessl on real problems and found that current SDD tools generate too many files for the size of the problem they solve. Applied to a small bug fix, Spec Kit produced 4 user stories and 16 acceptance criteria when one prompt would have been better. She uses the German word Verschlimmbesserung — “an improvement that makes things worse” — for SDD setups that generate more markdown than code.

Her concrete warning: reviewing markdown can be harder than reviewing code. If your specs are repetitive, contain code-like detail, or duplicate what the code already says, you’ve made your job harder, not easier.

Take this seriously. Spec Kit now offers a short path for small features and a separate, lighter workflow for bug fixes, but the warning still applies to anything you push through the full loop. We’ll cover countermeasures in Module 8 (Best Practices & Pitfalls).

Bryan Finster: “BDD with branding”

Finster’s argument is that the core insight of SDD — write specs first, derive code from them — has been agile/BDD wisdom for two decades. Calling it a paradigm shift in 2026 is marketing.

He’s right that the insight isn’t new. What’s new is:

  1. AI agents are now competent consumers of specs, not just humans.
  2. CI/CD has matured to the point where spec enforcement is routine, not exotic.
  3. Tooling for spec authoring and validation is finally good enough to be worth adopting.

Together, those three changes make spec-first practice more powerful and more enforceable than ever before. So the practice is more valuable now than it used to be, even if the idea has been around for years.

The honest position: SDD isn’t revolutionary, but it’s currently the most practical way to keep AI-generated code aligned with intent in a production codebase.

Why we’re adopting it anyway

Three reasons, in order of importance:

  1. It addresses a real problem. When AI agents generate code, the gap between “what we wanted” and “what we got” widens unless something keeps both ends explicit. Specs are that something.
  2. It scales review better than PR-time inspection. The Agoda study (10,000+ devs, March 2026) found that high-AI-adoption teams completed 21% more tasks but PR review time grew 91%. Pushing validation upstream into specs is one of the few ways to stop that bottleneck from breaking the team.
  3. It improves cross-functional collaboration. Done well, specs become the shared interface between PM, architecture, engineering, and QA. Done badly, they’re just more documentation. Modules 9 and 10 are about making sure we land on “done well.”

8.The mental shift to make

Engineers tend to enter SDD thinking it’s a tooling change: install SPECTRA, learn nine slash commands, done. That’s the version that produces Verschlimmbesserung.

The mental shift that actually unlocks the value:

In the “before” view, drift is inevitable and you accept it. In the “after” view, drift is a defect and you fix it.

You don’t need to fully internalize this on day one. You’ll get there by doing the labs. But notice when you’re slipping back into “the code is the truth, the spec is documentation” thinking — that’s the failure mode that turns SDD into bureaucracy.


Knowledge check

Spend 10–15 minutes on these. Click an answer to see whether you got it right and why.

Knowledge check 8 questions
1. Spec Kit says that in spec-driven development, specifications become executable. What does that mean in practice?
Answer: b — the spec stays in use
“Executable” is about the spec’s role, not its format: it keeps driving the work, and the result is checked against it, instead of being thrown away once coding begins. Option a is spec-as-source at its most extreme, which is rare. Option c is the anti-pattern — a spec full of code has lost the “what”. Option d confuses format with function.
2. Which of the following best describes the difference between a PRD and an SDD spec?
Answer: c — SDD specs are enforced
Length, location, and authorship are surface differences. The defining property is enforceability: an SDD spec is consumed by tooling and checks that flag it when reality diverges, while a PRD relies on people reading it and hoping the code matches.
3. Your team has a 25-line bug fix to make in a payment service. The bug is well understood, the change is mechanical, and you can verify correctness from the diff in two minutes. Should you write a spec for it with /speckit-specify?
Answer: b — No
This is exactly the case Augment’s “skip the spec” heuristic targets. Mechanical, well-understood, quickly reviewable work doesn’t justify the ceremony, and forcing a spec here is how SDD becomes bureaucracy. If you want a traceable record for a real defect, the bug-fix workflow in Module 6 is the lighter path — and it doesn’t need a spec.
4. Your team keeps each spec after it ships and uses it whenever that area changes. When a requirement changes, nobody edits a finished feature directory — you open a new one and link back to the old. Which pair names your convention?
Answer: b — spec-anchored + flow-forward
Two separate questions. Keeping the spec and using it for later changes is spec-anchored (spec-first would let it go after shipping). Leaving finished directories untouched and opening a new one is flow-forward. Living spec would edit spec.md first and regenerate the plan and tasks; flow-back would edit whichever artifact the change touches and reconcile by hand. Linking back answers flow-forward’s risk, scattered context — and the whole convention belongs in the constitution.
5. You’ve run /speckit-checklist, and it generated a custom checklist of twelve items about your spec. Who marks them [x], and when?
Answer: c — the reviewer, only when satisfied
Custom checklists are reviewer-owned: /speckit-checklist never ticks its own items, and a tick means a person decided the requirement is well written, not that code is done. That rules out a and b. Nor is it ignored (d): /speckit-implement counts unchecked items and asks before continuing, but never changes a checkbox. Don’t confuse this with the built-in checklists/requirements.md, which specify and clarify tick themselves — those ticks are the agent’s claim, not a sign-off.
6. /speckit-converge finishes by reporting that it appended three tasks under a Convergence section of tasks.md. What do you do next?
Answer: b — implement, then converge again
Converge is append-only: it leaves your code alone, never removing a line, and appending tasks is its only possible write. So there is nothing to revert (a) and nothing has been fixed yet (c). “Tasks appended” means work remains: build it and check again, as often as it takes (not d). Only Converged — with tasks.md left untouched — means it’s time for review and the pull request.
7. Read this snippet from a teammate’s spec for a file-upload feature:
“Create a function called uploadFile(file) that calls multer.single('file'), then writes to S3 using putObject(), then inserts a row into the uploads table with columns id, path, created_at.”
What’s wrong with it?
It’s prescribing implementation, not describing behavior. Function names, library calls, and database column names belong in the plan (/speckit-plan), not the spec (/speckit-specify). A behavior-level rewrite would be: “Users can upload files up to X MB; uploaded files are stored durably and tracked with timestamps; users can later list their own uploads.” Object storage, S3, and the table layout are choices the team makes later, when it writes the plan.
8. Why does the spec-driven workflow reduce review burden compared with reviewing AI-generated code at PR time? Name at least two points where a person decides before any code exists.
Each step produces a small, focused artifact that a person reviews before the next step starts, so divergence between intent and code can be caught at many checkpoints instead of one. Before any code exists, a person approves the constitution, reviews the spec, answers clarify’s questions, reviews the plan, ticks the checklist items they’re satisfied with, and fixes analyze findings at their source. Then they review each implement stage, and converge confirms nothing in the artifacts was left unbuilt. When everything piles up at PR review instead, reviewers either skim or burn out — both bad.

What’s next

In Module 2 you’ll look inside a good spec: the six-element framework for writing specs that don’t bloat or drift, and how the constitution keeps cross-cutting rules out of individual specs. You’ll finish with a short writing exercise — rewriting a deliberately bad spec into a good one.

After that, Module 3 moves into the first hands-on lab: installing SPECTRA and running the workflow on a small greenfield feature.