Learning objectives
By the end of this module, you will be able to:
- Explain why AI agents move the bottleneck upstream, and what it means for a spec to be executable rather than disposable scaffolding.
- Define a “spec” and tell it apart from a PRD, a design doc, or a detailed prompt.
- Separate the two persistence questions — how long a spec matters, and how the artifact set changes — and name the main risk of each model.
- Walk the full and short Spec Kit workflow, from constitution to converge, and say what each step produces.
- Say where a person decides at each human gate, and what the implement–converge loop and “Converged” mean.
- Place the workflow in SPECTRA’s seven agentic SDLC phases.
- Recognize when SDD is overkill, and articulate its main critiques honestly enough to push back when adoption is oversold.
1.Why this is happening now
For decades, code has been the source of truth in software development. We write requirements docs and they drift. We draw architecture diagrams and they rot. We write tests, often after the fact. The code — whatever it actually does — becomes the de facto truth of the system, and “what should this function do?” gets answered with “read the code.”
That stayed manageable as long as humans were the bottleneck on writing code. AI coding agents have changed that. When an agent can produce a working endpoint in 30 seconds, the bottleneck shifts upstream. The new constraint is: how clearly can you state what you actually want?
Spec-driven development is the answer to that question. Rather than typing increasingly elaborate prompts into a chat window, you write a structured specification first, and the agent uses it as the authoritative description of what to build.
Spec Kit, the open-source toolkit this course is built on, frames it as reversing an old relationship. A spec used to be scaffolding, taken down once the “real” coding began. In SDD the spec stays up: it is the input the agent generates from and the yardstick the result is measured against. Spec Kit’s word for that is executable.
The simplest way to put the inversion (from Piskala’s 2026 paper):
In spec-driven development, code is the implementation detail of the specification — not the other way around.
That’s the whole idea. Everything below is detail.
Three kinds of work SDD is for
Spec Kit names three development phases the method serves. You’ll do two of them in this course.
- Greenfield (zero to one). From high-level requirements to a spec, a plan, and something new. Module 3’s lab.
- Creative exploration. Several implementations of one spec — different stacks, architectures, or UX patterns — compared side by side. With the intent held in the spec, building twice is affordable.
- Brownfield (iterative enhancement). Adding features to an existing system, or modernizing a legacy one. Module 4’s lab. Each change brings its own small spec, and the areas that change most often gradually build up spec coverage.
2.What is a spec, exactly?
You’re going to hear “spec” used loosely in tutorials, vendor blog posts, and social media threads. Some people use it to mean “a slightly longer prompt.” Others mean “a regenerated copy of the codebase in markdown form.” Neither is what we mean.
For this course, a spec has three properties:
- Structured. It follows a known format, not just prose. Sections for user scenarios, requirements, success criteria, scope.
- Behavior-oriented. It describes what should happen and why, not how it’s implemented.
- Executable. It is consumed by tooling — AI agents, consistency checks, CI gates, contract tests — not just read by humans.
That third property is what separates a spec from a PRD or a design doc:
| Artifact | Read by | Enforced? | Update cadence |
|---|---|---|---|
| PRD | Product, engineering | No — humans interpret it | Often stale within weeks |
| Design doc | Engineering peers | No — point-in-time review | Rarely updated |
| Detailed prompt | One AI agent | No — vanishes after the chat | Per-conversation |
| SDD spec | AI agent, reviewers, automated checks | Yes — the agent builds from it, and converge, tests, and CI flag divergence | Deliberately, under the team’s persistence model (section 3) |
That’s the bar: when the code and the spec disagree, something notices — a converge pass, a failing test, a CI gate — and a person resolves it before merging.
What a spec is not
- It is not a code-shaped artifact. If your spec reads like pseudo-code, you’ve gone too far down toward “how” and lost the abstraction. Stack and architecture belong in the plan.
- It is not exhaustive. You don’t spec everything. You spec what’s needed to remove ambiguity for the current piece of work.
- It is not generated and forgotten. If you generate a spec, never review it, and let an agent implement against it, you have a vibe-coding loop with extra steps.
3.How long a spec lives — and how it changes
Not all SDD adoption looks the same, and most of the confusion online comes from people answering different questions without realizing it. There are two, and they are independent:
- How long does the spec matter? Only until the code exists, or for the life of the system?
- When requirements change, what happens to the artifacts you already have? Do you edit them, or leave them as history and start new ones?
3.1 How long the spec matters
| Level | What it means | Fits | Risk |
|---|---|---|---|
| Spec-first | Written before coding, used during the work, then allowed to go | Teams new to SDD; initial AI-assisted feature work, where the value is up-front clarity | Once the spec is dropped, the code quietly becomes the truth again — or a stale spec lingers that someone still trusts |
| Spec-anchored | Kept after implementation and used for future changes; when behavior changes, both change | Long-lived production systems, regulated domains, contracts other teams depend on | Upkeep: a kept spec nobody updates is worse than none, because people trust it |
| Spec-as-source | The only artifact humans edit; code is regenerated from it like a compiled binary | Mature, certified code-generation pipelines | Everything rests on a generator you trust to overwrite code; any hand edit is lost or breaks the model |
Spec-first is where almost every team should start. Spec-anchored is the sweet spot for most production systems: BDD with Cucumber and contract testing with Pact or Specmatic both predate the “SDD” label, and both are fundamentally spec-anchored. Spec-as-source is real in narrow domains — automotive control software generated from Simulink models for ISO 26262, for example — and tools like Tessl are exploring it more broadly, but it remains experimental for business software. If someone tries to sell you spec-as-source for your CRUD app, push back.
3.2 How the artifact set changes
The second question is about spec.md, plan.md, and tasks.md once they exist and a requirement moves. Spec Kit names three models:
| Model | When a requirement changes… | Fits | Risk |
|---|---|---|---|
| Flow-back | Edit whichever artifact (or the code) the change touches first, then reconcile the rest by hand | Small teams, fast iteration, plans expected to shift as you learn | Silent divergence — a lower-level artifact changes, the spec doesn’t, and nobody knows which to trust |
| Flow-forward | Leave finished feature directories untouched as a record; the new requirement gets a new feature directory | Audit trails, traceability, well-scoped features rarely revisited | Scattered context — related decisions spread across directories unless you name and link them well |
| Living spec | Edit spec.md first — it is the contract — then regenerate or revise the plan and tasks from it | A stable product contract; teams comfortable regenerating derived files | Lost rationale — decisions buried in a regenerated plan disappear unless recorded elsewhere, such as an architecture decision record |
3.3 A team convention, written in the constitution
Spec Kit doesn’t answer either question for you, and there is no CLI setting for it. The choice is a team convention: make it explicitly and record it in the constitution (.specify/memory/constitution.md), so every contributor — and every agent — reads the same rule. Two questions get you most of the way: are finished feature directories history or editable work areas? And is spec.md the single source of truth, or may plan.md and tasks.md become co-equal sources?
Where our team starts: spec-first for greenfield work, one small spec per change for brownfield work, spec-anchored where contracts are already stable (APIs you publish, modules other teams depend on). For the second question there is no house default — choose deliberately, per project.
4.The spec-driven workflow
Every SDD tool you’ll encounter — GitHub Spec Kit, Kiro, OpenSpec, Tessl — shares one spine: specify what, plan how, break it into tasks, implement. Spec Kit adds a constitution up front, three optional quality gates, and a verification step at the end. We use Spec Kit’s vocabulary because the labs run on SPECTRA, which is built on it: the commands and artifacts are Spec Kit’s, and SPECTRA installs them and layers its own agents around them.
speckit.specify, speckit.spectra.adr — and your coding agent decides how you type it. This course shows the trigger form Claude Code and GitHub Copilot use: a leading slash and dashes, as in /speckit-specify and /speckit-spectra-adr. Other agents differ: kiro-cli uses /speckit.specify, Codex uses $speckit-specify, and Kimi uses /skill:speckit-specify. Your agent’s own command list always shows the exact trigger.4.1 Once per project: the constitution
/speckit-constitution writes .specify/memory/constitution.md: the principles — coding standards, testing rules, architectural constraints — that every later step is checked against. Run it before the first feature, and again whenever a principle changes. SPECTRA calls the agent behind it Guardrails: standards encoded once instead of repeated in every prompt. Module 2 looks at what belongs in it.
4.2 The full path, step by step
| Step | What it produces | Where you stand |
|---|---|---|
/speckit-constitutiononce per project | The constitution | You approve every principle before it binds an agent |
/speckit-specify | spec.md — the what and why: user scenarios, requirements, success criteria, no tech stack — plus the built-in checklists/requirements.md, which specify and clarify tick themselves | You review behavior and scope, and strip out anything that is really “how” |
/speckit-clarifyoptional | Up to five targeted questions per pass; answers written into spec.md | You answer the questions |
/speckit-plan | plan.md — the how: stack, architecture, technical constraints — plus supporting files such as data-model.md | You review the design and its trade-offs |
/speckit-checklistoptional | A custom checklist under checklists/ — “unit tests for your requirements”: is the spec complete, clear, consistent? | You tick each item, only when satisfied |
/speckit-tasks | tasks.md in phases: Setup → Foundational → one per user story, in priority order → Polish; [P] marks parallel tasks | You check order, coverage, and scope creep |
/speckit-analyzeoptional | A read-only report of gaps and contradictions across spec, plan, and tasks | You fix each finding at its source, then re-run |
/speckit-implement | Code, and the tests the tasks call for; finished tasks marked [X] | You review each stage before the next |
/speckit-converge | Converged, or new tasks appended to tasks.md | You decide it’s done, then review and open the pull request |
Keeping specify and plan apart is the habit that matters most, and the one most engineers struggle with at first. Compare two prompts for the course’s to-do app:
/speckit-specify Let people filter the to-do list to show all, active, or completed items. The chosen filter should still be selected after a page reload.
/speckit-plan Next.js App Router, React, and TypeScript, as the app already uses. Keep the filter state in app/page.tsx next to the list, and persist it through lib/storage.ts.
The first says what a user gets. The second says how it will be built. When a specify prompt names a framework, a function, or a database column, it’s in the wrong step.
4.3 The short path
For a small feature, drop the three optional gates: once the constitution exists, run specify → plan → tasks → implement → converge. Only /speckit-specify is strictly required before /speckit-plan. SPECTRA’s roster lists clarify, checklist, and analyze as add-on agents you switch on as needed, next to the core agents that run the loop end to end. Switch them on whenever there is real ambiguity or the feature is headed for production — skipping them on a feature full of open questions just moves the questions into code review.
4.4 Implement and converge: the closing loop
/speckit-implement works through tasks.md phase by phase, in dependency order. Before it starts, it counts the items in every checklist; if any are unchecked, it asks whether to continue. It never ticks a box itself. On a large feature, scope each run and verify the result before starting the next:
/speckit-implement only execute tasks T001-T010, then stop and report progress
When implement finishes, run /speckit-converge. It compares the codebase against the spec, plan, and tasks, summarizes its findings by severity, and ends in one of two ways:
- Converged. Nothing the artifacts ask for is left unbuilt, and
tasks.mdis left exactly as it was. Move on to review and the pull request. - Tasks appended. It found gaps and added them as new tasks under a Convergence section of
tasks.md. Run/speckit-implementto build them, then converge again. Each pass should find less; repeat until it reports Converged.
Converge is append-only: it leaves your code alone, never removing a line, and appending tasks is the only write it can make. So “Converged” is a statement about coverage, not a code review — the build matches what was written down. Whether that is right is still your call.
4.5 Where you stand at each gate
SPECTRA’s short version is agents draft, people decide — agentic, not autonomous. Four gates are easy to get wrong:
- Clarify — you answer. The agent finds the ambiguity and asks; you decide. If you don’t know an answer yet, chase it down before planning rather than guessing.
- Checklist — you tick. A custom checklist from
/speckit-checklistbelongs to the reviewer: the command never marks its own items, and an item gets[x]only when a person is satisfied. The agent can help evaluate when asked, but never silently approves. A tick means “this requirement is well written,” not “this code is done.” The built-inchecklists/requirements.mdis different: specify and clarify tick and untick it themselves, so its ticks are the agent’s own claim, not a sign-off. - Analyze — you fix at the source. Analyze only reports. A requirement problem goes back to specify or clarify, a design problem to plan, a task-list problem to a fresh run of tasks; then re-run analyze until it’s clean. Hand-patching
tasks.mdto silence what is really a spec gap just hides the gap. - Implement — you review each stage. Read the code and the updated artifacts together. If the output doesn’t match the spec, fix whichever was wrong — code or spec — and keep the spec authoritative.
Because each step produces an artifact that constrains the next, each is also a review checkpoint. By the time you’re looking at a pull request, divergence has had several chances to be caught. When a team complains that “AI-generated code is hard to review,” it’s usually because they skipped straight to the code and tried to review only the output.
4.6 The workflow inside SPECTRA’s agentic SDLC
SPECTRA describes the whole lifecycle as an agentic SDLC in seven phases, each with a human gate. The Spec Kit workflow covers the first five. Every phase reads and writes the same durable artifacts — constitution, specs, plans, tasks, decision records — so no phase starts cold.
| Phase | What runs there | Human gate |
|---|---|---|
| 00 Foundation | /speckit-constitution | Architects and engineering leads approve every standard |
| 01 Plan | /speckit-specify, /speckit-clarify, /speckit-checklist | The product owner approves intent, scope, and business alignment |
| 02 Design | /speckit-plan | Architects approve the design and its decisions |
| 03 Implement | /speckit-tasks, /speckit-analyze, /speckit-implement | Engineers review every change |
| 04 Test | /speckit-converge, plus the tests written alongside the code in each story’s phase | QE decides what a finding means |
| 05 Deploy | SPECTRA’s pull-request agents, speckit.spectra.create-pr and speckit.spectra.review-pr | A maintainer merges |
| 06 Maintain | The bug-fix workflow — Module 6 | The team decides the fix, and whether the lesson becomes a standard |
Checklist runs after plan in the command sequence, but it tests the requirements, not the design — which is why it sits in Plan.
Two paths sit outside this sequence. Idea assessment (Module 5) happens before any of it: it decides whether Plan should start at all. Bug fixing (Module 6) is a separate path and doesn’t need SDD first — you can run it on a codebase that has no specs at all. Both are Spec Kit workflows, added to a project with specify extension add.
5.How SDD relates to things you already know
TDD
A unit test is a micro-spec. You write the test first, declaring what the function should do; then you write the function. SDD extends the same “specify before you build” thinking up to feature, system, and architectural scope.
Use both. TDD at the unit level, SDD at the feature/system level. They don’t compete.
BDD
Behavior-driven development is the most direct ancestor of modern SDD. Gherkin scenarios — Given / When / Then — are literally executable specs. Cucumber and SpecFlow have been turning them into automated tests for over a decade.
What changed: AI agents can now generate code from those scenarios, not just verify against them. This is why people like Bryan Finster (who has been doing BDD for years) call SDD “BDD with branding.” He’s not entirely wrong.
If your team already runs BDD, SDD is the smallest possible step from where you are. You’re already most of the way there.
Contract testing and contract-driven development
Pact and Specmatic have enforced agreements between services for years. Spec Kit folds the idea into SDD as contract-driven development: when a component exposes an interface that something else consumes — an API, a library, a CLI whose output gets parsed — agree on its observable obligations (inputs, outputs, errors, guarantees such as retry safety) before building either side. One side owns the contract, the agreement is settled during planning, and each side then builds and tests against it, whatever your architecture or repository layout.
PRDs and design docs
These are advisory. Humans read them and write code that hopefully matches. Drift is normal and expected; nobody is alarmed when a 6-month-old design doc no longer matches the code, because nothing in the system enforces alignment.
SDD specs are enforced: the agent builds from them, and a mismatch between spec and code is flagged and has to be resolved.
This is the cultural shift the team has to internalize: documents that nobody enforces become decoration. Documents that the workflow enforces become contracts.
Agile / Scrum
User stories with acceptance criteria are specs. The Definition of Done is a form of spec. You’re already specifying things. SDD asks you to do that with more discipline and to make the specs authoritative rather than advisory.
You don’t need to change your sprint structure to adopt SDD. You change what counts as “done” for a story and what artifacts get versioned.
6.When not to use SDD
There are tasks where spec ceremony costs more than it saves. The honest list, drawn from Augment’s “skip the spec” criteria:
| Skip the spec when… | …because |
|---|---|
| The work is exploratory or experimental | You don’t know what you want yet; specifying prematurely will lock in the wrong shape |
| A single prompt produces usable output | Spec overhead would exceed the work itself |
| The output can be reviewed in under five minutes | You can verify by reading; you don’t need a contract |
| You’re building a throwaway prototype | Maintenance benefits don’t apply |
| The change is mechanical or low-risk | Bumping a version number doesn’t need a spec |
The trigger heuristic that’s worth memorizing:
If you’d be annoyed to have the agent interpret requirements differently than you meant, write the spec. If you could fix the output in a quick follow-up prompt, skip it.
This is calibrated by stakes, not by intrinsic complexity. A two-line auth token check might warrant a spec because misinterpretation has security consequences. A 200-line refactor of a private function might not, because if the AI misreads, you’ll catch it on diff review.
Skipping the spec doesn’t have to mean skipping structure. A defect with a clear symptom has its own lighter path — Spec Kit’s bug workflow, which you’ll run in Module 6 — and an idea nobody has vetted yet is better assessed (Module 5) than specified.
7.The honest case against (and why we’re doing it anyway)
Spec-driven development has serious critics, and you should hear them before adopting any practice. This is the part most vendor blog posts skip.
Birgitta Böckeler (Thoughtworks): “Verschlimmbesserung”
Böckeler used Spec Kit, Kiro, and Tessl on real problems and found that current SDD tools generate too many files for the size of the problem they solve. Applied to a small bug fix, Spec Kit produced 4 user stories and 16 acceptance criteria when one prompt would have been better. She uses the German word Verschlimmbesserung — “an improvement that makes things worse” — for SDD setups that generate more markdown than code.
Her concrete warning: reviewing markdown can be harder than reviewing code. If your specs are repetitive, contain code-like detail, or duplicate what the code already says, you’ve made your job harder, not easier.
Take this seriously. Spec Kit now offers a short path for small features and a separate, lighter workflow for bug fixes, but the warning still applies to anything you push through the full loop. We’ll cover countermeasures in Module 8 (Best Practices & Pitfalls).
Bryan Finster: “BDD with branding”
Finster’s argument is that the core insight of SDD — write specs first, derive code from them — has been agile/BDD wisdom for two decades. Calling it a paradigm shift in 2026 is marketing.
He’s right that the insight isn’t new. What’s new is:
- AI agents are now competent consumers of specs, not just humans.
- CI/CD has matured to the point where spec enforcement is routine, not exotic.
- Tooling for spec authoring and validation is finally good enough to be worth adopting.
Together, those three changes make spec-first practice more powerful and more enforceable than ever before. So the practice is more valuable now than it used to be, even if the idea has been around for years.
The honest position: SDD isn’t revolutionary, but it’s currently the most practical way to keep AI-generated code aligned with intent in a production codebase.
Why we’re adopting it anyway
Three reasons, in order of importance:
- It addresses a real problem. When AI agents generate code, the gap between “what we wanted” and “what we got” widens unless something keeps both ends explicit. Specs are that something.
- It scales review better than PR-time inspection. The Agoda study (10,000+ devs, March 2026) found that high-AI-adoption teams completed 21% more tasks but PR review time grew 91%. Pushing validation upstream into specs is one of the few ways to stop that bottleneck from breaking the team.
- It improves cross-functional collaboration. Done well, specs become the shared interface between PM, architecture, engineering, and QA. Done badly, they’re just more documentation. Modules 9 and 10 are about making sure we land on “done well.”
8.The mental shift to make
Engineers tend to enter SDD thinking it’s a tooling change: install SPECTRA, learn nine slash commands, done. That’s the version that produces Verschlimmbesserung.
The mental shift that actually unlocks the value:
- Before: Code is the source of truth. Documents describe code.
- After: Specifications are the source of truth. Code implements specs.
In the “before” view, drift is inevitable and you accept it. In the “after” view, drift is a defect and you fix it.
You don’t need to fully internalize this on day one. You’ll get there by doing the labs. But notice when you’re slipping back into “the code is the truth, the spec is documentation” thinking — that’s the failure mode that turns SDD into bureaucracy.
Knowledge check
Spend 10–15 minutes on these. Click an answer to see whether you got it right and why.
/speckit-specify?spec.md first and regenerate the plan and tasks; flow-back would edit whichever artifact the change touches and reconcile by hand. Linking back answers flow-forward’s risk, scattered context — and the whole convention belongs in the constitution./speckit-checklist, and it generated a custom checklist of twelve items about your spec. Who marks them [x], and when?/speckit-checklist never ticks its own items, and a tick means a person decided the requirement is well written, not that code is done. That rules out a and b. Nor is it ignored (d): /speckit-implement counts unchecked items and asks before continuing, but never changes a checkbox. Don’t confuse this with the built-in checklists/requirements.md, which specify and clarify tick themselves — those ticks are the agent’s claim, not a sign-off./speckit-converge finishes by reporting that it appended three tasks under a Convergence section of tasks.md. What do you do next?tasks.md left untouched — means it’s time for review and the pull request.“Create a function calleduploadFile(file)that callsmulter.single('file'), then writes to S3 usingputObject(), then inserts a row into theuploadstable with columnsid,path,created_at.”
/speckit-plan), not the spec (/speckit-specify). A behavior-level rewrite would be: “Users can upload files up to X MB; uploaded files are stored durably and tracked with timestamps; users can later list their own uploads.” Object storage, S3, and the table layout are choices the team makes later, when it writes the plan.What’s next
In Module 2 you’ll look inside a good spec: the six-element framework for writing specs that don’t bloat or drift, and how the constitution keeps cross-cutting rules out of individual specs. You’ll finish with a short writing exercise — rewriting a deliberately bad spec into a good one.
After that, Module 3 moves into the first hands-on lab: installing SPECTRA and running the workflow on a small greenfield feature.