Learning objectives
By the end of this module, you will be able to:
- Explain why specs catch failures that unit tests miss.
- Distinguish the workflow’s gates, which test the artifacts, from contract, BDD, and property-based tests, which test the code — and use the SpecOps loop to improve both.
- Place SPECTRA’s testing agents where they genuinely help.
- Connect the role split (Product / Architecture / Engineering / Quality) to the human gates each role owns, and identify which role you fill.
- Anticipate the review-bottleneck risk that comes with parallel agent work, and its countermeasures.
- Lead a structured discussion with your team about how SDD would change your existing workflows.
1.Why specs catch what unit tests miss
The labs in Modules 3 and 4 produced working code with passing tests. That’s necessary but not sufficient. Unit tests verify individual functions; they don’t catch the failure modes that cause production incidents.
Three categories of failure that specs catch and unit tests don’t:
1.1 Architectural violations
The function works correctly but is in the wrong place. Auth checks in route handlers instead of middleware. Business logic in database access modules. Direct service-to-service calls that should go through a message bus.
Each instance is fine in isolation. The codebase eroding over time isn’t.
A constitution principle can state “auth checks run in middleware, not handlers,” and a CI rule can enforce it. /speckit-analyze treats a conflict with the constitution as critical before any code exists. Unit tests can’t catch it at all, because each individual handler “works.”
1.2 API contract drift between services
Service A’s caller assumes a userId field. Service A’s implementation now returns user_id. Both services have green CI; the integration is broken.
Tools like Specmatic and Pact validate that implementations match an OpenAPI or Pact spec in CI. A failing schema check fails the build before the broken contract reaches production. The Piskala paper cites a financial services case study where this approach reduced API integration cycle time by 75% — most of the integration time had been spent debugging spec/implementation mismatches that were now caught at spec review.
1.3 Security anti-patterns that emerge across boundaries
The single most-cited example in SDD literature: a payment endpoint without an idempotency key. Each individual function does its job correctly. Retry logic — also doing its job correctly — creates duplicate charges in production. The team patches the code. The next AI regeneration cycle reintroduces the same vulnerability because nothing written down encodes the constraint.
Write the constraint where every future change will read it — for a project-wide rule like “all payment-modifying endpoints require an idempotency-key header,” that’s the constitution — and back it with a CI check that fails when an endpoint is added without one. The vulnerability stays fixed.
The pattern: vulnerabilities that emerge from interactions between correctly-functioning components are exactly the ones unit tests can’t catch and specs can.
A note on the empirical case
The pro-SDD literature cites strong numbers — Augment references benchmarks where LLMs generate vulnerable code at 9.8–42.1% rates; Piskala cites controlled studies showing “error reductions of up to 50%” with human-refined specs. Take them with calibrated trust: they come from contexts that overlap with SDD but aren’t pure tests of it. The directional claim is solid; the exact percentages are weaker. The honest framing for your team: SDD measurably improves quality on dimensions unit tests don’t cover. The size of the improvement depends on your domain, your team practices, and what you were doing before.
2.The quality mechanisms
Each mechanism below addresses a different failure mode. Use them together; none alone is sufficient. Start with the ones already built into the workflow, because they test something your test suite never will: the artifacts themselves.
2.1 The workflow’s own gates test the artifacts
| Gate | What it tests | What it can’t tell you |
|---|---|---|
speckit.checklist | Requirement quality — is the spec complete, clear, consistent? Unit tests for your requirements. A person ticks each custom checklist item; the command never does. | Whether the code works. It never looks at the code. |
speckit.analyze | Whether spec, plan, and tasks agree — a task no requirement asked for, a plan choice the spec rules out, a conflict with the constitution. Read-only; each finding is fixed in the step that owns it. | Whether the spec describes what users actually need. |
speckit.converge | Whether the code covers everything the spec, plan, and tasks call for. Append-only: it can add tasks, never edit code. | Whether the behavior is right at runtime, at the edges, under load. That takes tests. |
These are the cheap, early gates: each catches a whole class of problem before a single test runs. None of them executes anything, so they complement the mechanisms below rather than replacing them.
SPECTRA adds a bridge between the two layers. speckit.spectra.test-plan runs after /speckit-specify and before /speckit-plan, writing a test-plan.md beside the spec — risks, test conditions traced to acceptance criteria, exclusions, exit criteria — for stakeholders to approve before anything is designed. It is approved, not tracked, so it carries no checkboxes; the approved plan goes to the planning command so the tasks include its tests.
2.2 Contract testing
A spec describes the contract between a service and its callers (or between an API and its consumers). A contract test verifies the implementation matches it. The contract is the source of truth; the implementation is what gets verified.
Tools: Specmatic (OpenAPI-based), Pact (consumer-driven contract testing). Both run in CI and fail builds where the implementation doesn’t match the contract.
When this is most valuable: services with multiple consumers, public APIs, microservice boundaries. Less valuable for pure UI work or single-consumer internal tools. Whether you need it is a project-level decision, which is why speckit.spectra.test-strategy (section 3) asks whether an external consumer exists and covers API contract testing as one of its layers.
2.3 BDD scenarios as executable acceptance tests
Gherkin scenarios — Given / When / Then — are executable specs at the feature level. Cucumber, SpecFlow, and similar tools turn them into automated tests. Spec Kit’s spec template already writes each user story’s acceptance scenarios in the same Given/When/Then shape, so the step from spec to scenario is short:
Feature: Adding a due date to a todo
Scenario: Setting a due date on creation
Given I am on the todo app
When I add a todo "Pay taxes" with due date "2026-04-15"
Then the todo "Pay taxes" should appear in my list
And it should display its due date "2026-04-15"
Scenario: Overdue todos are visually distinct
Given I have an uncompleted todo with due date "2025-12-01"
When I view my todo list
Then the todo should appear with the overdue treatment
The killer property: PM and engineering agree on what “done” means before any code is written. A feature is done when all scenarios pass. There’s no debate at the end of the sprint about whether the work matches what was asked for, because the scenarios are what was asked for.
When this is most valuable: features where business stakeholders care about behavior rather than implementation. UI features, business rules, customer-facing workflows. Less valuable for infrastructure, refactors, performance work.
2.4 Property-based testing for spec invariants
LLM-generated code is non-deterministic. The same spec can produce different code on different runs. Unit tests with hand-written examples may pass for one generation and fail for another.
Property-based tests assert that something holds for all inputs of a given shape, not just hand-picked examples. Tools: QuickCheck (Haskell), Hypothesis (Python), fast-check (JavaScript).
For example, in the to-do app: “for any todo with a due date in the past and completed = false, isOverdue(todo) returns true.” That’s an invariant — it must hold for any input matching the precondition. A property-based test generates hundreds of inputs and verifies the invariant for each.
When this is most valuable: code with clear invariants, especially anything involving math, data transformations, or state machines. Less valuable for purely UI-driven features.
2.5 The SpecOps feedback loop
Module 4 introduced this in a paragraph at the end of the lab. This is where it gets explained, because it is the mechanism that makes every other gate improve over time.
When a bug surfaces, first establish what actually happened. Spec Kit’s bug workflow (Module 6) gives you a diagnosis backed by evidence; speckit.spectra.defect-rca goes further, to root cause and preventive actions. Then classify the bug:
- Spec-to-implementation gap: the spec was clear, the code diverged. Strengthen the gates that should have caught it — a missing test, a contract check, a CI rule, a constitution principle that analyze and PR review can check against.
- Intent-to-spec gap: the spec itself was incomplete or wrong. Improve elicitation — the questions you push through clarify, the domains your checklists cover, the sections in your spec template (Module 7 shows how to override it for the whole team).
Both land back in the shared context. In SPECTRA’s terms this is the Maintain gate: the team decides the fix, and whether the lesson becomes a standard. This is the mechanism that turns SDD from a one-time productivity bump into a compounding capability. Without it, you plateau. With it, every bug strengthens the system.
A practical habit: in retros, classify recent bugs into the two categories and track the ratio. Many spec-to-implementation gaps mean your validation needs work. Many intent-to-spec gaps mean your spec-writing process does. Each tells you a different thing to fix.
3.Test coverage strategies under SDD
Decide the strategy once. speckit.spectra.test-strategy runs once, alongside the constitution: it reads the project, asks only what the repository can’t answer, and writes <artifact-root>/test-strategy/TEST_STRATEGY.md — unit, integration, API contract, and end-to-end testing, plus a coverage floor the codebase can actually hold. It drafts the constitution amendment that would make the strategy binding and hands it to /speckit-constitution, never editing the constitution itself. Architects and engineering leads approve it at the Foundation gate. Three patterns then come up feature by feature:
3.1 Test-first specs
Include test requirements as part of the spec. The agent generates tests alongside the implementation. The spec defines the contract; the tests prove it. When tests are requested, /speckit-tasks puts them inside each user story’s phase, next to the code they prove.
This is TDD adapted for AI-generated code. The mental model: you don’t trust the agent’s implementation, but you trust the spec. So the tests should be derivable from the spec, not from the implementation — which is exactly what a test plan from speckit.spectra.test-plan records.
If your spec is well written (Module 2’s six elements, including verification criteria), your test suite is mostly already there. Just have the agent codify it.
3.2 Edge cases in the spec, not discovered in implementation
A bad pattern that’s hard to break: write a happy-path spec, get a happy-path implementation, discover an edge case during manual testing, patch the code.
The edge case should have been in the spec. If you’re discovering edge cases during implementation, your spec is incomplete. Update the spec, then update the code.
Practical exercise during spec writing: ask “what would break this?” for each requirement. Empty input. Maximum input. Concurrent operations. Network failures. Timezone changes. Rounding boundaries. Each potential break either becomes an explicit verification criterion or a deliberate out-of-scope decision. Both are fine; silence isn’t. /speckit-clarify with a focus area, and a checklist aimed at edge cases, are the workflow’s tools for asking.
3.3 Validation gates in CI — the same gates, but stricter
Whatever validation your team runs on human-written code, run on AI-generated code: minimum coverage thresholds, SAST, dependency scanning, secrets detection, integration tests, regression tests.
The honest reason: AI-generated code has higher baseline error rates than carefully written human code. Treating it as if it had human-equivalent quality skips the validation that made human-written code reliable. If your CI doesn’t have these gates today, the moment you start generating code with AI is the moment to add them.
A gate only protects you while people believe it. Once a suite fails intermittently, the team re-runs instead of investigating, and real regressions hide in the noise. speckit.spectra.flaky-test-detector finds flaky tests by reading the test source — unconditional sleeps, un-awaited async calls, shared state, the real clock — without running anything, and with your explicit go-ahead fixes the cause. A retry, a skip, or a loosened assertion never counts as a fix.
4.Team roles under SDD
The cultural shift is the actual work. The InfoQ piece by Krishnan is the most explicit on this point: the most significant impact of SDD is cultural, not technical. The shift is from one-way instructions (“do X”) to dialogue (“here’s the intent — let’s converge on what we’re building”). Specs become the shared interface where product, architecture, engineering, and quality collaborate.
If your team treats SDD as “engineers write better prompts,” you’ll get token-usage improvements and not much else. Treat it as a way to restructure how PM, architect, and engineer collaborate, and you unlock the bigger benefit: directing parallel agent work without losing coherence. SPECTRA makes the split concrete. Its agentic SDLC has a human gate in every phase — agents draft, people decide — and each gate belongs to a role:
| Role | Owns | Human gates it stands at |
|---|---|---|
| Product | What and why | 01 Plan — approves intent, scope, and business alignment. Answers clarify’s questions, and decides whether an assessed idea goes ahead at all (Module 5). |
| Architecture | How | 00 Foundation — with engineering leads, approves every standard before it binds an agent. 02 Design — approves the plan and the decisions it records. |
| Engineering | Tasks and code | 03 Implement — reviews every change; an agent’s diff is read like a person’s. 05 Deploy — a maintainer merges. |
| Quality | The harness | 04 Test — decides what a finding means and whether a fix is safe to land. |
| The whole team | Lessons | 06 Maintain — decides the fix, and whether the lesson becomes a standard. |
4.1 Product owns “what”
Business context, user value, acceptance criteria. Product usually works in the backlog (Jira, Linear, Azure DevOps), not in Git — and doesn’t have to move. SPECTRA’s speckit.spectra.brd takes a raw requirement, typed or handed over as a document, and turns it into a specify-ready business requirements document.
Product writes the intent layer of the spec: outcomes, in-scope behavior, business constraints. They don’t write the technical sections. They review and challenge them, and they are the person clarify’s questions are really addressed to.
A PM’s ticket with acceptance criteria is already most of the spec’s first section; it just becomes the canonical input rather than advisory text.
4.2 Architecture owns “how”
Technical approach, repository boundaries, integration patterns, cross-cutting constraints. Architecture encodes these into a reusable harness — the constitution (SPECTRA calls the agent that writes it Guardrails) and architecture decision records from speckit.spectra.adr — rather than manually decomposing each story.
Architecture’s biggest contribution under SDD isn’t writing individual specs; it’s curating the constitution and the validation gates so that every spec inherits sound technical decisions. A 50-line constitution improvement cascades through every future spec.
If your team doesn’t have a designated architect, this role still needs to exist — a rotating senior engineer, a tech lead, a small “tech council.” Without it, every spec becomes a one-off, and architectural drift accelerates under SDD.
4.3 Engineering owns “tasks”
Per-repository implementation tasks, tightly coupled to the codebase, living with the code. This is what most engineers were doing before SDD; the change is that the what and the how are now explicit upstream artifacts rather than implicit assumptions.
The big shift for engineers: less time writing code, more time writing specs and reviewing artifacts at every phase. This feels uncomfortable to engineers who derive identity from typing code. The honest reframe: you’re still doing engineering work, at a higher level of abstraction. The code is now an output, not the primary deliverable.
This isn’t deskilling. Writing a precise, testable spec an agent can implement correctly is a harder skill than writing the code yourself.
4.4 Quality owns harness validation
Krishnan’s key shift: QA validates the harnesses and validation mechanisms, not the finished implementations.
In the old model, QA caught bugs at the end. In the SDD model, the validation gates and the spec templates are the QA artifact. In SPECTRA’s terms, that means the test strategy, each feature’s test plan, and often the reviewer’s seat on the custom checklists — all settled before code exists. QA’s job is to make those better, so that a wider class of bugs can’t reach production in the first place.
Visual regressions, complex user flows, and exploratory testing still happen by hand, but the bulk-validation work shifts upstream. QA is often the slowest function to absorb this, because its identity has been “find bugs late”; “make the harness catch more, sooner” is a genuinely different mental model.
4.5 Specialised roles
Security, performance, infrastructure: each layers its own context onto incoming stories through the harness, rather than serving as a manual review gate on every change.
Example: security writes its policies (auth requirements, data handling rules, encryption standards) into the constitution. /speckit-analyze then flags any spec, plan, or task list that conflicts with them before implementation, and speckit.spectra.review-pr checks each pull request against the same constitution. If the spec is fine, security signs off without being a bottleneck. If it has issues, they’re caught at spec review, not at PR time.
This scales much better than the old model where every PR needed a security reviewer. Specialised reviewers become consultants on the harness rather than gatekeepers in the workflow.
5.Multi-repo and multi-team coordination
The single biggest gap in current SDD tooling. Spec Kit’s conventions are per repository — even Module 8’s spec-of-specs roadmap ties slices together inside one repo. Modern enterprise architectures span microservices, shared libraries, and infrastructure repos.
Krishnan’s recommended pattern:
- Business context (the “what”) lives at the backlog level, visible to the whole org.
- Parent stories get decomposed into repository-specific sub-issues during the design phase. Each sub-issue becomes a spec in its own repo.
- Technical implementation details live with the code in each repo.
- Architects maintain repository-boundary documentation that agents can superimpose onto incoming stories to generate appropriate sub-issues per repo.
If your team is single-repo today, this is mostly a future concern. If you span multiple repos, plan for it now: the longer you wait, the more divergent each repo’s specs and constitutions become.
6.The review bottleneck
The most counterintuitive risk in SDD adoption: it can make your team slower, not faster, if you don’t address review capacity.
The Agoda study (10,000+ developers, 1,255 teams, March 2026) is sobering. High-AI-adoption teams completed 21% more tasks and merged 98% more PRs — but PR review time increased by 91%. From the InfoQ piece on the study:
If we keep ourselves preoccupied verifying AI output, the backlog is starved for fresh ideas. A review-based approach alone cannot scale.
Why this happens: agents generate code far faster than humans can carefully review it. Each PR is bigger, more frequent, and less familiar to the reviewer. Reviewers become the bottleneck.
SDD’s countermeasures push validation upstream:
- Spec review — clarify and the custom checklists — catches intent-level issues before code is written. Cheaper than catching them at PR.
- Plan and tasks review, with
/speckit-analyzecross-checking them, catches design issues before code is written. - Verifier passes (Module 8) —
/speckit-converge, and an adversarial verifier for high-stakes work — check generated code against the artifacts before a person does. - Validation gates (CI checks tied to specs) catch divergence automatically, so the reviewer can focus on judgment-level concerns.
- A conformance-aware first pass at the PR.
speckit.spectra.review-prreviews the pull request against the spec, plan, tasks, ADRs, and constitution it carries — catching a task marked done but never built, or a change no requirement authorized — and proposes only blockers and majors by default. The human reviewer decides what gets published.
Without these, AI adoption plus manual PR review equals the Agoda outcome. This is also why the role split matters. If quality engineering moves upstream into the harness, and architecture moves upstream into the constitution, code review becomes a smaller fraction of the work — not because it’s done less carefully, but because the upstream work makes it easier.
7.Discussion prompts for your team
Spend 15 minutes on these in a team setting (or a Slack thread, or an async doc — whatever fits how your team operates). The goal is to surface the specific frictions your team will hit, not to produce a polished document.
- Where in our current workflow is the spec implicit? When a PM hands a ticket to engineering, what’s understood that isn’t written down? Whose responsibility is it that those assumptions don’t drift?
- Who would approve our constitution today? Realistically — a name, not a role. That person stands at the Foundation gate. If you can’t name someone, that’s the first organizational gap to close before serious adoption.
- What’s our current review bottleneck? PR review time, design review time, security review time? Of those, which would SDD push upstream, and what’s the work to make that happen?
- What’s our spec-rot risk? If we adopt SDD and the spec becomes part of every PR, what would actually keep specs current six months later? Who enforces it? What gate would catch it?
- Where would we hit the multi-repo problem? Are there features today that span repos? How do we coordinate them currently? How would SDD make that easier or harder?
- What’s the smallest pilot we could run? Pick one team, one type of feature, four weeks. What would success look like? What would failure look like? What would we measure?
The last question is the most important. The next module shows what the lifecycle looks like once this way of working is the default — but the questions above need answers from your team specifically. No external framework can substitute for that judgment.
What's next
Module 10, AI-DLC — the final module — is about the lifecycle itself. AWS’s AI-Driven Development Life Cycle folds the SDLC into three stages, replaces sprints with bolts measured in hours or days, and puts humans and AI in the room together at every stage. The module maps SPECTRA’s seven-phase agentic SDLC, and the commands you already know, onto it.
If this module was about how SDD changes collaboration within the existing lifecycle, the final module is about what happens when the lifecycle itself catches up.