Module 8

Best Practices & Pitfalls

Patterns that consistently produce good outcomes, how to size work so the agent doesn’t lose the thread mid-implementation, the nine failure modes that sabotage adoption, and a critique exercise on three deliberately broken specs.

Time 60 min
Format Reading + critique exercise
Prerequisites Modules 1–4

Learning objectives

By the end of this module, you will be able to:

  1. Apply the human-reviewability test to decide when a spec is too big.
  2. Run the workflow’s quality gates as decisions with a named owner, not as formalities.
  3. Use the adversarial verifier pattern, and tier AI models by role.
  4. Size implementation to the agent’s context window, reaching for a spec of specs only when nothing lighter works.
  5. Recognize the nine most common pitfalls before they consume a sprint.
  6. Critique a poorly written spec and suggest specific, actionable fixes.

1.Best practices that consistently work

The labs in Modules 3 and 4 exposed you to a lot of small judgment calls, and Modules 5 to 7 added the work around a feature. This section names the patterns that repeatedly produce good outcomes across teams that have been doing SDD for a while.

1.1 Human reviewability is the master test

If you find yourself skimming a spec thinking “the AI probably got it right,” the feature is too large.

That single sentence (paraphrased from intent-driven.dev) is the most useful heuristic in SDD. It applies at every phase. If the spec is too long to read carefully, slice it. If the plan has too many components to think through, break it. If the task list is too long to review, put fewer tasks in each spec.

Two practical sub-rules:

  1. Hard cap on per-feature spec content. Around 500 words of substantive content (not counting boilerplate). Hard cases get exceptions, but treat them as exceptions.
  2. Decompose using INVEST. Independent, Negotiable, Valuable, Estimable, Small, Testable. If a feature fails any of these, it’s a candidate for slicing into smaller specs.

The human-reviewability test isn’t a guideline; it’s the thing that prevents the markdown-monster failure mode. Take it seriously. Section 2 covers the agent’s side of the same problem: a feature can be perfectly reviewable by you and still too big for one implementation run.

1.2 Specify just enough to remove ambiguity

Piskala calls this the “Golden Rule” of SDD: use the minimum specification rigor that removes ambiguity for your context. Not maximum. Not comprehensive. Minimum.

The trigger from Augment is worth committing to memory:

If you'd be annoyed to have the agent interpret requirements differently than you meant, write the spec. If you could fix the output in a quick follow-up prompt, skip it.

Three implications:

1.3 Treat every gate as a decision, not a formality

Spec Kit’s full path puts a quality gate between most steps. A gate only works if someone actually stands at it:

GateWho decidesThe habit that makes it work
/speckit-clarifyWhoever holds the intent — usually the product ownerAt most five questions per pass, answers written into spec.md. Run it again with a focus area rather than hoping one pass caught everything.
/speckit-checklistThe reviewerA custom checklist item is ticked [x] only by a person; the agent helps evaluate when asked. (The built-in checklists/requirements.md is different: specify and clarify tick it themselves, so read those ticks as the agent’s claim, not a sign-off.) /speckit-implement asks before continuing past unchecked items and never ticks any.
/speckit-analyzeThe owner of the step the finding belongs toFix each finding at its source — requirements in specify or clarify, design in plan, the task list by re-running tasks — then re-run analyze until clean. Hand-editing tasks.md to quiet a finding just hides it.
/speckit-implementThe engineer reviewing the diffVerify each stage — run the app, run the tests, read the change — before starting the next.
/speckit-convergeThe engineerIt either reports Converged or appends tasks under a Convergence section of tasks.md. Implement those and converge again until Converged. Skip it and a “done” feature ships with a requirement nobody built.

Agents draft, people decide. A gate nobody owns is just one more step the agent runs on its own.

1.4 Use adversarial agent patterns when stakes are high

For high-stakes features (security boundaries, payment flows, anything in regulated domains), a single agent reviewing its own output is overconfident by design. Use multiple agents with opposing roles:

The Verifier’s opposing goal is the load-bearing part. Implementing agents are trained to be helpful, which means they’re optimistic about their own output. A Verifier with the opposing instruction — “find every way this output fails to match the spec” — catches mistakes the implementor missed.

You already have a narrow, built-in version: /speckit-converge exists only to find where the code falls short of the artifacts, and it can’t fix its way to a pass, because its only possible write is appending tasks. For high-stakes work, go further and give a second agent the explicit brief to break the implementation.

A practical signal you need this pattern: you’ve shipped a bug the spec clearly ruled out, the agent generated it anyway, and review missed it. The pattern also forces your spec to be more explicit, because the Verifier has to be able to check it. Bring the discipline upstream.

1.5 Tier your models by role

Different phases need different model capabilities, and using the same model for everything either wastes money or sacrifices quality.

RoleModel tierWhy
Spec writing / refiningYour most capable modelErrors here propagate to every later phase
Plan and tasksYour most capable modelSame logic — design errors are expensive
ImplementationA strong coding model a tier below your best — faster and cheaperCode generation is well served by mid-tier models; the spec and plan already constrain the output heavily
VerificationA fast, inexpensive model for mechanical checks; step back up for adversarial review of high-stakes workChecking is a narrower task than generating; smaller models do routine checks well at much lower cost

The mistake to avoid: using a top-tier model for implementation and a cheap one for spec authoring. That’s exactly backward. A bad spec can’t be saved by a good implementation; a good spec can be implemented well by a mid-tier model.

Model names and prices change every few months; the shape of this table doesn’t. If your team is locked into a single model, default to “good enough across all phases” rather than the cheapest option — the spec phase has the most leverage and is least forgiving.

1.6 Pair specs with project-level context

Every serious SDD tool has a place for durable project context: Spec Kit’s constitution at .specify/memory/constitution.md (SPECTRA calls the agent that writes it Guardrails), Kiro’s steering files, the more generic AGENTS.md convention. Use it.

The right division (revisiting Module 2):

Don’t cram constitution-level conventions into every spec “to be safe.” That’s how specs bloat into markdown monsters. Trust the constitution. If a convention is missing, fix the constitution once instead of repeating it in every spec.

Two maintenance habits. Write down only what is already true or explicitly agreed — invented standards become noise in every later plan and analysis. And review it quarterly: a stale constitution is worse than none because it teaches the agent the wrong patterns.

1.7 Apply systems thinking for cross-feature analysis

Specs are usually authored one feature at a time. But features interact in non-linear ways. The classic example (intent-driven.dev): an inventory-alert feature combined with an existing notification system can produce notification storms during peak traffic — a problem only visible if you analyze the specs together as a system, not as isolated features.

Practical implication: when a spec touches a system area that other specs also touch (notifications, auth, data egress), do a brief cross-feature review before implementing: “do any other features in this area cause the new one to behave badly, or vice versa?” SPECTRA’s speckit.spectra.impact helps by reporting what a proposed change would actually touch, with cited path:line evidence. The Architecture role from Module 9 owns this review; if nobody holds that role yet, designate someone for the duration of the feature.

1.8 Commit at every phase boundary

Commit your spec, plan, and tasks separately, before generating code. This gives you natural rollback points and makes the artifacts greppable in Git history.

The minimum hygiene:

git commit -m "spec: <feature>"                 # after specify + clarify
git commit -m "plan: <feature>"
git commit -m "tasks: <feature>"                # once analyze is clean
git commit -m "feat: <feature> (spec-driven)"   # once converge says Converged

If the implementation goes sideways, git reset --hard <tasks-commit> throws away the generated code and you can re-run /speckit-implement without losing the upstream work. If the task list itself was the problem, reset to the plan commit and re-run /speckit-tasks first.

2.Sizing work to the context window

Section 1.1 sizes work for the human reviewer. This section sizes it for the agent, and the two limits are not the same.

2.1 Why big features fail mid-implement

The pattern is recognizable. Specify, plan, and tasks go smoothly. Then, deep into a long /speckit-implement run, the agent starts to drift: it skips tasks, marks work done that isn’t, contradicts a plan decision it followed an hour ago. It often happens right around the moment the agent compacts its context — summarizing the conversation so far to make room.

The cause is context-window exhaustion. A run that tries to hold the spec, the plan, the task list, and every file it has touched fills the window, and the model’s grip on the early material weakens. Compaction then swaps detail for a summary, and the detail it drops is often the constraint that mattered. The fix is not a better prompt. It is a smaller run.

2.2 Four options, lightest first

Each option adds overhead, so stop at the first one that works.

1 — Scope each implement run. /speckit-implement considers whatever you type after it, so you can bound a run by task range or by phase:

/speckit-implement only execute tasks T001-T010, then stop and report progress
/speckit-implement only execute Phase 3 (User Story 1), then stop

Completed tasks are marked [X] in tasks.md, so the next run resumes where this one stopped. Between runs, stand at the gate: run the app and the tests, read the diff. This works with any agent, and it is enough for most features.

2 — Delegate parallel tasks to sub-agents. If your coding agent supports sub-agents, ask implement to hand each [P] task to one. Each sub-agent starts with one task and the parts of the plan it needs, so the main session never fills up; [P] tasks touch different files and don’t depend on each other, which is what makes them safe to farm out.

/speckit-implement delegate each [P] task in this phase to a sub-agent

3 — Combine both. Scope the run to one phase and delegate its [P] tasks: the scope keeps the main session small, delegation keeps each task small.

4 — A spec of specs. When even one phase is too much for one run, the feature is really several features. This is the last resort.

2.3 The spec of specs

A spec of specs splits an epic into slices that are each specified on their own — own spec.md, plan.md, and tasks.md — tied together by a roadmap. Suppose that after the due-date lab the team wants to turn the to-do app into a lightweight planner: repeating to-dos, reminders, a calendar view, snoozing. That’s an epic. Start with a roadmap pass — a short planning conversation with the agent, not a full spec: state the whole epic; find slices that are independently testable (building any one leaves something you can demonstrate); give each one line of intent and an explicit scope boundary saying what is deferred to a sibling; order them by dependency; and record the result.

The roadmap names and orders the slices; it doesn’t design them. It is an ordinary Markdown file under version control — specs/<epic>/roadmap.md, or a top-level ROADMAP.md for an epic that cuts across the whole product:

# Roadmap: Planner

Turn the to-do app into a lightweight planner, one shippable slice at a time.

Status: planned · in-progress · done

| ID | Slice            | Intent                                    | Scope boundary                               | Depends on | Status      | Spec                       |
|----|------------------|-------------------------------------------|----------------------------------------------|------------|-------------|----------------------------|
| R1 | Repeating to-dos | Repeat a to-do daily, weekly, or monthly  | Next one appears on completion; no reminders | —          | done        | specs/002-repeating-todos/ |
| R2 | Reminders        | Remind the user before a to-do falls due  | One reminder per to-do; no snooze            | —          | in-progress | specs/003-reminders/       |
| R3 | Calendar view    | See to-dos laid out by due date           | Read-only; no drag-to-reschedule             | —          | in-progress | specs/004-calendar-view/   |
| R4 | Snooze           | Push a reminder back by a chosen interval | Reminders only; the due date never moves     | R2         | planned     | —                          |

Spec C in the critique exercise below is what you get when a roadmap is written as if it were one spec.

3.Pitfalls that consistently sabotage SDD adoption

The labs probably surfaced one or two of these. The full list:

Pitfall 1 — Over-specification

The spec reads like pseudo-code. Function names, exact error messages, schema column names appear before the plan phase. The abstraction benefit is lost; the agent has nothing to figure out, so the spec is just a slow way of typing the code.

Symptom: you find yourself thinking “this would be easier to just write.”

Cure: strip implementation detail. Ask: does this constrain what should happen, or how? If “how,” it belongs in /speckit-plan.

Pitfall 2 — Spec rot

You meant to keep your specs, but nobody updates them when the code changes. Six months in, the specs are lies, and every agent that reads them is misled.

Symptom: specs in your repo from a year ago that nobody has read since they merged.

Cure: decide how your specs age (Module 1) — spec-first, spec-anchored, or spec-as-source — and record the choice in the constitution so the agents work to it too. If you keep specs, the spec edit is part of the change, not a follow-up. At review time, speckit.spectra.review-pr checks a pull request against the spec it carries and flags a change no requirement authorized.

Pitfall 3 — Specification theater (the markdown monster)

You’re producing elaborate specs that nobody actually reviews. Specs get rubber-stamped because they’re long, and a long doc nobody read is just expensive overhead.

Symptom: specs are getting longer over time and review comments are getting shorter.

Cure: apply the human-reviewability test (1.1). If specs are too long to review, slice them. Reward concision in code review.

Pitfall 4 — False confidence from passing spec tests

A passing spec test only proves code matches spec. If the spec is wrong, the code is faithfully wrong. A green CI on a faulty spec is more dangerous than a red CI on a missing spec, because it produces undeserved trust. The same goes for /speckit-converge: Converged means the code matches the artifacts, no more.

Symptom: a bug ships that “should have been caught” but the spec test passed.

Cure: when bugs surface, classify them with the SpecOps loop (Module 9): intent-to-spec versus spec-to-implementation. Intent-to-spec bugs mean your elicitation needs work; treat that as a higher-priority fix than the bug itself.

Pitfall 5 — Tooling cargo-culting

You’re running the full path — every optional gate, every time — on every change, including mechanical ones. The tool has become a religion.

Symptom: small bug fixes that take a day instead of an hour.

Cure: the “when not to use SDD” criteria from Module 1 are real. Small, clear features can take the short path (1.2). A defect usually doesn’t need a feature spec at all — Spec Kit’s bug workflow from Module 6 is sized for it. Empower engineers to choose.

Pitfall 6 — AI-generated spec bloat

You let the agent write the first-draft spec and accept it without editing. The agent over-includes “to be safe.” The spec is verbose, repetitive, and harder to review than the code it directs.

Symptom: specs averaging 1500+ words for routine features.

Cure: treat agent drafts as input, not output. Edit ruthlessly. The team that writes the shortest clear spec wins.

Pitfall 7 — Premature comprehensiveness

You try to spec the whole system upfront, often when adopting SDD on a brownfield codebase (the temptation Module 4 warned about). The result: a giant spec nobody reads, written before you’ve learned anything from doing SDD on real changes.

Symptom: “we’ll roll out SDD once we finish writing the system spec.”

Cure: never write the system spec. Spec changes incrementally. Frequently-touched areas accumulate spec coverage organically. Premature comprehensiveness is how SDD adoptions die.

Pitfall 8 — “SpecFall”

You install SPECTRA, run a training, and don’t change anything else about how product, architecture, and engineering collaborate. SDD becomes a technical artifact in an unchanged process. Krishnan calls this “SpecFall” — analogous to the classic “Scrumerfall” failure where teams adopt Scrum ceremonies on top of waterfall thinking and wonder why nothing changes.

Symptom: PMs still write Jira tickets exactly as before, engineers translate them into specs in isolation, nothing about cross-functional collaboration has changed.

Cure: treat SDD adoption as a process change, not a tooling change. Specs become the shared interface where PM, architecture, and engineering meet. Module 9 covers this in depth.

Pitfall 9 — Gates in name only

The workflow has all its gates and nobody stands at them. Custom checklists are ticked by whoever ran them, analyze findings vanish because someone edited tasks.md instead of the spec, one long implement run goes straight to a pull request, and converge never runs.

Symptom: every custom checklist item is [x] and nobody can remember ticking one.

Cure: name an owner for each gate (1.3). Only a person ticks a custom checklist item; a finding is fixed in the step that owns it; each implement stage is verified; and the pull request waits until converge says Converged.


Critique exercise

Spend 20 minutes on this. Three deliberately broken specs follow. For each, write a short critique:

Sample critiques are revealable below each spec — work through these yourself before peeking.

Spec A — “User profile picture”

Add the ability for users to upload a profile picture. The function should be called uploadProfilePicture(userId, file). It should resize the image using sharp's .resize(256, 256) method, then save it to S3 using s3.putObject({ Bucket: 'profile-pics', Key: `${userId}.jpg`, Body: resizedBuffer }). Update the users table by setting profile_picture_url = `https://profile-pics.s3.amazonaws.com/${userId}.jpg` for that user.

What’s wrong: This is the over-specification pitfall (Pitfall 1) in its purest form. The “spec” is mostly an implementation. It names a function (uploadProfilePicture), specifies a library (sharp), provides exact API call shapes (s3.putObject(...)), and even hardcodes the URL pattern. There’s nothing for a plan or implementation phase to add.

It’s also missing every element except a thin sketch of in-scope work:

  • No outcomes (what does the user actually experience? Where do they upload? When does the picture appear?)
  • No out-of-scope items (cropping? animations? video?)
  • No verification criteria (what’s the failure case? What if the file isn’t an image?)
  • No constraints (size limits? Allowed formats? Auth?)

How to fix it: rewrite at the behavior level. Something like:

Outcomes: a signed-in user can upload an image to use as their profile picture. The new picture replaces any existing one. The picture appears in the user’s avatar slot wherever it’s currently shown, within 5 seconds of upload.

In scope: JPEG and PNG only, max 5MB, single picture per user, replaces previous on upload.

Out of scope: cropping UI, multiple historical pictures, removing without replacing, gravatar fallback changes.

The implementation details (sharp, S3, schema column) belong in the plan, not the spec.

Spec B — “Better search”

Users want a better search experience. Improve the search to be faster, more relevant, and more user-friendly. Should handle typos and synonyms. The search bar should give a great experience overall. Make it work well on mobile too. Should integrate with the existing search infrastructure where possible. Performance should be acceptable.

What’s wrong: This is verification by vibe (Module 2’s Failure 3 — roughly the opposite of Pitfall 1). The spec is full of words like “better,” “faster,” “more relevant,” “great experience,” “acceptable” — none of which are verifiable. It also lacks every other element of a good spec.

Three specific failures:

  1. No measurable outcomes. “Better” doesn’t tell the agent (or a reviewer) what to build. Faster than what? More relevant by what definition?
  2. Vague scope. Search across what data? Posts? Users? Products? All of the above?
  3. No verification criteria. What test would prove this is done?

How to fix it: make every adjective measurable. “Faster” becomes “p95 search latency under 100ms (currently 400ms).” “More relevant” becomes “top-3 results for the test query set match human-rated top results in 80% of cases (currently 60%).” “Handle typos” becomes “single-character typos in user names produce the correct result.” “Mobile” becomes “the search input is reachable and usable at viewports as narrow as 375px.”

This is exactly the spec the gates exist for. /speckit-clarify would spend its five questions on these adjectives, and a checklist item such as “Is ‘faster’ quantified?” fails on sight. Without measurability, the spec just gives the agent vibes to chase. You’ll get something. It probably won’t be what you wanted. You won’t be able to prove either way.

Spec C — “Reports module”

Build a reports module that lets administrators generate reports about user activity. Reports should include: daily active users, weekly active users, monthly active users, retention curves (1-day, 7-day, 30-day), revenue per user, churn rate, average session length, conversion rates from each funnel step, geographic distribution of users, device type breakdown, and exportable CSV downloads of all metrics. Reports should be schedulable (daily, weekly, monthly) and emailable. Each report should have visual charts. Charts should support drill-down by clicking. Admins should be able to define custom reports by combining metrics. Custom reports should be shareable. Custom reports should support permissioning so different admins can see different reports.

Out of scope: Real-time reports (we’ll add streaming later)

What’s wrong: This is premature comprehensiveness (Pitfall 7) and a failed human-reviewability test (Best Practice 1.1). The “feature” is at least eight features in a trench coat: scheduled reports, exportable reports, custom report builder, charting, drill-down, permissioning, sharing, geographic analytics, conversion funnel analytics. Each one is its own feature, and together they would overwhelm any single implement run.

The out-of-scope list has one item, which is a tell: when in-scope is sprawling and out-of-scope is empty, the author hasn’t drawn the boundaries deliberately. They’ve just listed every feature they could think of.

How to fix it: decompose using INVEST and MoSCoW. A reasonable first slice, R1, might be only “daily active users” — one metric, one chart, no scheduling, no custom reports, no permissioning, no email. Ship it. Then layer:

  • R2: weekly + monthly active users (extends the existing UI)
  • R3: CSV export
  • R4: scheduled email delivery
  • R5: revenue per user
  • R6: retention curves
  • … etc.

Each slice is independently reviewable, independently testable, and independently shippable. None of them is bigger than 500 words.

The original spec is worth keeping, but as a roadmap — the spec of specs from section 2.3, with a stable ID, a scope boundary, and dependencies for each slice — not as a single spec. Specs are not the right document for a roadmap. A roadmap is. Don’t conflate them.


4.Putting it together

The patterns and pitfalls above are descriptive, not prescriptive. They describe what tends to happen, not what you must do. When in doubt, return to the principle behind them:

The goal of SDD is to make intent explicit and enforceable, with the minimum overhead that achieves that goal.

Everything else is implementation detail. Concision over completeness. Review over generation. Slicing over comprehensiveness. Small runs over heroic ones. The pitfalls are mostly variations on “we forgot to apply that principle.” When you notice you’re drifting into a pitfall, ask whether you’re optimizing for the principle or for the appearance of doing SDD. Drop whatever isn’t serving the principle.


What's next

Module 9 turns from individual practice to team practice — how the workflow’s own gates and SPECTRA’s testing agents fit alongside the quality mechanisms you already know, and how SDD changes the collaboration between product, architecture, engineering, and quality, down to who stands at which human gate. It’s the most important module for anyone in a leadership or coordination role, and it’s also the module where most SDD adoptions either succeed or quietly fail.