Module 6

Bug Fix — Assess, Fix, Verify

Take a real defect in the lab app from a support ticket to a verified, committed fix with Spec Kit’s bug workflow. Diagnose it without touching code, fix only what the diagnosis covers, and don’t call it done until someone has re-run the reproduction.

Time 75 min
Format Concepts + hands-on lab
Prerequisites Modules 1–4; SPECTRA installed

Learning objectives

By the end of this module, you will be able to:

  1. Decide whether work belongs in the bug workflow, the SDD loop, or an idea assessment.
  2. Explain the trust boundary of each bug stage — speckit.bug.assess, speckit.bug.fix, speckit.bug.test — and why only one of them may edit code.
  3. Review an assessment at the human gate, and approve it or send it back.
  4. Check that a fix stayed inside its assessment.
  5. Read verified, partial and failed honestly, and supply the evidence a partial is missing.
  6. Know when to escalate to SPECTRA’s speckit.spectra.defect-rca, and how to deliver a verified fix with /speckit-spectra-create-pr.

Commands use the Claude Code / GitHub Copilot trigger form (/speckit-bug-assess); Module 1 has the spelling for other agents.

1.Is it a bug? Choosing the workflow

Modules 3 and 4 ran the SDD loop. Module 5 decided whether an idea deserves the loop at all. This module covers the work most teams do most often: something that should work doesn’t.

The work in front of youUseBecause
Existing behaviour is broken — a crash, a wrong result, a regressionThe bug workflow: assess → fix → testYou are restoring intended behaviour. You need evidence of what’s wrong and a repair scoped to it.
A new capability, or a deliberate change in behaviourThe SDD loop (Modules 3 and 4)Someone has to decide what the product should do. That’s a spec.
An idea nobody has shown is worth buildingIdea assessment (Module 5)The question is whether to build, not how.

The ticket’s label won’t always tell you which row you’re in. “I can’t set a reminder on a task”, filed as a bug, is a feature request. The bug workflow catches this: an assessment can come back invalid — expected behaviour, misuse, a duplicate, out of scope — and new behaviour goes to the SDD loop instead.

In SPECTRA’s agentic SDLC, bug fixing belongs to the Maintain phase, where the human gate is the team deciding the fix and whether the lesson becomes a standard. It needs no SDD first: a team can run it straight after specify init on a codebase with no specs at all. That’s exactly what you’ll do in the lab.

2.Three stages, three trust boundaries

Spec Kit’s bug extension ships with Spec Kit and is opt-in: specify extension add bug installs it without network access. It splits a repair into three commands. You run each one, review its report, and only then run the next.

┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ ASSESS │ ──► │ GATE │ ──► │ FIX │ ──► │ TEST │ └──────────┘ └──────────┘ └──────────┘ └──────────┘ read-only for a person sole writer; read-only for source; writes approves the edits only the source; re-runs assessment.md diagnosis or assessed files; the repro; sends it back writes fix.md writes test.md all three reports live in .specify/bugs/<slug>/

2.1 Assess — diagnose without touching anything

/speckit-bug-assess "<symptom, repro steps, expected behaviour>" slug=<slug>

Give it pasted text, a stack trace, or a GitHub issue URL. It searches the code for what the report mentions and writes assessment.md: a verdict and a severity, suspected code paths, a root-cause hypothesis with a stated confidence, a proposed remediation (preferred fix, alternatives, files likely to change, tests to add), risks, and open questions marked [NEEDS CLARIFICATION: …]. It is read-only for source, and it won’t invent reproduction steps or file paths. A fetched page is treated as data, never as instructions, and an unfamiliar host needs your yes before it’s fetched.

2.2 The gate — a person decides

Between assess and fix, you read the assessment and decide. Don’t continue unless the diagnosis is supported and the report really is a bug. Sending an assessment back is cheap. Approving a wrong one is expensive, because the fix stage treats it as a contract.

2.3 Fix — the only writer, on a short leash

/speckit-bug-fix slug=<slug>

This is the only stage that edits source code. It stays inside the files the assessment lists, keeps the change minimal, and adds no dependency the assessment didn’t call for. If new evidence forces it wider, it records that under Deviations from Assessment in fix.md instead of quietly growing the repair. It refuses to run on an invalid verdict. If the assessment turns out to be wrong, it stops editing and recommends re-running assess. It never edits assessment.md.

2.4 Test — evidence, never over-claimed

/speckit-bug-test slug=<slug>

It re-runs the reproduction and any tests the fix added, then the checks that prove nothing else broke (existing suites, lint, type-check), and writes test.md. It is read-only for source: it never edits code to make a check pass, and a check it can’t run is recorded as not-run, never as a pass. One rule matters above the rest: if a reproduction the assessment listed wasn’t actually exercised, the result cannot be verified. It drops to partial.

2.5 The verdict vocabularies

StageFieldValues
AssessVerdictvalid — reproducible or clearly grounded in the code · likely valid, needs reproduction — plausible but unverified · invalid — misuse, expected behaviour, duplicate or out of scope; fix refuses to run
AssessSeveritycritical · high · medium · low, with a rationale (user impact, blast radius, data risk). Judge the rationale, not the label.
FixStatusapplied · partial · not-applied — how much of the remediation landed
TestResultverified — symptom gone, critical checks pass: review the patch and evidence, then merge · partial — some verification missing or inconclusive: supply the evidence and test again · failed — symptom persists or something broke: revisit the diagnosis or fix, then test again. Each check reads pass, fail, skipped or not-run.

A partial fix is incomplete work. A partial test result is incomplete evidence — often about a perfectly good fix.

2.6 Slugs

The slug names the folder .specify/bugs/<slug>/ and is normalised to lowercase kebab-case. Leave it off and assess asks for one; fix and test reuse it from the same session or from the only candidate on disk. No report is overwritten without your confirmation. Pass slug= every time anyway — it removes a guess.

2.7 Why a green suite isn’t proof

A test suite checks what somebody thought to test, and the bug exists because nobody thought of this case. A suite that was green before the fix and is green after it says nothing about the reported symptom. The only evidence about the report itself is the original reproduction, re-run on the fixed code. That is why the test stage downgrades to partial without it, however many other checks passed.

2.8 Keep the three reports together

The assessment says what broke and why. fix.md says what changed and where it departed from the plan. test.md says what was run, and by whom. Commit all three with the fix and a reviewer can follow the repair from the reported symptom through to the evidence without asking you anything.

3.Where SPECTRA fits

The bug workflow is Spec Kit’s, and everything in Section 2 works the same in a plain Spec Kit project. SPECTRA installs the toolchain (spectra install) and adds agents on either side of the fix: one goes deeper than triage, the other delivers the result.

3.1 Triage or root cause?

Four commands in this course sound alike and answer different questions:

CommandQuestion it answersWrites
speckit.assess.* (Spec Kit assess)Is this idea worth building?.specify/assessments/<slug>/
speckit.bug.assess (Spec Kit bug)What is broken, and what is the smallest credible fix?.specify/bugs/<slug>/assessment.md
speckit.spectra.impact (SPECTRA)What would this change actually touch in the codebase?<artifact-root>/impact-analysis/
speckit.spectra.defect-rca (SPECTRA)Why did this defect happen, and what prevents it recurring?<artifact-root>/defect-rca/

speckit.bug.assess is triage: a diagnosis that exists to drive an immediate, bounded fix. speckit.spectra.defect-rca is a durable root-cause record:

speckit.bug.assessspeckit.spectra.defect-rca
DepthSuspected code paths and one hypothesis, with a confidenceAt least five hypothesis branches; each probe names its layer, from symptom down to systemic cause; ruled-out hypotheses stay in the record
HistoryLooks only at this reportChecks earlier analyses at intake and gives each of their preventive actions a verdict: completed or not completed (both need a citation), or undeterminable
ActionsA remediation for the fix stageCorrective and preventive actions in separate sections, each with an owner and a way to verify it
Changes code?No — fix doesNever. It fixes nothing and writes no tests.

Escalate to defect-rca when a defect recurs, when it’s serious, or when it reached production. Use both: fix the instance with the bug workflow, then write the RCA so the lesson outlives the ticket.

3.2 Delivering the fix

SPECTRA offers /speckit-spectra-create-pr through an after_implement hook. The bug extension registers no hooks, so nothing makes that offer after a bug fix. Once the fix is verified and you’ve reviewed it, run it yourself. It works on any branch, builds a bug branch’s PR body from its commits and diff, and asks before every commit, every push, and the PR itself.

4.Lab setup — a fresh copy (10 min)

★
Lab codebase — lab-codebase.zip
The same Next.js + TypeScript to-do app as Module 4. This lab needs it exactly as shipped, so start from a fresh copy.
Download

The app carries a real, latent defect. You’ll take it from a support ticket to a verified, committed fix, stopping at a human gate after each agent runs.

4.1 Unzip a new copy and take a baseline

Don’t reuse your Module 4 folder: its changes may hide or alter the defect. The zip ships no .gitignore, so add one before the baseline commit:

unzip lab-codebase.zip -d ~/sdd-bugfix-lab
cd ~/sdd-bugfix-lab/lab-codebase
npm install
printf 'node_modules/\n.next/\nnext-env.d.ts\n' > .gitignore
git init && git add -A && git commit -m "baseline"

4.2 Install the toolchain

spectra install
specify extension add bug
git add -A && git commit -m "Add Spec Kit, SPECTRA and the bug extension"

spectra install offers to run specify init --here --force; accept, and pick your usual agent. specify extension add bug may print “Configuration may be required”. The extension has no configuration, so ignore it. Committing the toolchain separately keeps your fix commit readable. Then restart your agent in this folder. You don’t need a constitution, a spec or a plan for this module.

5.The ticket, and your own reproduction (10 min)

5.1 Read the ticket

#212 — To-do app won’t load

“Pulled the to-do app this morning and all I get is ‘Application error: a client-side exception has occurred’. Every reload. I’ve restarted the dev server, reinstalled node_modules and pulled again — no change. A colleague on the same commit says it works for her. Only other thing I can think of: yesterday I tried a different to-do tutorial, also on localhost:3000.”

Follow-up: “Someone told me to look in DevTools under Local Storage. For localhost:3000 there’s one entry, todos, and its value is {"legacy":true}. No idea where that came from.”

A symptom, an environment detail, a clue — but no cause. Finding that is the assessment’s job.

5.2 Reproduce it yourself first

See it before any agent does. You learn what “fixed” must look like, a working reproduction earns a firmer verdict, and the same steps will prove the fix later.

npm run dev

Open http://localhost:3000 and add a task to confirm the app works. Then, in DevTools → Console, put your browser into the reporter’s state and reload:

localStorage.setItem("todos", '{"legacy":true}')

If the console blocks pasting, type allow pasting first, or edit the value under Application → Local Storage instead.

Next.js’s dev overlay reports a Runtime TypeError, todos.map is not a function, in app/components/TodoList.tsx. Behind it, the page reads “Application error: a client-side exception has occurred…”. The console shows Uncaught TypeError: todos.map is not a function at TodoList, just below the leftover Loaded todos: log. Reload, and you get the same again. The app never recovers.

Leave the dev server running and the bad value in place. To get a working app back at any time, run localStorage.removeItem("todos") and reload.

Hands off.You may already suspect where this lives. Don’t fix it yourself. The point is to see whether the assessment finds the cause, and to practise reviewing what it says.

6.Assess — and stand at the gate (12 min)

6.1 Run the assessment

Give the agent the symptom, your reproduction, and the expected behaviour:

/speckit-bug-assess
Symptom: the to-do app shows "Application error: a client-side exception
has occurred" on every load and never recovers. Console: Uncaught
TypeError: todos.map is not a function, at TodoList. The reporter had
been running a different to-do tutorial on localhost:3000.
Reproduction:
1. npm run dev, open http://localhost:3000
2. In the console: localStorage.setItem("todos", '{"legacy":true}')
3. Reload. The crash appears on this and every later reload.
Expected: the app loads instead of crashing, showing an empty list if it
cannot use what is stored.
slug=corrupt-storage

It writes .specify/bugs/corrupt-storage/assessment.md and suggests /speckit-bug-fix slug=corrupt-storage as the next step. Not yet. Run git status: nothing should have changed outside .specify/bugs/.

6.2 Review the assessment

Human gate: approve the diagnosis before any code changes. Check each part:

Send it back if the preferred remediation guards TodoList.tsx or wraps the list in an error boundary. That hides the crash but lets the bad value into the page’s state, and the next time someone adds a task, [newTodo, ...prev] in app/page.tsx throws. Fix where bad data enters, not where it explodes.

# Bug Assessment: Crash on load when stored todos are not a list

- **Slug**: corrupt-storage
- **Verdict**: valid
- **Severity**: medium

## Suspected Code Paths
- `lib/storage.ts:15` - `JSON.parse(raw) as Todo[]` returns any valid
  JSON as if it were a list; nothing checks its shape
- `app/components/TodoList.tsx:23` - `todos.map` is where it throws

## Root Cause Hypothesis
loadTodos() guards against unparseable JSON, but not against parseable
JSON of the wrong shape. Confidence: high.

## Proposed Remediation
**Preferred**: in loadTodos(), use the parsed value only if it is an
array, keeping entries with the Todo shape; otherwise return [].
**Alternatives**: a guard in TodoList or an error boundary - both hide
the crash and leave the bad value in page state.
**Files likely to change**: `lib/storage.ts`
**Tests to add or update**: loadTodos with a non-array, a malformed
entry and a valid list [NEEDS CLARIFICATION: no test runner exists]

## Risks & Considerations
- After the fallback, the save effect (app/page.tsx:19-21) writes []
  under "todos", replacing whatever was stored there

Yours will be worded differently, and your severity may differ. Judge it on whether each section is specific, cites real code, and matches what you saw in the browser.

6.3 The risk worth arguing about

Lines 19–21 of app/page.tsx save on every change, including the render straight after loading. So once loadTodos falls back to an empty list, the page writes [] over the other app’s {"legacy":true}. The fix quietly destroys another app’s data. If Risks doesn’t say so, raise it. Then decide:

Either answer is defensible. Not deciding isn’t. To send the assessment back, edit assessment.md yourself, or re-run /speckit-bug-assess with the constraint and the same slug and confirm the overwrite. Approve only when you’d be happy for the fix to do exactly what the assessment says, and nothing more.

7.Fix — bounded by the assessment (8 min)

7.1 Branch, then fix

git switch -c fix/corrupt-storage
/speckit-bug-fix slug=corrupt-storage

It starts by restating what it’s about to change; check that against what you approved. Then it edits, runs the project’s local checks, and writes fix.md.

7.2 Review the diff

Human gate: engineers review every change.

git status
git diff

The Todo type and saveTodos are untouched. A type guard goes above loadTodos, and loadTodos changes inside its existing try/catch:

function isTodo(value: unknown): value is Todo {
  if (typeof value !== "object" || value === null) return false;
  const item = value as Record<string, unknown>;
  return (
    typeof item.id === "string" &&
    typeof item.text === "string" &&
    typeof item.completed === "boolean" &&
    typeof item.createdAt === "string"
  );
}

export function loadTodos(): Todo[] {
  if (typeof window === "undefined") return [];
  try {
    const raw = window.localStorage.getItem("todos");
    if (!raw) return [];
    const parsed: unknown = JSON.parse(raw);
    return Array.isArray(parsed) ? parsed.filter(isTodo) : [];
  } catch {
    return [];
  }
}
  • Typing the parsed value unknown makes the compiler insist on a check before use — the opposite of as Todo[], which told it to trust whatever came out of storage.
  • The repeated "todos" string stays. Extracting a constant is a fair clean-up, but it isn’t this bug.
  • Filtering drops malformed entries, and the save effect then persists the list without them — the same kind of trade-off as 6.3. If your assessment didn’t mention it, it deserves a line under Risks.

If the diff strays outside the assessment without a recorded deviation, don’t approve it. git restore . resets the tracked files (the untracked reports stay); tighten the assessment if it left room for the overreach, and run fix again. It asks before overwriting fix.md.

8.Test — and close the gap yourself (10 min)

8.1 Run the verification

/speckit-bug-test slug=corrupt-storage

8.2 Expect partial — and read why

The most likely result is partial, and that’s the stage working. There’s no test script, so there’s no suite to run. A type-check (npx tsc --noEmit) or a build passes — but both passed before the fix too, because as Todo[] told the compiler to trust the stored data. That’s Section 2.7 in the flesh. The real proof is the reproduction, a browser step, and without browser tooling the agent can’t run it. So that row reads not-run or skipped, and the result can’t honestly be verified.

Got verified? Check how the reproduction was exercised. A Node snippet that calls loadTodos with a fake window tests a function, not the page the reporter saw. Whether that’s enough is your call.

8.3 Run the reproduction yourself

Human gate: QE decides what a finding means — and here, you are also the evidence. With the dev server still running:

  1. Set the bad value again (localStorage.setItem("todos", '{"legacy":true}')) and reload. The app shows its empty-list message: no overlay, no console error.
  2. Run localStorage.getItem("todos"). With the sample fix it returns "[]" — the overwrite from 6.3, now visible. Confirm it matches what you decided.
  3. Add two tasks, tick one, delete the other, reload. What’s left persists.
  4. Optional: store one good entry and one junk entry, reload, and note what your fix does. The sample fix keeps the good one.
    localStorage.setItem("todos", '[{"id":"1","text":"still here","completed":false,"createdAt":"2026-01-01T00:00:00.000Z"},{"oops":true}]')

8.4 Supply the evidence and test again

/speckit-bug-test slug=corrupt-storage
I re-ran the reproduction by hand in the browser on the fixed code:
after setting todos to {"legacy":true} and reloading, the app shows the
empty list with no error, and todos now holds []. Add, toggle, delete
and reload all behave as before.

Confirm when it asks to overwrite test.md. The new report should carry your evidence, attributed to you. Some agents will now say verified; others keep partial because they still didn’t run the reproduction themselves. Both are honest. A verified that nobody can trace to a reproduction someone actually ran is not.

9.Deliver (5 min)

9.1 Commit the fix with its paper trail

git add lib/storage.ts .specify/bugs/corrupt-storage
git commit -m "Fix crash when stored todos are not a list"

Add any other file your fix.md lists. With the reports in the same commit, a reviewer can follow symptom → diagnosis → change → evidence.

9.2 The pull request — in a real repository

Human gate: a maintainer merges. With a GitHub origin and an authenticated gh, you would now run:

/speckit-spectra-create-pr

It checks gh and the remote, works out the base branch, asks before pushing, drafts the body from your commits and diff, and asks once more before creating the PR. The lab has no remote, so skip it — or run it and watch it stop without changing anything.

10.Stretch: take it to root cause (20 min, optional)

Ask SPECTRA’s RCA agent the question the assessment didn’t: why did this happen, and what stops it recurring? Stay on fix/corrupt-storage:

/speckit-spectra-defect-rca The to-do app crashed on every load with
"Application error: a client-side exception has occurred" (console:
Uncaught TypeError: todos.map is not a function, at TodoList) for a
developer whose browser held {"legacy":true} under the localStorage key
"todos", left there by another tutorial app on localhost:3000. A fix is
in the latest commit on this branch.

It names the channel it resolved (a plain description), then reads the constitution (still the unfilled template) and the code. It records the commit it analysed and asks which one users were running: answer that they had the baseline, and the fix isn’t released. It searches docs/defect-rca/ for earlier analyses and records that it found none. Then it shows a hypothesis tree and asks at most five questions the repository can’t answer, each with a note on why it matters. Ask it to synthesise when you’re satisfied. It writes docs/defect-rca/001-<slug>.md and an index, and nothing else.

Compare its document with assessment.md:

That lesson is what SPECTRA’s Maintain gate is for. A team might make it a constitution rule — “data read from browser storage or the network is validated before use, never cast” — so every future plan and review inherits it. defect-rca never writes that line; the team decides, through /speckit-constitution. And the next storage crash would match this RCA on lib/storage.ts, prompting the question of whether its preventive actions were ever done.

11.Optional: the orchestrated path

Prefer to drive the stages from the terminal? Spec Kit’s first-party bugfix bundle pairs the bug extension with a bugfix workflow. It needs Spec Kit 1.0.9 or later and network access. Read the bundle before you install it:

specify bundle info bugfix
specify bundle install bugfix
specify workflow run bugfix \
  --input report="<text or issue URL>" --input slug="<slug>"

The run pauses after assess at a gate named review-assessment. Approve, and fix then test run; reject, and the run aborts. specify workflow status lists runs, and specify workflow resume <run_id> continues a paused one.

It guarantees the gate that matters most — no code changes until someone approves the diagnosis — but it’s the only gate. Fix and test run back to back, so you review the diff, fix.md and test.md together at the end. To try it, use a second fresh copy of the lab.

Knowledge check

Knowledge check7 questions
1. A ticket filed as a bug says: “I can’t set a reminder on a task — please fix.” The app has never had reminders. What’s the right route?
Answer: c
Nothing is broken; reminders never existed. New behaviour needs someone to decide what the product should do, which is a spec. Run through assess, this would most likely come back invalid, and the fix stage refuses to run on that verdict.
2. /speckit-bug-test ran a type-check and a build, and both pass. The assessment’s reproduction is a browser step the agent couldn’t run. What is the overall result?
Answer: b
A listed reproduction that wasn’t exercised caps the result at partial, however green everything else is. It isn’t failed, because nothing showed the symptom persisting. not-run describes a single check (the reproduction row), not the overall result.
3. Halfway through /speckit-bug-fix, the agent discovers the real cause is in a file the assessment never mentioned. What should it do?
Answer: c
When the assessment proves wrong, the fix stage stops editing, records what it found in fix.md, and sends you back to assess. It never edits assessment.md: that is the contract a person approved. A new diagnosis needs a new trip through the gate.
4. An assessment comes back valid with an open [NEEDS CLARIFICATION] about whether to add a test runner as part of the fix. When should that be settled?
Answer: a
With a valid verdict, the fix stage proceeds with the remediation as written; it only pauses on open questions when the verdict is likely valid, needs reproduction. The test stage records evidence and makes no scope decisions. The gate is where you settle it.
5. The same class of storage crash has now reached production twice. Which command adds what speckit.bug.assess doesn’t?
Answer: b
Recurrence, severity and production are the cues to escalate. defect-rca checks earlier analyses and whether their preventive actions were done, digs past the first plausible code path, and records preventive actions separately from the fix. impact maps what a proposed change would touch; assess.intake starts an idea assessment.
6. Your test report says verified and you’ve reviewed it. Why hasn’t SPECTRA offered to open a pull request?
Answer: a
The automatic offer fires after /speckit-implement, which a bug fix never runs. create-pr works on any branch — on a bug branch it drafts the body from your commits and diff — and nothing else has to run first.
7. In the lab, the fixed app writes [] over whatever was stored under todos. Make the case for accepting that, the case against, and say where the decision belongs.
For: the app owns that key on its origin, the stored value was already unusable, the collision only happens on developer machines, and it’s the smallest fix. Against: data loss can’t be undone; the same fallback would silently wipe real users’ lists the day a release changes the stored shape; and a safer remediation (don’t save until the user changes something, or keep a copy first) is cheap. Where it belongs: in the assessment, before the fix runs — under Risks if accepted, in the remediation and file list if rejected. Then the fix is bounded by a decision a person made at the gate, and a reviewer can see it was deliberate rather than finding it as a surprise deviation.

What's next

You’ve now run Spec Kit’s workflows as they ship: the SDD loop, idea assessment, and bug fixing. Module 7, Customization, shows how to make them your own with extensions, presets, overrides, workflows and bundles — including why restyling the bug reports you just read takes a command override or a preset rather than a template edit.