Learning objectives
By the end of this module, you will be able to:
- Tell discovery (“should we build it?”) from delivery (“how do we build it?”), and place assessment before Plan in SPECTRA’s agentic SDLC.
- Run the five assessment stages one at a time, reviewing each artifact at its human gate.
- Hold research to the evidence rules: evidence against,
ASSUMPTIONtags, no invented citations. - Read the scorecard, say what a
gorequires, and explain what follows each verdict. - Resolve a clarification by refining an artifact instead of re-running a stage.
- Ground an assessment with
speckit.spectra.impact, and placespeckit.spectra.brdafter ago. - Hand a
goto/speckit-specifyyourself.
1.Discovery before delivery
Everything so far in this course answers one question: how do we build it? In Modules 3 and 4, somebody had already decided the feature was worth building, and you took that on trust. That decision is where a lot of wasted engineering starts. An idea sounds reasonable in a meeting, becomes a ticket, then a spec, and three weeks later the team learns nobody needed it. SDD makes delivery disciplined. It does nothing to stop you delivering the wrong thing well.
Idea assessment is the step in front: should we build it at all? Spec Kit ships it as the assess extension, and keeps the two questions on separate tracks:
| Discovery | Delivery | |
|---|---|---|
| Question | Should we build it? | How do we build it? |
| Spec Kit tooling | The assess extension: intake → research → define → shape → decide | The SDD loop: constitution → specify → … → implement → converge |
| Writes | .specify/assessments/<slug>/ — never source code | specs/<NNN-feature>/, then code |
| Ends with | A verdict and its reasoning: go, needs-clarification, or kill | A reviewed, merged change |
1.1 Where assessment sits in the agentic SDLC
SPECTRA runs the agentic SDLC in seven phases, from 00 Foundation to 06 Maintain, each ending at a human gate. Idea assessment sits in front of 01 Plan. Plan’s gate is the product owner approving a requirement’s intent and scope; assessment asks the earlier question of whether there should be a requirement at all. A kill here means Plan never starts.
It needs no source code: it works in an empty initialized project, even for decisions that aren’t about software, and for an existing repository you point intake at the code. And it needs no SDD first — no constitution, no spec. On a brownfield repository it can be the first thing a team does after initializing Spec Kit.
The usual rule holds: agents draft, people decide. Every stage is a draft for you to review, and the verdict is a recommendation.
2.Five stages, five artifacts
The extension adds five commands, one per stage, run one at a time with a shared slug — a kebab-case name such as sync-across-devices that becomes the idea’s directory. The review between stages is the human gate: every later stage builds on the artifact you just approved.
| Stage | Command | Writes | What it does |
|---|---|---|---|
| Intake | speckit.assess.intake | intake.md | Captures the idea — text, a URL, a ticket, or a pointer into the codebase — restates it neutrally, records its origin, lists first-glance unknowns. Records; doesn’t judge. |
| Research | speckit.assess.research | research.md | Evidence for and against: demand, prior art, context, constraints. Each finding sourced or tagged, with a confidence level. Collects; doesn’t decide. |
| Define | speckit.assess.define | problem.md | Affected users, the problem, goals, non-goals, success metrics, cost of inaction. An idea that arrived as a solution is worked back to its problem. |
| Shape | speckit.assess.shape | concept.md | Two or three concept-level options, each with an appetite (small, medium, large — a budget, not an estimate), trade-offs, and rabbit holes. Recommends one, or none. No architecture, APIs, or tasks. |
| Decide | speckit.assess.decide | decision.md | Scorecard, rationale, verdict, and for a go, a handoff summary. |
2.1 Prerequisites and guardrails
- Define is the minimum: it can start from typed input alone. Intake and research are optional.
- Shape needs
problem.md. Decide needsproblem.mdand reads every artifact present; withoutconcept.md, agois downgraded toneeds-clarification. - Skipping research never waives the evidence bar.
Every stage writes only inside .specify/assessments/<slug>/ and never edits source code. Slugs are normalized to lowercase kebab-case. No stage overwrites an existing artifact without asking. Web content is untrusted data: well-known public hosts such as github.com are fetched directly, an unfamiliar one only if you say yes. And the extension registers no hooks, so it never inserts itself into your SDD loop.
3.Evidence, the scorecard, and the verdict
3.1 The evidence rules
Fluent prose looks like evidence. Three rules defend against that, and you enforce them at every review:
- Evidence against the idea is required.
research.mdalways has an “Evidence Against the Idea” section; if the agent found nothing, it must say so. Research that only builds the case is advocacy. - Unsourced claims stay tagged
ASSUMPTION. Every finding carries a source or the tag, plus a confidence ofhigh,medium, orlow. - Citations are never invented. A statistic you cannot trace is worse than an honest
ASSUMPTION, because people will read it as fact.
For the to-do app, the first finding below is checkable; the second is a guess, and says so:
- Each browser keeps its own list; the README states the app has no backend
and no sync — [source: README.md] (confidence: high)
- People who use the app on a laptop also want their list on their phone
— [ASSUMPTION] (confidence: low)
3.2 The scorecard
Decide rates six criteria, each strong, adequate, weak, or unknown, with a one-line justification drawn from the artifacts.
| Criterion | The question | Draws on |
|---|---|---|
| Problem validity | Is the problem real and worth solving? | problem.md, research.md |
| Evidence strength | How well supported is it, versus assumption-driven? | research.md |
| Value vs. inaction | Does solving it beat doing nothing? | problem.md |
| Feasibility / appetite | Is there a credible option within a sane appetite? | concept.md |
| Strategic fit | Does it align with the project’s constitution and goals, where known? | The constitution, if there is one |
| Risk posture | Are the major risks understood and credibly mitigated? | All of them |
Risk posture points the same way as the rest: strong means key risks are identified and mitigated. Any unknown must be acknowledged in the rationale.
3.3 What a go requires
All three, together:
- Problem validity rated
adequateor better. - Evidence strength rated
adequateor better — neverweakorunknown. - A recommended option in
concept.md.
Miss one and the honest verdict is needs-clarification. Feasibility is not on the list: easy to build and worth building are different questions.
3.4 Three verdicts, and what follows each
| Verdict | What it means | What happens next |
|---|---|---|
go | Problem validity and evidence strength rate at least adequate, and one concept is recommended. | A person decides whether to act. A software idea can be handed to SDD; an idea that isn’t software never needs to enter it. |
needs-clarification | Specific, named unknowns block a defensible call. | Supply the missing facts in the affected artifacts, then revise decision.md. |
kill | Not worth pursuing now. The decisive reason is stated plainly. | Keep the record. Stopping is a successful outcome. |
Take the last row seriously: a process that never kills anything is a rubber stamp. A written-down kill saves the build cost, and if the facts change later, the record shows which assumption to revisit.
3.5 Refine, don’t re-run
Each stage normally runs once. A [NEEDS CLARIFICATION: …] marker is a gap to fill, not a reason to regenerate: edit the Markdown yourself, or ask the agent to update a named artifact, keeping source and confidence tags. Then have decision.md revised if the blocker clears. Re-running is for exceptions such as a wrong slug, and never overwrites without asking. One trap: decide’s report names a stage to revisit — read that as which artifact to refine.
3.6 The handoff is manual
No hook connects assessment to SDD. A go writes a handoff summary into decision.md — problem, chosen approach, scope, success metrics, open questions — and if you decide to build, you pass it to /speckit-specify yourself. Deciding to build is a human call, and it should look like one.
4.Where SPECTRA fits
The assess extension is Spec Kit’s, added with specify extension add assess, and it works the same in a plain Spec Kit project. SPECTRA adds agents on either side of it. First, untangle four commands that all sound like “assess”:
| Command | Question it answers | Writes |
|---|---|---|
speckit.assess.* (Spec Kit assess) | Is this idea worth building? | .specify/assessments/<slug>/ |
speckit.bug.assess (Spec Kit bug) | What is broken, and what is the smallest credible fix? | .specify/bugs/<slug>/assessment.md |
speckit.spectra.impact (SPECTRA) | What would this change actually touch in the codebase? | <artifact-root>/impact-analysis/ |
speckit.spectra.defect-rca (SPECTRA) | Why did this defect happen, and what prevents it recurring? | <artifact-root>/defect-rca/ |
This module uses the first and third; Module 6 uses the second. <artifact-root> is docs/ unless your constitution declares an Artifact root: line.
4.1 Impact: what the idea would actually touch
Research looks outward, at users and prior art. speckit.spectra.impact looks inward: give it one paragraph describing what should be true after the change, and it reads the project, asks at most five questions the code cannot answer, and writes a numbered analysis where every finding carries a path:line citation and a confidence level. It states how much it read and what it searched for without finding, derives its rating from fixed triggers, never claims there is no impact, and makes no network requests.
That evidence feeds research (cited constraints and evidence against, instead of assumptions about the code), shape (realistic appetites and rabbit holes), and decide (grounded feasibility and risk ratings). Know its limit: impact can show an idea is expensive, never that anyone wants it. Code carries no evidence of demand.
4.2 BRD: structure a go before specify
After a go, you can hand the summary straight to /speckit-specify. When the requirement needs more structure first — prioritized user journeys and testable business requirements for a product owner to approve in Plan — run /speckit-spectra-brd on it. It writes a specify-ready BRD under <artifact-root>/brd/, asks questions only where there are gaps, never invents scope, and tells you to run specify rather than running it.
Several of these documents could each hold the same success metrics. Give each a single job:
| Artifact | Its job | Success metrics |
|---|---|---|
problem.md, decision.md | Why: problem, evidence, verdict | First written here, as the decision’s yardstick. After handoff, a record — don’t edit it to match later changes. |
| Impact analysis | What it would touch; risk; rollback | None. Evidence, not requirements. |
| BRD (optional) | Journeys, requirements, scope | Carried over from the handoff summary, not rewritten; can link back to decision.md. |
spec.md | What to build | Success Criteria. From here on, metrics change only here. |
Metrics flow one way — problem.md, handoff, BRD, spec — and change only in the newest artifact in that chain.
Lab — Should the to-do app sync across devices?
lab-codebase.zipYou will assess two ideas. “Sync to-dos across devices” sounds like something every to-do app should have, and runs straight into what this codebase is. “Show only open tasks”, the stretch, is small enough to reach a go. Nothing in this lab edits source code or runs the app. Prompts use the Claude Code / Copilot spelling; Module 1 lists the others.
5.Set up (5 min)
5.1 Open the project
Open your Module 4 project, where SPECTRA is already installed, and commit any leftover work so this lab’s files show up cleanly in git status. If you skipped Module 4, start fresh:
unzip lab-codebase.zip -d ~/sdd-assess-lab
cd ~/sdd-assess-lab/lab-codebase
git init && git add -A && git commit -m "baseline"
spectra install
Accept the offer to initialize Spec Kit, pick your agent, and restart it. No npm install needed.
5.2 Add the assess extension
specify extension add assess
The install ends with ⚠ Configuration may be required, pointing at .specify/extensions/assess/. Ignore it; the extension ships no configuration file. Then restart your agent so it loads the five new commands.
specify extension listshowsassess, and typing/speckit-assessin your agent offers all five stages.- If the add fails, run
spectra version: the course needs Spec Kit 1.0.9 or later, andspectra updatebrings it current.
6.Intake: capture the idea (5 min)
/speckit-assess-intake "Sync to-dos across devices, so a task I add on my laptop is there on my phone." This is an idea for this repository; read enough of the codebase to record what it relates to. slug=sync-across-devices
The agent writes .specify/assessments/sync-across-devices/intake.md. If it asks for a slug, you left off slug=.
Human gate: you confirm the idea was recorded faithfully, before anyone evaluates it. Check intake.md:
- It quotes your words, and the restatement is neutral — neither “a great addition” nor “infeasible”.
- “Raised by” and “Trigger” are
[NEEDS CLARIFICATION: …]unless you supplied them. An invented trigger (“users have been asking”) is fabrication; correct it. - Unknowns are listed but not answered — which devices, whose account, two devices editing one task — and there is no solution or verdict.
7.Research: evidence for and against (7 min)
/speckit-assess-research slug=sync-across-devices
The agent may search the web. If it asks to fetch a host it doesn’t recognize, the default is no; say yes only for a site you know.
Human gate: you decide whether this research actually challenges the idea. Check research.md:
- Every finding carries
[source: …]or[ASSUMPTION], plus a confidence level. - “Evidence Against the Idea” has substance. The README alone supplies some: no backend, no auth, no sync.
- “Users & Demand” is honest. A course app has no real users, so expect
ASSUMPTIONand low confidence. An invented survey result is a failure; an admitted gap is not. - Spot-check one listed source: the claim is really there.
If something fails, ask the agent to fix that part of research.md. Don’t re-run the stage.
8.Ground it in the code with impact (10 min)
8.1 Run the impact analysis
/speckit-spectra-impact A person who uses the to-do app on more than one device sees the same tasks on each: a task added, completed, or deleted on one device appears on the others without any manual export or import. Today each browser keeps its own separate list.
When it asks whether this repository is the only system involved, answer 1 — just this repository; the document records that you said so. It may then ask up to five questions, one at a time, each with a recommendation. Answer as the product owner, or press Enter to accept it (recorded as defaulted — not confirmed). It writes a numbered analysis and a folder index under docs/impact-analysis/, or your declared artifact root.
Human gate: the analysis is always a draft; you decide what it means. Check that:
- Findings cite
path:linein files you recognize, such as the localStorage helpers inlib/storage.ts. - It names what it searched for and did not find — server routes, network calls, authentication. That evidenced absence proves there is no backend to build on.
- The rating names its trigger, coverage is stated, and effort is a coupling heuristic, not days.
- If the security and privacy lens fired, it routes to a named agent rather than analysing. That agent may not exist yet; a routed item is a handoff, not a promise.
8.2 Fold the findings into research
Decide weighs the artifacts in the assessment directory, so evidence that lives only in docs/ won’t reach it. Bring it in:
Add the relevant findings from docs/impact-analysis/<NNN>-<name>.md to
.specify/assessments/sync-across-devices/research.md, under "Data & Constraints"
and "Evidence Against the Idea". Keep each path:line citation as the source,
list the analysis under Sources, and use research.md's own high/medium/low
confidence scale. Add nothing the analysis does not say, and keep every
existing source and confidence tag.
- The new findings carry
path:linesources, and the demand findings are unchanged — code is no evidence anyone wants sync.
9.Define, then shape (8 min)
9.1 Define the problem
/speckit-assess-define slug=sync-across-devices
Human gate: you confirm the problem is stated as a problem, not a feature. Check problem.md:
- It doesn’t say “build sync”. It says who is stuck: someone using more than one device can’t see the same tasks on each.
- Users that research didn’t support are marked
[NEEDS CLARIFICATION: …], and non-goals bound the idea (sharing lists with other people, real-time collaboration). - Success metrics are measurable, even with an “unknown” baseline, and the cost of inaction is honest.
9.2 Shape the options — and talk about appetite
/speckit-assess-shape slug=sync-across-devices
Here the idea meets the codebase. The app has no backend, so real sync means building one: somewhere to store tasks, a way to know whose they are, and a rule for when two devices change the same task. Before reading concept.md, decide: how much is this worth to a 200-line, single-browser app? A day? A month? That is the appetite — a budget you set, not an estimate you are handed.
Human gate: check that shaping stayed at concept level and took appetite seriously:
- Two or three genuinely different options, including a smallest-thing-that-could-work and, where sensible, do nothing or buy instead of build — each with appetite, trade-offs, and rabbit holes.
- No endpoints, schemas, or task lists. If they crept in, ask the agent to remove them.
- The recommendation ties back to
problem.md’s goals — or recommends none, which is valid.
- Export and import a file (small). No backend. But it isn’t really sync: it is manual, and an old file can overwrite newer tasks.
- A hosted sync service (medium). Less to build, but accounts, a vendor, ongoing cost, and task data leaving the device.
- Our own backend with accounts (large). Rabbit holes: conflicting edits, identity, hosting, privacy.
- Do nothing. Each browser keeps its own list, as the README already says.
10.Decide (5 min)
/speckit-assess-decide slug=sync-across-devices
Expect needs-clarification or kill. Read the scorecard like this:
- Problem validity and Evidence strength first — they gate a
go. With demand resting on assumptions, Evidence strength shouldn’t be aboveweak. - Every justification points at an artifact. “Feasibility: weak” should cite the missing backend, not a feeling. Every
unknownis acknowledged. - The verdict follows the ratings. Weak evidence plus
gois an over-claim; push back.
For this idea, needs-clarification means named unknowns block the call — typically real demand, and the appetite for a backend. kill means cost beats value: a backend, identity, and hosting, for demand nobody has shown (and if your Module 4 constitution kept “no backend”, Strategic fit agrees). A go is a red flag; check Evidence strength.
Human gate: you are the product owner. The verdict is a recommendation; accepting it, or overriding it with a reason, is your call. Commit now so the next step’s changes are easy to review (adjust the second path if you declared another artifact root):
git add .specify/assessments docs/impact-analysis
git commit -m "Assess sync across devices"
A sample. Your ratings will differ; what matters is that each is justified from the artifacts and the verdict follows the rules.
| Criterion | Rating | Justification |
|---|---|---|
| Problem validity | weak | Plausible, but no user signal for this app; demand is ASSUMPTION. |
| Evidence strength | weak | Codebase findings cited from the impact analysis; demand uncited. Research confidence: low. |
| Value vs. inaction | unknown | Each browser keeps its own list; no data on how many people hit that. |
| Feasibility / appetite | weak | Beyond manual export, every option needs storage, identity, and a server the app lacks (impact: localStorage only, no API routes, no auth). |
| Strategic fit | weak | The Module 4 constitution declares localStorage-only persistence, no backend. |
| Risk posture | weak | Conflicting edits, task data leaving the device, running a service — named, none mitigated. |
Verdict: needs-clarification. Blocking: [NEEDS CLARIFICATION: Do people use this app on more than one device?] and [NEEDS CLARIFICATION: What appetite will we accept for building and running a backend?] Revisit: research, shape.
A kill is equally defensible (“cost exceeds value: a backend, accounts, and hosting for demand nobody has evidenced”). The indefensible answer is go.
11.Refine instead of re-running (5 min)
Pick one [NEEDS CLARIFICATION: …] marker you can answer and resolve it. Appetite is the easiest, and it lives in concept.md. For this exercise you are the product owner, so a fact you supply is sourced to you:
As product owner, I have confirmed the appetite for this idea: at most one week
of work, and no hosted service we would have to run or pay for. Record that in
.specify/assessments/sync-across-devices/concept.md, sourced to me with today's
date, and adjust any wording there that depended on the old appetite. Then check
whether this resolves a blocker in decision.md; if it does, revise the scorecard,
rationale, and verdict in place. Keep every source and confidence tag, and do not
regenerate either file.
Or edit concept.md yourself — replace the marker with the fact and [source: product owner, <date>] — then ask the agent to revise decision.md.
With that appetite no real sync option fits, so expect the verdict to harden, often to kill. Resolving an unknown doesn’t always move an idea forward; sometimes it closes the question. That is the funnel doing its job.
Human gate: you supplied the fact; the agent must not invent anything around it. Run git diff:
- A few targeted lines changed in
concept.mdanddecision.md— not two rewritten files. - The marker is gone, the fact carries its source, and the rationale says what changed and why the verdict did or didn’t move.
- Nobody re-ran shape or decide. If the agent offers to, it must ask before overwriting — say no.
12.Stretch: take a small idea to go, and hand it off (10 min)
Now a proportionate idea. Today finished tasks stay in the list, struck through. Run each stage separately and review each artifact, as before:
/speckit-assess-intake "Let me hide completed tasks so I only see what is still open." This is an idea for this repository. slug=show-open-tasks
/speckit-assess-research slug=show-open-tasks Include prior art from other to-do apps, such as the TodoMVC project on github.com.
/speckit-assess-define slug=show-open-tasks
/speckit-assess-shape slug=show-open-tasks
/speckit-assess-decide slug=show-open-tasks
This should score very differently: small appetite, nothing leaving the browser, a comfortable fit with how the app works, plenty of prior art. A go is likely. If you get needs-clarification, find the criterion that blocked it — usually Evidence strength — and refine that artifact.
Human gate: a go says the idea is worth specifying, not that you must build it. If you go ahead, the handoff is yours:
/speckit-specify Use the handoff summary in .specify/assessments/show-open-tasks/decision.md to specify a way to show only the open tasks in the to-do list.
Or structure it as a business requirement first: run the BRD agent on the decision, then give the BRD it writes to /speckit-specify. Pick one route, so the metrics live once in the chain.
/speckit-spectra-brd .specify/assessments/show-open-tasks/decision.md
- The new
spec.mdtakes its success criteria from the handoff instead of inventing new ones. - The out-of-scope items survived, and the carried-forward questions are open questions, ready for
/speckit-clarify.
Stop there; don’t run /speckit-plan. The point is the moment discovery hands over to delivery.
13.Optional: run the assessment as a workflow
Spec Kit also ships the process as a first-party bundle: the assess extension plus an assess workflow. It needs network access and Spec Kit 1.0.9 or later, and sits outside the 75 minutes. Read what a bundle installs before installing it:
specify bundle info assess
specify bundle install assess
specify workflow run assess --input idea="Let me reorder tasks by dragging them" --input slug="reorder-tasks"
It runs the five stages back to back, then stops at a final review-verdict gate: approve completes the assessment, reject aborts it. A go is still handed off manually. You trade the per-stage reviews for one review of all five artifacts at the end, so a bad intake is caught late. Use a fresh slug: stages won’t overwrite existing artifacts in an unattended run. The workflow suits triaging a queue of ideas the same way every time; the stage-by-stage path suits contested ideas, and learning.
Knowledge check
strong, Evidence strength weak, and a concept is recommended. What is the most positive verdict it can honestly record?go needs Problem validity and Evidence strength at adequate or better, plus a recommended concept. Weak evidence caps the verdict at needs-clarification however strong the rest is. kill is possible, but not the most positive.research.md lists six findings, all in favour, and has no “Evidence Against the Idea” section. What do you do?needs-clarification because concept.md has an appetite marker. The product owner now says: one week, no hosted service. What next?/speckit-spectra-impact add that the assess stages don’t produce themselves?show-open-tasks reaches go. What happens next?review-verdict gate. Crossing into delivery is a deliberate human step.kill. A teammate says the hour was wasted, because nothing got built. What do you tell them?kill is the deliverable. For an hour of reading, we learned that sync means building and running a backend, accounts, and conflict handling the app doesn’t have (cited in the impact analysis), for demand nobody has evidenced (every demand claim is an ASSUMPTION), against a constitution that says localStorage only. Otherwise we would have found out weeks into a plan. And because decision.md records the decisive reason, whoever proposes sync again — say, once real users ask — can reopen it with new evidence against a named assumption instead of restarting the argument.What's next
This module asked whether something new should be built. Module 6, Bug Fix, turns to something already built that is broken. You will take a real defect in a fresh copy of the lab app through Spec Kit’s bug workflow — assess, fix, verify — with a diagnosis before any code changes, a fix that stays inside what the diagnosis covers, and verification against the original reproduction.