Module 5

Idea Assessment — Should We Build It?

Not every idea deserves a spec. Take two ideas for the to-do app through Spec Kit’s five-stage idea assessment, back them with SPECTRA’s cited codebase evidence, and reach a verdict you can defend — including the one that says don’t build it.

Time 75 min
Format Concepts + hands-on lab
Prerequisites Modules 1–4; SPECTRA installed

Learning objectives

By the end of this module, you will be able to:

  1. Tell discovery (“should we build it?”) from delivery (“how do we build it?”), and place assessment before Plan in SPECTRA’s agentic SDLC.
  2. Run the five assessment stages one at a time, reviewing each artifact at its human gate.
  3. Hold research to the evidence rules: evidence against, ASSUMPTION tags, no invented citations.
  4. Read the scorecard, say what a go requires, and explain what follows each verdict.
  5. Resolve a clarification by refining an artifact instead of re-running a stage.
  6. Ground an assessment with speckit.spectra.impact, and place speckit.spectra.brd after a go.
  7. Hand a go to /speckit-specify yourself.

1.Discovery before delivery

Everything so far in this course answers one question: how do we build it? In Modules 3 and 4, somebody had already decided the feature was worth building, and you took that on trust. That decision is where a lot of wasted engineering starts. An idea sounds reasonable in a meeting, becomes a ticket, then a spec, and three weeks later the team learns nobody needed it. SDD makes delivery disciplined. It does nothing to stop you delivering the wrong thing well.

Idea assessment is the step in front: should we build it at all? Spec Kit ships it as the assess extension, and keeps the two questions on separate tracks:

DiscoveryDelivery
QuestionShould we build it?How do we build it?
Spec Kit toolingThe assess extension: intake → research → define → shape → decideThe SDD loop: constitution → specify → … → implement → converge
Writes.specify/assessments/<slug>/ — never source codespecs/<NNN-feature>/, then code
Ends withA verdict and its reasoning: go, needs-clarification, or killA reviewed, merged change

1.1 Where assessment sits in the agentic SDLC

SPECTRA runs the agentic SDLC in seven phases, from 00 Foundation to 06 Maintain, each ending at a human gate. Idea assessment sits in front of 01 Plan. Plan’s gate is the product owner approving a requirement’s intent and scope; assessment asks the earlier question of whether there should be a requirement at all. A kill here means Plan never starts.

It needs no source code: it works in an empty initialized project, even for decisions that aren’t about software, and for an existing repository you point intake at the code. And it needs no SDD first — no constitution, no spec. On a brownfield repository it can be the first thing a team does after initializing Spec Kit.

The usual rule holds: agents draft, people decide. Every stage is a draft for you to review, and the verdict is a recommendation.

2.Five stages, five artifacts

The extension adds five commands, one per stage, run one at a time with a shared slug — a kebab-case name such as sync-across-devices that becomes the idea’s directory. The review between stages is the human gate: every later stage builds on the artifact you just approved.

Discovery — every file is written under .specify/assessments/<slug>/ intake research define shape decide intake.md research.md problem.md concept.md decision.md go needs-clarification kill a person decides; hand off manually to /speckit-specify refine the named artifact, then revise decision.md stop, and keep the reasoning — a successful outcome
Five stages, one artifact each, one of three verdicts. You review every artifact before running the next stage.
StageCommandWritesWhat it does
Intakespeckit.assess.intakeintake.mdCaptures the idea — text, a URL, a ticket, or a pointer into the codebase — restates it neutrally, records its origin, lists first-glance unknowns. Records; doesn’t judge.
Researchspeckit.assess.researchresearch.mdEvidence for and against: demand, prior art, context, constraints. Each finding sourced or tagged, with a confidence level. Collects; doesn’t decide.
Definespeckit.assess.defineproblem.mdAffected users, the problem, goals, non-goals, success metrics, cost of inaction. An idea that arrived as a solution is worked back to its problem.
Shapespeckit.assess.shapeconcept.mdTwo or three concept-level options, each with an appetite (small, medium, large — a budget, not an estimate), trade-offs, and rabbit holes. Recommends one, or none. No architecture, APIs, or tasks.
Decidespeckit.assess.decidedecision.mdScorecard, rationale, verdict, and for a go, a handoff summary.

2.1 Prerequisites and guardrails

Every stage writes only inside .specify/assessments/<slug>/ and never edits source code. Slugs are normalized to lowercase kebab-case. No stage overwrites an existing artifact without asking. Web content is untrusted data: well-known public hosts such as github.com are fetched directly, an unfamiliar one only if you say yes. And the extension registers no hooks, so it never inserts itself into your SDD loop.

3.Evidence, the scorecard, and the verdict

3.1 The evidence rules

Fluent prose looks like evidence. Three rules defend against that, and you enforce them at every review:

  1. Evidence against the idea is required. research.md always has an “Evidence Against the Idea” section; if the agent found nothing, it must say so. Research that only builds the case is advocacy.
  2. Unsourced claims stay tagged ASSUMPTION. Every finding carries a source or the tag, plus a confidence of high, medium, or low.
  3. Citations are never invented. A statistic you cannot trace is worse than an honest ASSUMPTION, because people will read it as fact.

For the to-do app, the first finding below is checkable; the second is a guess, and says so:

- Each browser keeps its own list; the README states the app has no backend
  and no sync — [source: README.md] (confidence: high)
- People who use the app on a laptop also want their list on their phone
  — [ASSUMPTION] (confidence: low)

3.2 The scorecard

Decide rates six criteria, each strong, adequate, weak, or unknown, with a one-line justification drawn from the artifacts.

CriterionThe questionDraws on
Problem validityIs the problem real and worth solving?problem.md, research.md
Evidence strengthHow well supported is it, versus assumption-driven?research.md
Value vs. inactionDoes solving it beat doing nothing?problem.md
Feasibility / appetiteIs there a credible option within a sane appetite?concept.md
Strategic fitDoes it align with the project’s constitution and goals, where known?The constitution, if there is one
Risk postureAre the major risks understood and credibly mitigated?All of them

Risk posture points the same way as the rest: strong means key risks are identified and mitigated. Any unknown must be acknowledged in the rationale.

3.3 What a go requires

All three, together:

Miss one and the honest verdict is needs-clarification. Feasibility is not on the list: easy to build and worth building are different questions.

3.4 Three verdicts, and what follows each

VerdictWhat it meansWhat happens next
goProblem validity and evidence strength rate at least adequate, and one concept is recommended.A person decides whether to act. A software idea can be handed to SDD; an idea that isn’t software never needs to enter it.
needs-clarificationSpecific, named unknowns block a defensible call.Supply the missing facts in the affected artifacts, then revise decision.md.
killNot worth pursuing now. The decisive reason is stated plainly.Keep the record. Stopping is a successful outcome.

Take the last row seriously: a process that never kills anything is a rubber stamp. A written-down kill saves the build cost, and if the facts change later, the record shows which assumption to revisit.

3.5 Refine, don’t re-run

Each stage normally runs once. A [NEEDS CLARIFICATION: …] marker is a gap to fill, not a reason to regenerate: edit the Markdown yourself, or ask the agent to update a named artifact, keeping source and confidence tags. Then have decision.md revised if the blocker clears. Re-running is for exceptions such as a wrong slug, and never overwrites without asking. One trap: decide’s report names a stage to revisit — read that as which artifact to refine.

3.6 The handoff is manual

No hook connects assessment to SDD. A go writes a handoff summary into decision.md — problem, chosen approach, scope, success metrics, open questions — and if you decide to build, you pass it to /speckit-specify yourself. Deciding to build is a human call, and it should look like one.

4.Where SPECTRA fits

The assess extension is Spec Kit’s, added with specify extension add assess, and it works the same in a plain Spec Kit project. SPECTRA adds agents on either side of it. First, untangle four commands that all sound like “assess”:

CommandQuestion it answersWrites
speckit.assess.* (Spec Kit assess)Is this idea worth building?.specify/assessments/<slug>/
speckit.bug.assess (Spec Kit bug)What is broken, and what is the smallest credible fix?.specify/bugs/<slug>/assessment.md
speckit.spectra.impact (SPECTRA)What would this change actually touch in the codebase?<artifact-root>/impact-analysis/
speckit.spectra.defect-rca (SPECTRA)Why did this defect happen, and what prevents it recurring?<artifact-root>/defect-rca/

This module uses the first and third; Module 6 uses the second. <artifact-root> is docs/ unless your constitution declares an Artifact root: line.

4.1 Impact: what the idea would actually touch

Research looks outward, at users and prior art. speckit.spectra.impact looks inward: give it one paragraph describing what should be true after the change, and it reads the project, asks at most five questions the code cannot answer, and writes a numbered analysis where every finding carries a path:line citation and a confidence level. It states how much it read and what it searched for without finding, derives its rating from fixed triggers, never claims there is no impact, and makes no network requests.

That evidence feeds research (cited constraints and evidence against, instead of assumptions about the code), shape (realistic appetites and rabbit holes), and decide (grounded feasibility and risk ratings). Know its limit: impact can show an idea is expensive, never that anyone wants it. Code carries no evidence of demand.

4.2 BRD: structure a go before specify

After a go, you can hand the summary straight to /speckit-specify. When the requirement needs more structure first — prioritized user journeys and testable business requirements for a product owner to approve in Plan — run /speckit-spectra-brd on it. It writes a specify-ready BRD under <artifact-root>/brd/, asks questions only where there are gaps, never invents scope, and tells you to run specify rather than running it.

Several of these documents could each hold the same success metrics. Give each a single job:

ArtifactIts jobSuccess metrics
problem.md, decision.mdWhy: problem, evidence, verdictFirst written here, as the decision’s yardstick. After handoff, a record — don’t edit it to match later changes.
Impact analysisWhat it would touch; risk; rollbackNone. Evidence, not requirements.
BRD (optional)Journeys, requirements, scopeCarried over from the handoff summary, not rewritten; can link back to decision.md.
spec.mdWhat to buildSuccess Criteria. From here on, metrics change only here.

Metrics flow one way — problem.md, handoff, BRD, spec — and change only in the newest artifact in that chain.

Lab — Should the to-do app sync across devices?

★
Lab codebase — lab-codebase.zip
The same Next.js to-do app as Module 4. Use your Module 4 project if you have it; download only if you are starting fresh.
Download

You will assess two ideas. “Sync to-dos across devices” sounds like something every to-do app should have, and runs straight into what this codebase is. “Show only open tasks”, the stretch, is small enough to reach a go. Nothing in this lab edits source code or runs the app. Prompts use the Claude Code / Copilot spelling; Module 1 lists the others.

5.Set up (5 min)

5.1 Open the project

Open your Module 4 project, where SPECTRA is already installed, and commit any leftover work so this lab’s files show up cleanly in git status. If you skipped Module 4, start fresh:

unzip lab-codebase.zip -d ~/sdd-assess-lab
cd ~/sdd-assess-lab/lab-codebase
git init && git add -A && git commit -m "baseline"
spectra install

Accept the offer to initialize Spec Kit, pick your agent, and restart it. No npm install needed.

5.2 Add the assess extension

specify extension add assess

The install ends with ⚠ Configuration may be required, pointing at .specify/extensions/assess/. Ignore it; the extension ships no configuration file. Then restart your agent so it loads the five new commands.

6.Intake: capture the idea (5 min)

/speckit-assess-intake "Sync to-dos across devices, so a task I add on my laptop is there on my phone." This is an idea for this repository; read enough of the codebase to record what it relates to. slug=sync-across-devices

The agent writes .specify/assessments/sync-across-devices/intake.md. If it asks for a slug, you left off slug=.

Human gate: you confirm the idea was recorded faithfully, before anyone evaluates it. Check intake.md:

7.Research: evidence for and against (7 min)

/speckit-assess-research slug=sync-across-devices

The agent may search the web. If it asks to fetch a host it doesn’t recognize, the default is no; say yes only for a site you know.

Human gate: you decide whether this research actually challenges the idea. Check research.md:

If something fails, ask the agent to fix that part of research.md. Don’t re-run the stage.

8.Ground it in the code with impact (10 min)

8.1 Run the impact analysis

/speckit-spectra-impact A person who uses the to-do app on more than one device sees the same tasks on each: a task added, completed, or deleted on one device appears on the others without any manual export or import. Today each browser keeps its own separate list.

When it asks whether this repository is the only system involved, answer 1 — just this repository; the document records that you said so. It may then ask up to five questions, one at a time, each with a recommendation. Answer as the product owner, or press Enter to accept it (recorded as defaulted — not confirmed). It writes a numbered analysis and a folder index under docs/impact-analysis/, or your declared artifact root.

Human gate: the analysis is always a draft; you decide what it means. Check that:

8.2 Fold the findings into research

Decide weighs the artifacts in the assessment directory, so evidence that lives only in docs/ won’t reach it. Bring it in:

Add the relevant findings from docs/impact-analysis/<NNN>-<name>.md to
.specify/assessments/sync-across-devices/research.md, under "Data & Constraints"
and "Evidence Against the Idea". Keep each path:line citation as the source,
list the analysis under Sources, and use research.md's own high/medium/low
confidence scale. Add nothing the analysis does not say, and keep every
existing source and confidence tag.

9.Define, then shape (8 min)

9.1 Define the problem

/speckit-assess-define slug=sync-across-devices

Human gate: you confirm the problem is stated as a problem, not a feature. Check problem.md:

9.2 Shape the options — and talk about appetite

/speckit-assess-shape slug=sync-across-devices

Here the idea meets the codebase. The app has no backend, so real sync means building one: somewhere to store tasks, a way to know whose they are, and a rule for when two devices change the same task. Before reading concept.md, decide: how much is this worth to a 200-line, single-browser app? A day? A month? That is the appetite — a budget you set, not an estimate you are handed.

Human gate: check that shaping stayed at concept level and took appetite seriously:

  • Export and import a file (small). No backend. But it isn’t really sync: it is manual, and an old file can overwrite newer tasks.
  • A hosted sync service (medium). Less to build, but accounts, a vendor, ongoing cost, and task data leaving the device.
  • Our own backend with accounts (large). Rabbit holes: conflicting edits, identity, hosting, privacy.
  • Do nothing. Each browser keeps its own list, as the README already says.

10.Decide (5 min)

/speckit-assess-decide slug=sync-across-devices

Expect needs-clarification or kill. Read the scorecard like this:

  1. Problem validity and Evidence strength first — they gate a go. With demand resting on assumptions, Evidence strength shouldn’t be above weak.
  2. Every justification points at an artifact. “Feasibility: weak” should cite the missing backend, not a feeling. Every unknown is acknowledged.
  3. The verdict follows the ratings. Weak evidence plus go is an over-claim; push back.

For this idea, needs-clarification means named unknowns block the call — typically real demand, and the appetite for a backend. kill means cost beats value: a backend, identity, and hosting, for demand nobody has shown (and if your Module 4 constitution kept “no backend”, Strategic fit agrees). A go is a red flag; check Evidence strength.

Human gate: you are the product owner. The verdict is a recommendation; accepting it, or overriding it with a reason, is your call. Commit now so the next step’s changes are easy to review (adjust the second path if you declared another artifact root):

git add .specify/assessments docs/impact-analysis
git commit -m "Assess sync across devices"

A sample. Your ratings will differ; what matters is that each is justified from the artifacts and the verdict follows the rules.

CriterionRatingJustification
Problem validityweakPlausible, but no user signal for this app; demand is ASSUMPTION.
Evidence strengthweakCodebase findings cited from the impact analysis; demand uncited. Research confidence: low.
Value vs. inactionunknownEach browser keeps its own list; no data on how many people hit that.
Feasibility / appetiteweakBeyond manual export, every option needs storage, identity, and a server the app lacks (impact: localStorage only, no API routes, no auth).
Strategic fitweakThe Module 4 constitution declares localStorage-only persistence, no backend.
Risk postureweakConflicting edits, task data leaving the device, running a service — named, none mitigated.

Verdict: needs-clarification. Blocking: [NEEDS CLARIFICATION: Do people use this app on more than one device?] and [NEEDS CLARIFICATION: What appetite will we accept for building and running a backend?] Revisit: research, shape.

A kill is equally defensible (“cost exceeds value: a backend, accounts, and hosting for demand nobody has evidenced”). The indefensible answer is go.

11.Refine instead of re-running (5 min)

Pick one [NEEDS CLARIFICATION: …] marker you can answer and resolve it. Appetite is the easiest, and it lives in concept.md. For this exercise you are the product owner, so a fact you supply is sourced to you:

As product owner, I have confirmed the appetite for this idea: at most one week
of work, and no hosted service we would have to run or pay for. Record that in
.specify/assessments/sync-across-devices/concept.md, sourced to me with today's
date, and adjust any wording there that depended on the old appetite. Then check
whether this resolves a blocker in decision.md; if it does, revise the scorecard,
rationale, and verdict in place. Keep every source and confidence tag, and do not
regenerate either file.

Or edit concept.md yourself — replace the marker with the fact and [source: product owner, <date>] — then ask the agent to revise decision.md.

With that appetite no real sync option fits, so expect the verdict to harden, often to kill. Resolving an unknown doesn’t always move an idea forward; sometimes it closes the question. That is the funnel doing its job.

Human gate: you supplied the fact; the agent must not invent anything around it. Run git diff:

12.Stretch: take a small idea to go, and hand it off (10 min)

Now a proportionate idea. Today finished tasks stay in the list, struck through. Run each stage separately and review each artifact, as before:

/speckit-assess-intake "Let me hide completed tasks so I only see what is still open." This is an idea for this repository. slug=show-open-tasks
/speckit-assess-research slug=show-open-tasks Include prior art from other to-do apps, such as the TodoMVC project on github.com.
/speckit-assess-define slug=show-open-tasks
/speckit-assess-shape slug=show-open-tasks
/speckit-assess-decide slug=show-open-tasks

This should score very differently: small appetite, nothing leaving the browser, a comfortable fit with how the app works, plenty of prior art. A go is likely. If you get needs-clarification, find the criterion that blocked it — usually Evidence strength — and refine that artifact.

Human gate: a go says the idea is worth specifying, not that you must build it. If you go ahead, the handoff is yours:

/speckit-specify Use the handoff summary in .specify/assessments/show-open-tasks/decision.md to specify a way to show only the open tasks in the to-do list.

Or structure it as a business requirement first: run the BRD agent on the decision, then give the BRD it writes to /speckit-specify. Pick one route, so the metrics live once in the chain.

/speckit-spectra-brd .specify/assessments/show-open-tasks/decision.md

Stop there; don’t run /speckit-plan. The point is the moment discovery hands over to delivery.

13.Optional: run the assessment as a workflow

Spec Kit also ships the process as a first-party bundle: the assess extension plus an assess workflow. It needs network access and Spec Kit 1.0.9 or later, and sits outside the 75 minutes. Read what a bundle installs before installing it:

specify bundle info assess
specify bundle install assess
specify workflow run assess --input idea="Let me reorder tasks by dragging them" --input slug="reorder-tasks"

It runs the five stages back to back, then stops at a final review-verdict gate: approve completes the assessment, reject aborts it. A go is still handed off manually. You trade the per-stage reviews for one review of all five artifacts at the end, so a bad intake is caught late. Use a fresh slug: stages won’t overwrite existing artifacts in an unattended run. The workflow suits triaging a queue of ideas the same way every time; the stage-by-stage path suits contested ideas, and learning.

Knowledge check

Knowledge check7 questions
1. Where does idea assessment sit in SPECTRA’s agentic SDLC, and what does it need first?
Answer: b
Assessment decides whether Plan should start at all, and it works in an empty initialized project with no SDD artifacts. Bug fixing is what belongs to Maintain.
2. A scorecard rates Problem validity and Feasibility strong, Evidence strength weak, and a concept is recommended. What is the most positive verdict it can honestly record?
Answer: c
A go needs Problem validity and Evidence strength at adequate or better, plus a recommended concept. Weak evidence caps the verdict at needs-clarification however strong the rest is. kill is possible, but not the most positive.
3. A research.md lists six findings, all in favour, and has no “Evidence Against the Idea” section. What do you do?
Answer: d
The section is required every time, and “none found” must be said explicitly. You fix it by refining the artifact; skipping research never waives the evidence bar, and regenerating everything is not how gaps are filled.
4. The verdict is needs-clarification because concept.md has an appetite marker. The product owner now says: one week, no hosted service. What next?
Answer: a
Clarifications are resolved by refining the affected artifact, then the decision. Re-running is for exceptions such as a wrong slug, and asks before overwriting. Handing off an unresolved decision skips the gate.
5. What does /speckit-spectra-impact add that the assess stages don’t produce themselves?
Answer: c
Impact traces what a change would disturb, with citations and a stated coverage boundary. It cannot show demand — code says nothing about whether anyone wants a feature. The verdict belongs to decide, and impact never creates or links a spec.
6. show-open-tasks reaches go. What happens next?
Answer: b
The extension registers no hooks, and decide only prepares the handoff; it never writes a spec. Even the workflow stops at its review-verdict gate. Crossing into delivery is a deliberate human step.
7. Your sync assessment ends in kill. A teammate says the hour was wasted, because nothing got built. What do you tell them?
The kill is the deliverable. For an hour of reading, we learned that sync means building and running a backend, accounts, and conflict handling the app doesn’t have (cited in the impact analysis), for demand nobody has evidenced (every demand claim is an ASSUMPTION), against a constitution that says localStorage only. Otherwise we would have found out weeks into a plan. And because decision.md records the decisive reason, whoever proposes sync again — say, once real users ask — can reopen it with new evidence against a named assumption instead of restarting the argument.

What's next

This module asked whether something new should be built. Module 6, Bug Fix, turns to something already built that is broken. You will take a real defect in a fresh copy of the lab app through Spec Kit’s bug workflow — assess, fix, verify — with a diagnosis before any code changes, a fix that stays inside what the diagnosis covers, and verification against the original reproduction.