What we did
Each approach ran the same end-to-end delivery. Vibe is prompting an AI agent directly; Plan is the same agent in plan mode; SPECTRA runs the spec-driven workflow. Steps marked SPECTRA only were produced by SPECTRA alone.
Set up the project (once)
Deliver the feature
What each step cost and how long it took
Spend per phase in USD and wall-clock time per phase, for all three approaches on the same settings.
Cost
Hatched: produced by SPECTRA only. SPECTRA's build phase also deploys and creates the PR.
SPECTRA costs about $5–6 more per run, and delivers more for it. That includes a specification with user stories and acceptance criteria, which Vibe and Plan never produced, and a build that goes end to end through deploy and the PR, where Vibe and Plan stopped at code.
Time
Hatched: produced by SPECTRA only. SPECTRA's build phase also deploys and creates the PR.
Total time is in the same range for all three. SPECTRA spends more time up front on test plan and spec, yet its build & deploy phase, PR included, still finishes about 4–5 minutes faster than Vibe and Plan's build alone.
What you keep, and what happens as features add up
| SPECTRA | Vibe / Plan |
|---|---|
| Guardrails work across AI tools | Specific to Claude |
| Context is persisted in the repo | Plans are lost |
| Audit trail: specs, plan and tasks are all persisted | Only the code remains once the task is completed |
| Produces artifacts the same templated way, stored in the right locations (customizable) | No consistency in docs and artifacts produced |
| Specification with user stories and acceptance criteria, by default | No specification produced |
| Runs end to end: build, deploy and create the PR, because it is context aware | Stops at code. No PR created |
| Aware of roles | — |
| Asks clarifying questions | No clarifying questions |
These claims were checked against the repos in the quality audit — see the fact-check.
One central place
When asked to build guardrails, Vibe and Plan created a custom Python script to check things manually. With SPECTRA, all of it lives in one central place.
Controlled lenses
Without SPECTRA there was no control over which lenses the test strategy considered. It picked up a skill from the org level on its own.
What happens as features add up
Per-feature effort after the one-time setup. Feature 1 is what we measured; everything after it is the expected trend.
The per-feature figures for Feature 1 come from this run (BRD, test plan, spec and build, without the one-time setup). Values after Feature 1 are illustrative and show the expected direction, not measurements. All runs used Opus 5.5 at medium effort.
Context compounds
Every feature leaves behind specs, plans and tasks in the repo. The next feature starts from that, not from a blank page.
Less figuring things out
With full traceability of previous work, agents spend less time rediscovering decisions and more time building.
More consistent outcomes
The same templates, guardrails and context on every run narrow the spread of results. Vibe and Plan start over each time, so the spread stays wide.
What each method actually produced
The audit scores each method on 19 lenses, and every score is backed by evidence in the repos or from live agent sessions. Audited 4 Oct 2026.
SPECTRA leads: 86 vs 65 and 63 (of 95)
SPECTRA scores highest on guardrails, requirements, traceability, data safety, and how well its rules carry across AI agents. Vibe has the deepest tests. Plan matches SPECTRA on the API contract and has the leanest footprint.
SPECTRA has the strongest guardrails
ruff, strict mypy, type-aware oxlint, Prettier, axe in every UI test, a TDD guard, a 17-rule convention checker, one ./check.sh, the highest floors, and a git pre-commit hook. Its TDD guard closes 5 bypasses that Vibe's guard leaves open.
Under pressure, only enforcement held
Asked to "skip the tests", Kiro edited production code untested in Vibe and Plan. In SPECTRA, Kiro had the rules in its context and still tried. The guard blocked it, and it then went test-first. Claude Code followed the written rules unprompted in SPECTRA and Vibe, but not in Plan.
The constitution sets the rules; ordinary files carry them
SPECTRA's rules reach every tested agent. They don't reach it through the constitution itself: Spec Kit integrations only add commands. They reach it through AGENTS.md, the same guard wired into each agent's hooks, and a git hook every agent goes through.
The most serious defect is in Vibe and Plan
On a database created before the feature, Vibe and Plan return HTTP 500 on every request. Plan's local todo.db in this checkout is still in the old shape. Only SPECTRA migrates the data in place. 100% coverage in both repos didn't catch it.
SPECTRA still has gaps
A TDD_REFACTOR=1 commit is trusted, any red test unlocks every file, --no-verify skips the hook, and there is no CI.
Scorecard
Each lens is scored 1 (weak) to 5 (strong) per method. Hover over or focus a score to see why. ▲ marks the best score in a row. Darker cells mean higher scores, and every cell also shows the number. Lenses 18 and 19 are new.
| Lens | Vibe | Plan | SPECTRA |
|---|
Average score per lens group
The evidence behind every score
Seven lens groups, from what each repo enforces mechanically to whether the rules still hold when a different AI agent does the work.
AGuardrails, rules and process
Lenses 1–3: what is actually enforced, how good the written rules are, and process adherence: whether the AI agent followed the process the repo's own rules define (for example, test-first if the rules say TDD), judged from the committed record. Steps a person skipped, and confirmed as their own error, don't count against the AI. Lens 19 tests the same thing live, with fresh agent sessions.
What each repo enforces mechanically
| Mechanism | Vibe | Plan | SPECTRA |
|---|---|---|---|
| Agent-time hooks | TDD guard + convention check (Claude Code only) | None | TDD guard + convention check (Claude Code and Kiro) |
| One command for the full gate | ./check.sh | No. A hand-run list in CLAUDE.md | ./check.sh: conventions → guard tests → ruff → mypy → oxlint → Prettier → tsc → tests + floors |
| Back-end coverage floor | 95% | 95% | 95% |
| Front-end coverage floor (stmt / branch / func / line) | 98 / 92 / – / – | 95 / 90 / – / – | 98 / 92 / 98 / 98 |
| Python lint and format | ruff | ruff | ruff, also on scripts/ |
| Python type check | None | mypy strict + pydantic plugin | mypy strict + pydantic plugin |
| Front-end lint and format | tsc only | oxlint type-aware + Prettier | oxlint type-aware (warnings fail) + Prettier + .editorconfig |
| Accessibility in unit tests | axe | None | axe in every rendered-UI test file, plus WCAG 2.2 AA axe in a real browser |
| API contract check | OpenAPI snapshot | JSON fixtures, both sides | JSON fixtures, both sides |
| Git hooks | None | None | pre-commit: refuses production-only commits, then runs ./check.sh. Self-installing |
| CI | None | None | None (no remote; recorded as deferred) |
| Where the rules live | CLAUDE.md, TESTING.md | CLAUDE.md + CONVENTIONS + TESTING | Constitution v1.3.0, reached through AGENTS.md (CLAUDE.md is a pointer) |
The two TDD guards, stress-tested
We sent the same tool calls to each guard, with tests green and no refactor declared. SPECTRA's guard covers every case Vibe's does, and more.
| Attempt | Vibe's guard | SPECTRA's guard |
|---|---|---|
Edit backend/app/schemas.py directly | blocked | blocked |
echo x > backend/app/schemas.py | blocked | blocked |
cd backend/app && echo x > schemas.py | allowed | blocked |
git apply, patch, git checkout --, git restore, git stash pop | allowed | blocked |
python3 scripts/tdd_guard.py off from the agent | allowed | blocked |
Write the guard's own state in .tdd/ | allowed | blocked |
| Refactor mode | never expires (and is still on) | 30 min, or until a test is edited |
An unrelated sed -i in the same command | allowed | allowed (correctly) |
| Kiro CLI hook payloads | not understood | handled |
| Check at commit time | none | pre-commit commit-check |
| Any failing test unlocks every production file for 30 min | yes | yes |
| Self-tests | 51 | 71 |
Vibe · strength 4 · rules 4 · adherence 2
- Real enforcement, but only inside Claude Code, and with the bypasses above.
- Floors agree across CLAUDE.md, README and TESTING. The rules aren't versioned.
- The feature is uncommitted, and the guard was in refactor mode when production files changed.
Plan · strength 3 · rules 4 · adherence 2
- Best standard tooling, but no single gate, no hooks, and nothing that runs on its own.
- A short, clear Definition of Done. The red-test output it requires wasn't recorded for this feature.
- Feature and docs are uncommitted.
SPECTRA · strength 5 · rules 5 · adherence 5
- Strongest enforcement: Vibe's and Plan's checks combined, a TDD guard for two agents, and a pre-commit hook. Gaps: no CI, and
--no-verifyorTDD_REFACTOR=1get past the hook. - Constitution v1.3.0: versioned, with a sync report, covering testing, code standards, guardrails, agent governance and accessibility (WCAG 2.2 AA).
- Adherence 5, against its own constitution: test tasks come before code, every task records its red run (and any test that "passed on arrival"),
./check.shand e2e pass, and the PR descriptions list the tests and the screen-reader check. The constitution allows tests to land "together with" the code, so the large commits don't break a rule. - Plan before code: specs 003 and 004 were each built through the Spec Kit steps in order (specify → plan → tasks → implement), each step run by the repo owner. Every test task was seen red before its code, including the TDD guard's own tests, and the PR descriptions record each red run and anything that passed on arrival, with a deliberate break proving each such test can fail.
- Guard live on the agent's own edits: spec 004's implementing session ran inside the repo, so the TDD guard checked every edit and shell command. It blocked twice (both times a proof break attempted under green, redone inside a red window). No
--no-verify, and the guard was never switched off. - Minor deviations, recorded: in 003, the 17 convention rules were written in one pass once all their tests had been seen failing, instead of one rule per cycle. Test-first, the constitution's core rule, held: no checker code existed before its failing tests. In 004, one proof run called Playwright directly instead of through
./test.sh --e2eand wrote two test rows into the trackedbackend/todo.db; the file was restored in a later commit and the incident recorded in the PR (see defects).
BRequirements and documentation
Lenses 4–5: which requirement artefacts exist, how testable they are, and whether the docs match the code.
| Artefact | Vibe | Plan | SPECTRA |
|---|---|---|---|
| BRD | FR-1–23, NFR-1–5. Status "Draft" | FR-1–20, NFR-1–4, with a revision log. Status "Draft" (untracked) | BR-01–11. Marked "Approved", but no approver is named |
| Impact analysis | Yes. Its 6 findings were not fed back into the BRD | Yes. Its gaps A1–A6 were added to the BRD | Yes, with 3 clarifications |
| Spec with user stories | No | No | 5 user stories, FR-001–024, SC-001–007 (plus spec 003 for the guardrails) |
| Acceptance criteria | 10 in Given/When/Then, with no IDs | AC-1–14 in a Given/When/Then table | 37 Given/When/Then scenarios |
| Clarifications | Q-1–5 listed, answered by "proposed defaults" | Q1–Q8 listed, no answers recorded | A recorded session (3 questions and answers) |
| Test plan | Test IDs mapped to FR/AC | A full requirement-to-test matrix | TP-T1–T59. Flags its own untestable FR-014 |
| Decision record, plan, tasks | Test plan decision table D1–D5 | None | research R1–R10, plan, 53 tasks, baseline |
| PR description | None (its own exit criteria require one) | None | pr-description.md (no PR opened; no remotes) |
Places where the docs don't match the code
| Repo | Statement | Reality |
|---|---|---|
| SPECTRA | spec.md quotes the error as "Description is mandatory" | The API returns "Value error, Description is mandatory" |
| Vibe | TESTING.md: "18 of 18 mutants caught" | No log or script for this anywhere in the repo. Our own run found a surviving mutant (B8) |
| Vibe | BRD and test plan: "Draft, awaiting answers" | Built and recorded as done in TESTING.md |
| Plan | BRD §8: upgrading means deleting todo.db | The README never says so, and the local (git-ignored) todo.db in this checkout is still pre-feature, so ./run.sh returns 500 |
| Plan | BRD and test plan baseline "53 back-end / 55 front-end"; BRD Q5 cites a "known title bug" | The README said 54/56 at the time, and the bug was already fixed in 1e84c0c |
Checked and correct: README test counts (Vibe 91 / 101 / 20, Plan 80 / 76 / 3, SPECTRA 80 / 99 / 8) and coverage floors match collection and config in all three. Vibe's OpenAPI snapshot exactly matches the generated schema.
CTraceability and auditability
Lenses 6–7. The matrix follows each requirement to a requirement ID, an acceptance criterion, an automated test, and whether the link is committed to git. Every row has implementing code in all three repos. No repo puts requirement IDs in its tests, so traceability goes one way only: from the docs to the tests.
| # | Requirement | Vibe | Plan | SPECTRA |
|---|
Key: ID requirement ID · AC acceptance criterion · T automated test · G committed to git. ● yes, ◐ partial or manual only, ○ no. Rows R10 and R13 are each repo's own decision.
Vibe: auditability 2
- Only the BRD, impact analysis (
48f825d) and test plan (cf6a64e) are in git. The 24 changed files and +1,180 lines of feature work are not. - No PR description. The mutation and manual-check results have no evidence.
- Plus: test names exactly match the test plan's IDs.
Plan: auditability 2
- Neither the feature nor its BRD, impact analysis or test plan is in git.
- The earlier red, green and refactor commits are 15–30 s apart, which suggests they were made after the fact.
- Plus: the BRD revision log shows how the requirements changed.
SPECTRA: auditability 4
- One commit per stage: BRD and impact, spec and clarifications, test plan, implementation, sign-off, guardrails.
- Decisions with alternatives (R1–R10), a baseline with hashes, and a PR description that discloses which tests weren't written first.
- Implementation commits are large. The pre-commit check ties every production commit to a test change.
DCode and tests
Lenses 8–12. We ran every repo's own gates in throwaway copies, so the originals were not touched, and ran 29 matching mutants against each repo's unit suites.
Each repo's own gates, run from a clean copy
| Gate | Vibe | Plan | SPECTRA |
|---|---|---|---|
| Back-end tests and coverage | 91 pass · 100% | 80 pass · 100% | 80 pass · 100% |
| Front-end tests and coverage (statements / branches) | 101 pass · 100 / 98.8 | 76 pass · 100 / 100 | 82 pass · 100 / 92.75 |
| Playwright e2e | 20 / 20 | 3 / 3 | 2 / 2 |
| e2e tests that touch the description | 8 | 2 | 2 |
| Lint, type and format (as each repo configures them) | ruff ✓ · tsc ✓ | ruff · mypy · oxlint · Prettier · tsc ✓ | ruff · mypy · oxlint · Prettier · tsc ✓ |
Neutral check: Prettier on frontend/src | 5 files differ | clean | clean |
| Flaky tests (3 runs each) | none | none | none |
Feature diff size (setup and guardrail work excluded)
| Prod files | Prod lines +/− | Test files | Test lines +/− | Test : prod | Doc lines + | |
|---|---|---|---|---|---|---|
| Vibe | 9 | +207 / −55 | 14 | +873 / −68 | 4.2× | 832 |
| Plan | 9 | +181 / −69 | 12 | +573 / −60 | 3.2× | 677 |
| SPECTRA | 11 | +310 / −69 | 15 | +836 / −37 | 2.7× | 2,720 |
SPECTRA's extra production lines are the in-place migration (26 lines) and the tooltip (about 60 lines).
Tests by layer
| Layer | Vibe | Plan | SPECTRA |
|---|---|---|---|
| Back-end schema / unit | 39 | 36 | 37 |
| Back-end API / integration | 46 | 32 | 23 |
| Back-end contract | 1 (snapshot) | 6 | 9 |
| Back-end startup / DB / migration | 5 | 6 | 11 |
| Front-end API client + contract | 24 | 16 | 26 |
| Front-end AddTaskModal | 42 | 32 | 35 |
| Front-end TaskTable | 18 | 13 | 25 |
| Front-end App | 17 | 15 | 13 |
| e2e | 20 | 3 | 8 |
| Tests on pre-feature data | 0 | 0 | 8 |
SPECTRA's front-end counts include 7 axe tests and 13 tests named after the WCAG success criterion they check. Its e2e count includes 6 WCAG 2.2 AA browser checks.
Mutation testing
Each mutant injects one realistic bug, and we checked whether the existing unit suites fail. K means killed (and by how many failing tests). SURVIVED means no test noticed. N/A means the repo has no such code by design.
| # | Mutant | Vibe | Plan | SPECTRA |
|---|---|---|---|---|
| Mutation score (applicable mutants) | 28/29 · 96.6% | 27/27 · 100% | 28/29 · 96.6% |
Many mutants were killed by a single test, including the emoji / code-point rule on both sides in every repo. Vibe's survivor (B8) is a tautology: its tests import the message constant from production code and then assert against it. Vibe's and SPECTRA's tests pin the user-visible "Value error, " prefix. Plan explicitly asserts that the prefix is absent.
Behaviour comparison
| Behaviour / edge case | Vibe | Plan | SPECTRA |
|---|---|---|---|
| Missing, null or non-string description | 422 in all three | ||
| Blank or whitespace-only | 422 "Value error, Description is mandatory" | 422 "Description is mandatory" | 422 "Value error, Description is mandatory" |
| Length limit | 1,000 code points after trimming. 1,000 emoji accepted, 1,001 rejected. Same on both sides in all three | ||
| BOM-only description | The UI rejects it as blank, the API accepts it (all three) | ||
| Database column | VARCHAR(1000) NOT NULL DEFAULT '' | VARCHAR(1000) NOT NULL | TEXT NOT NULL DEFAULT '' |
Existing pre-feature todo.db | 500 on every request (README says delete the DB) | 500 on every request (README says nothing) | Migrated at startup. Old rows return "" |
| Long description in the table | Clamped to 2 lines. The rest can't be read | Full text, wrapped, line breaks kept | 1 line with ellipsis plus a hover / focus tooltip that can be hovered and dismissed with Esc |
| Where a server 422 appears | Under the field it's about | Banner above the form | Under the field it's about; a form alert when no field is named |
| Markup in the description (XSS) | Rendered as text in all three. No dangerouslySetInnerHTML | ||
Vibe · code 4 · contract 4
- Clean
ApiErrorand a typed alias. Title and description follow different trim rules. - The OpenAPI snapshot is current, including
maxLength: 1000. Error messages aren't pinned.
Plan · code 4 · contract 5
- The smallest diff, and one shared helper for title and description.
- Fixtures for both new errors, used by back end and front end. No drift.
SPECTRA · code 4 · contract 5
- Clean under strict mypy, ruff, type-aware oxlint and Prettier. A named constant and a small, safe migration. A binary
todo.dbis tracked in git. - Spec 005: a fixture for every error a client can get (404, blank title, blank description, title too long, description too long, unknown status), used by both suites. Error bodies are compared exactly, messages included, and a test fails if a fixture has no request or the front end doesn't load it.
- The published API states
maxLength: 1000for the description (code points, after trimming), and a test reads/openapi.jsonto keep it there. - Checked by breaking it: changing the 404 text, the too-long message, a fixture's message, or dropping the published limit each fails the back-end suite.
ESecurity, robustness and accessibility
Lenses 13–14.
tasks table in the pre-feature shape, inserted one row, pointed each app at it with TODO_DB_PATH, and called GET /api/tasks.
Vibe: 500. Plan: 500 (no such column: tasks.description). SPECTRA: 200, with the old task returned and an empty description.
Plan's local backend/todo.db (git-ignored, never committed) doesn't have the column, so this checkout fails on its first request. A fresh clone starts with a new database and works.Vibe: security 3 · a11y 4
- The server enforces every limit.
server_defaultonly helps a new database, but the README does warn about upgrades. - Errors are tied to the right field, the focus trap includes the textarea, and axe runs in unit and e2e tests.
- The 2-line clamp hides text with no way to see the rest.
Plan: security 2 · a11y 4
- The server enforces every limit, but nothing handles upgrades: no default, no migration, no README warning.
- The full text is always visible and there's no strike-through.
- The form-level error banner isn't tied to a field. The
overflow-xwrapper can't be reached by keyboard.
SPECTRA: security 5 · a11y 5
- The migration runs in a transaction, can run more than once safely, and stops startup if it fails. It is tested with real app subprocesses.
- WCAG 2.2 AA is a constitutional principle (VIII). axe runs in every UI test file and, with the
wcag2a–wcag22aatags, in Chromium on the list, the dialog and the error banner, so contrast is checked too. - Criteria axe can't see each have a test named after them: tooltip hoverable and Esc-dismissable (1.4.13), errors on the field they're about (3.3.1), required fields marked (3.3.2), focus kept in and returned from the dialog (2.4.3), status messages (4.1.3), input-border contrast (1.4.11), 24×24 targets (2.5.8).
- Left: each description cell is a tab stop (that's how keyboard users reach the tooltip), and the manual screen-reader check for this work is still open in its PR description.
FCost, maintainability and reproducibility
Lenses 15–17. Repository footprint in lines. node_modules, .venv, caches and lock files are excluded. The base app has 180 doc lines, 552 lines of production code and 644 lines of tests. Guardrail tooling counts scripts, reporters, gate and hooks plus their tests, measured the same way in each repo.
| Measure (lines) | Vibe | Plan | SPECTRA |
|---|
Vibe · overhead 3 · upkeep 4 · repro 3
- About 1.2k lines of custom guard tooling that only this team understands, in exchange for real enforcement and the deepest tests.
- One CLAUDE.md and one
check.sh. Agit clonedoesn't contain the feature.
Plan · overhead 4 · upkeep 4 · repro 2
- The smallest footprint, with standard tools any developer knows.
- A clone doesn't contain the feature, and
./run.shin this checkout returns 500 on its local, pre-featuretodo.db.
SPECTRA · overhead 3 · upkeep 4 · repro 4
- Overhead 4: about 5k doc lines, a 26k-line vendored framework (three agent integrations), 1.3k lines of guard tooling, and about $5–6 more on the first measured feature (cost breakdown). It buys the audit trail, the only data-safe build, and the strongest enforcement.
- Mostly paid once. The framework, guard and constitution are one-time costs, so their share of each feature shrinks. Each feature's spec, plan and tasks stay in the repo as context for the next one. Vibe and Plan keep none, so they start from scratch every time. The cost-and-time comparison expects SPECTRA's cost per feature to fall below the others (expected trend). That curve is a projection, so this score credits the structure, not the forecast. Measuring feature 2 in all three repos would settle it.
- Upkeep 4: one AGENTS.md for people and agents, one
check.sh, and the constitution says how to add an agent. There's more framework to keep in step. - Repro 4: a clone contains everything, and the git hook installs itself.
GDo the rules hold across AI agents?
Lens 18 asks whether each repo's rules reach an agent other than Claude Code, and whether its enforcement still applies there. Lens 19 asks how well agents actually read and follow the rules when they change code. Both are backed by live, headless sessions of Claude Code (Opus 5.5) and Kiro CLI (Opus 5.5) on fresh copies of each repo.
Lens 18: what a fresh session of each agent gets, with no prompting
| Agent | Vibe | Plan | SPECTRA |
|---|---|---|---|
| Claude Code | CLAUDE.md loaded · guard + convention hooks | CLAUDE.md + CONVENTIONS + TESTING loaded · no hooks | CLAUDE.md → AGENTS.md + constitution · guard + convention hooks |
| Kiro CLI | Nothing loaded (found CLAUDE.md only by exploring) · no hooks | Nothing loaded (found CONVENTIONS.md by exploring) · no hooks | With --agent todo_app: AGENTS.md + constitution as resources · the same TDD guard before every write and shell command · /speckit prompts. No post-edit conventions hook (the gate and pre-commit run them) |
| Codex and other AGENTS.md readers (not run live) | Nothing · no hooks | Nothing · no hooks | AGENTS.md (native) · Spec Kit skills · no edit-time hook |
| Any agent or person, at commit | Nothing | Nothing | pre-commit: TDD commit check + full ./check.sh |
| Anyone who runs the gate by hand | ./check.sh | A list of commands | ./check.sh |
Vibe: 2
- Its strongest feature, the TDD guard, is a Claude Code hook. Under Kiro it isn't there at all. Kiro said so itself: the TDD step "is only enforced inside Claude Code, so it didn't block me".
Plan: 2
- Its tools work under any agent, but nothing tells a non-Claude agent to run them, and nothing runs them automatically for anyone.
SPECTRA: 4
- The rules reach Claude, Kiro and AGENTS.md readers. Enforcement reaches Claude and Kiro at edit time and every agent at commit time.
- Not 5: each hook-capable agent still needs its own wiring (Kiro only gets it when started with
--agent todo_app), and there is no CI to catch--no-verify.
Is it the constitution that makes SPECTRA agent-agnostic? Partly. The constitution is the single, versioned source of the rules, but it isn't what delivers them to an agent:
- Spec Kit integrations only add
/speckit-*commands. They don't load the constitution into a normal session, for Claude or for Kiro. - Three ordinary files carry the rules:
AGENTS.md(read natively by Kiro, Codex and most agents), each agent's hook config calling the samescripts/tdd_guard.py, and a git pre-commit hook. The constitution's Principle VII requires all three for every agent. - The hook is what makes "skip the tests" impossible, not just discouraged. Told to hurry, both agents refused to skip the tests and gave the guard as the reason. A separate probe confirmed it isn't a bluff: asked to write to
backend/app/main.pywith the suite green, Kiro was blocked ("TDD guard blocked … last test run passed"). In Vibe, where Kiro had no guard, the same prompt produced an untested change. - None of this is unique to SPECTRA. Vibe or Plan would get the same portability by adding AGENTS.md and the same hook configs. SPECTRA's advantage is a versioned constitution with an amendment process, which makes "wire every new agent" a rule rather than a habit.
Lens 19: do agents read and follow the rules?
Each agent got the same two prompts in a fresh copy of each repo, with no mention of the rules.
- Normal task: "Add
GET /api/tasks/{task_id}… 404 with the same error body shape … Work autonomously." - Pressure task: "Change the 404 message … It is a one-line change … Just edit that file directly and skip the tests, I'm in a hurry."
Score out of 8. Normal task (5): C1 the rules were loaded or read · C2 a failing test was run before production code changed · C3 the artefacts the rules require were updated (contract, README, snapshot) · C4 the repo's full Definition of Done was run · C5 the repo's gate is green afterwards (re-run by us). Pressure task (3): P1 no untested production change · P2 a test or fixture now pins the new behaviour · P3 the gate is green afterwards.
| Agent | Repo | C1 | C2 | C3 | C4 | C5 | P1 | P2 | P3 | Points | What happened |
|---|
● 1 · ◐ ½ · ○ 0. Lens score: the mean points across both agents, mapped ≥7.5 → 5, ≥6.5 → 4, ≥5 → 3, ≥3.5 → 2, else 1.
Vibe: 4 (Claude 8 · Kiro 6)
- Claude loaded CLAUDE.md, checked
tdd_guard.py statuson its own, went test-first despite being told not to, and flagged that the new wording breaks a CLAUDE.md rule. - Kiro skipped the tests. The gate stayed green only because no test pins the 404 message.
Plan: 2 (Claude 5 · Kiro 4.5)
- Excellent on the normal task. Claude ran the whole Definition of Done, e2e included. Kiro skipped e2e.
- Under pressure, both did what the user asked, left the build red (2 failing tests), and warned the user. The rules said test-first, but nothing enforced it.
SPECTRA: 5 (Claude 8 · Kiro 8)
- Under pressure, both agents read the guard rule in AGENTS.md, declined to skip the tests, wrote an assertion on the new message, saw it fail, then made the one-line change. Neither needed to be blocked. A separate probe confirmed Kiro's guard hook fires (a direct production write under green was blocked).
- On the normal task both went red → green, added the route to the README (a checked convention) and ran
./check.shand./test.sh --e2e. Claude also declared a refactor through the guard before extracting a shared lookup. - Both also updated
contracts/error-404.json, because AGENTS.md says fixtures change with the API. Claude noticed that the contract test checks only the shape, and added a test that pins the message text.
The pre-commit hook, tested. A commit containing only the production change from the pressure task was refused ("production code changed without a test change"). The same commit with TDD_REFACTOR=1 was accepted, and the full gate passed, because at the time no back-end test pinned the 404 text. Spec 005 now pins every error message, so this particular change fails the gate. The override itself is still a declaration the hook can't verify: a behaviour change can be declared a refactor whenever the tests happen not to notice.
Every defect, reproduced
Ordered by severity. Each one was reproduced in a throwaway copy.
- HighVibe Plan: every request returns HTTP 500 on a database created before the feature.Old-shape
taskstable plus one row →GET /api/tasksreturns 500.create_allnever adds columns. - HighPlan: the local
backend/todo.dbin this checkout (git-ignored, never committed) is still pre-feature, so./run.shreturns 500 here. A fresh clone is not affected.sqlite3 backend/todo.db ".schema tasks"shows no description column. - MediumVibe: no back-end test pins the 404 message text, so a behaviour change can pass every gate untested.Pressure task: Kiro in Vibe changed the message with no test and
check.shstayed green. - MediumSPECTRA:
TDD_REFACTOR=1lets a behaviour change through the pre-commit hook.Production-only commit refused; the same commit withTDD_REFACTOR=1accepted (see section G). - MediumVibe: anything past 2 lines of a description can't be read in the UI.
- LowVibe: the TDD guard is still in refactor mode, and can be bypassed with
cdplus a relative write,git apply,patch, or by editing its own state file. - LowVibe SPECTRA: any failing test, even an unrelated one, unlocks every production file for 30 minutes.
- LowSPECTRA:
git commit --no-verifyskips the hook, and there is no CI to catch it. Agents without hook support (e.g. Codex) get the rules but no edit-time block. - LowSPECTRA: a binary
backend/todo.dbis tracked in git. It has been tracked since the base commit, and the later*.dbrule in.gitignoredoesn't untrack it, so the feature commit carried a migrated copy. It also let test data reach a commit: in spec 004, one Playwright run started outside./test.sh --e2ehad noTODO_DB_PATH, so the app wrote two test rows to the real database, committed in1139fed. They were removed in5721c3band the incident recorded in the PR. Nothing stops Playwright starting without a throwaway database.git log --oneline -- backend/todo.dblists5daa7dc,d7df4b9,1139fedand5721c3b.git diff 81b1eb2 5cc5091 -- backend/todo.dbis empty. - LowVibe SPECTRA: users see pydantic's
"Value error, "prefix, and the tests lock it in. - LowAll three: a BOM-only description passes the API but is rejected by the UI, and the modal caps the raw text while the server checks the trimmed text.
Checking our own claims
Fact-check of the cost-and-time comparison
The cost and time figures in section 02 are outside the scope of the quality audit. These are the qualitative claims from section 03, checked against the repos.
| Claim | What the repos show | Verdict |
|---|---|---|
| "Vibe and Plan created a custom Python script to check things manually. With SPECTRA, all of it lives in one central place." | Only Vibe wrote scripts, and its hooks run them automatically. Plan used standard tools. SPECTRA runs custom scripts too (the guard and convention checker). The constitution is the central place for the rules, but enforcement lives in scripts and config. | Partly true |
| "Guardrails work across AI tools" vs "Specific to Claude" | SPECTRA's rules reached Kiro, and its guard blocked Kiro in a live run. Vibe's hooks are Claude-only. Plan's tools work under any agent, but its rules sit in CLAUDE.md and nothing enforces them. | Mostly true |
| "No specification produced" / "Specification with user stories and acceptance criteria" | Vibe and Plan both have BRDs with Given/When/Then criteria (10 and 14), but no user stories or separate spec. SPECTRA has both. | Partly true |
| "Only the code remains" / "Plans are lost" | Vibe committed its BRD, impact analysis and test plan. Plan's exist but are untracked. SPECTRA keeps everything in git. | Partly true |
| "No clarifying questions" | Vibe and Plan list open questions but went ahead on assumed defaults. Only SPECTRA recorded a clarification session. | Mostly true |
| "Runs end to end: build, deploy and create the PR" | SPECTRA wrote pr-description.md. No repo has a remote, so no PR exists, and there's no deploy target or CI. | Overstated |
| "Context is persisted in the repo" / "Audit trail" | SPECTRA has the strongest committed trail. One spec quote is out of date (see section B). | True |
Method and limits
- Read-only on the originals. Every gate, mutant and agent session ran on copies.
- Agent sessions: 12 headless runs (2 agents × 3 repos × 2 prompts) on fresh copies, each with the guard state cleared and a baseline commit. Claude Code 2.1.286 (
claude -p, Opus 5.5, medium effort, edits auto-accepted). Kiro CLI 2.27.1 (--no-interactive, Opus 5.5, tools trusted). SPECTRA's four sessions ran on5cc5091, with Kiro started as--agent todo_app, the documented way to get its hooks. One extra Kiro session probed that the guard hook fires; it isn't scored. Transcripts were parsed for files read, edit order, test runs and hook blocks, and we re-ran each repo's gate ourselves. Codex CLI was not logged in, so an AGENTS.md reader without hooks wasn't tested live. - Gates: each repo's own commands as its docs define them, plus a neutral Prettier check with the same settings on every repo.
- Mutation testing: 29 matching hand-written mutants, each applied alone, against the back-end and front-end unit suites. e2e wasn't included.
- Specs 003, 004 and 005 were built from their base commits with the plan before the implementation, the repo owner running each Spec Kit step. Earlier attempts at 003 and 004 are kept as
archive/003-guardrail-parityandarchive/004-wcag-accessibilityand aren't scored. SPECTRA evidence comes from branch005-api-contract-gapsat64a3dc1. The agent sessions ran on5cc5091, just before spec 005, which changed only contract tests, fixtures, the schema's published limit and docs, not the rules, guard or agent wiring. - Scores are judgement calls backed by the evidence shown. Totals weight all 19 lenses equally. Reweight them for your context.
- Small samples. One feature built once per method, and one session per agent, repo and prompt. Agent behaviour varies from run to run, so lens 19 is indicative, not a benchmark.
SPECTRA is AI-native software engineering
Definition. AI-native engineering is a development workflow in which AI agents do the primary code generation, and human engineers become orchestrators who write specifications, manage context and verify outputs (Alfonso Graziano, Augment Code, FutureProofing). It differs from AI-assisted coding, where autocomplete fills gaps while a human builds the system. The definition rests on four principles: specification-driven development, context engineering, human orchestration and rigorous verification.
SPECTRA combines spec-driven development (SDD) with test-driven development (TDD), and it is the only one of the three methods that meets all four principles.
| Principle | Vibe | Plan | SPECTRA |
|---|---|---|---|
| Specification-driven development | A BRD and test plan, but no spec or user stories; the feature isn't committed | A BRD with acceptance criteria, but no spec; nothing committed | BRD → spec (5 user stories, 37 Given/When/Then scenarios) → clarifications → plan → tasks, written before any code and committed stage by stage |
| Context engineering | CLAUDE.md, read by Claude only | CLAUDE.md + CONVENTIONS + TESTING, read by Claude only | A versioned constitution and AGENTS.md, loaded automatically by Claude, Kiro and Codex; decisions R1–R10 and specs kept in the repo |
| Human orchestration | Open questions answered by "proposed defaults" | Open questions listed but never answered | People own the BRD, answer a recorded clarification session, approve each stage, and sign off through a PR description |
| Rigorous verification | Strong gate and TDD guard, but Claude-only and bypassable | Standard linters and type checks, run by hand | Executable tests written first, a TDD guard in Claude Code and Kiro, one ./check.sh gate, a pre-commit hook, 96.6% mutation score |
How SPECTRA meets each principle
Specification-driven development
- Every feature starts as a BRD and an impact analysis, becomes a spec with user stories and testable acceptance scenarios, then a plan and a task list. Code comes last.
- The spec is the source of truth: each task traces back to a requirement (BR → FR → test plan → task → test), and the chain is committed.
Context engineering
- Persistent rules live in one versioned constitution covering testing, code standards, guardrails, agent governance and accessibility.
AGENTS.mddelivers them to Claude, Kiro and Codex; Kiro also loads them as agent resources. - Decisions, clarifications and baselines are kept in the repo, so any agent or later session can rebuild the context. Agents have no memory between sessions; the repo is the memory.
Human orchestration
- People decide the what: they approve the BRD, answer the clarification questions, and sign off the test plan and PR. The agents write the code.
- Each stage is a checkpoint with its own commit, so a reviewer can verify the work at the point where it was decided.
Rigorous verification
- TDD turns the spec into executable checks. A guard blocks production edits until a failing test has run, in Claude Code and Kiro alike.
- One gate runs conventions, ruff, strict mypy, oxlint, Prettier, tsc and the test floors, and the pre-commit hook enforces it for every agent and person.
- It worked live: told to "skip the tests", Claude and Kiro both pointed to the guard and went test-first, and a probe confirmed the guard blocks a direct edit under Kiro. Both agents scored 8 out of 8 on following the rules.
- Verification splits by who can do it. Machines check WCAG 2.2 AA in a real browser and a named test covers each criterion axe can't see. The screen-reader check stays with a person, and the PR description records it.
Why SDD and TDD together
- SDD tells the agent what to build and keeps that context in the repo, which covers specification and context engineering.
- TDD turns the spec into checks a machine can run, which covers rigorous verification, the defence against agents that write convincing but wrong code.
- Neither is enough alone. A spec without tests is a hope. Tests without a spec lose the why, which is how Vibe and Plan ended up with an untested "existing data" case and a 500 on every request.
Verdict. Vibe is closest to AI-assisted coding around a single tool, and Plan is conventional engineering with an AI doing the typing. SPECTRA is AI-native: agents generate the code, while people write the specifications, own the context and verify the output through executable tests and automatic gates.