Lightweight Spec-Driven Development (Core Principles Shaped to an AI-Native Meta-Repo)
1. Executive Summary
Spec-driven development (SDD) puts a written, reviewed specification between intent and code, and has an AI agent implement against that document rather than against a conversation. It is commonly adopted as a framework: a toolkit with its own sequence of phases, commands, templates, and governing documents, designed for a default context that is most often a single repository with one developer driving one agent.
This architecture takes the other route. It keeps the ideas SDD rests on, leaves the packaging behind, and derives the process from the context it runs in: an AI-native team workspace (RA-003) that gives agents the whole system in one tree, features that routinely cross service boundaries, and an organisational tracker that stays. Lightweight here means framework-free, not ceremony-free. The process has structure — three tiers, separate stages, two change regimes — but no toolkit, and how much of that structure a change meets depends on the change.
The shape rests on one premise: nothing in the flow is written by hand, and nothing is accepted without its owner's review. Requirements, design, tasks, code, test scenarios and release notes are generated by AI agents; a person owns every stage, works through it with the agent, and answers for the result. This is AI-assisted engineering, not delegation. On that premise:
- Three tiers of process weight — a full specification, a one-page task, or a direct fix, proposed at intake and decided by a person. The tier follows where the engineering happens, not how far the change propagates.
- A specification is a folder of Markdown files in the workspace: requirements, design, tasks, a changelog, and an index that links out to the tracker. Agents read and write it, so its layout can drift without affecting the result.
- Each stage runs in its own session — requirements, design, implementation, testing, review. A full team or one engineer directing a team of agents can run the flow; what protects quality is the separation, not the headcount.
- Tasks are logical units of work, sized for a focused agent context and linked back to the requirement and design sections they implement.
- Two change regimes. Requirements and design are edited in place, and every change is logged. A completed task is never edited; rework arrives as a new task. Specifications stay alive after release, and defects route back to them.
- Delivery is recorded in the specification. Release preparation maps commit codes to specifications, so each one shows which of its parts have reached which environment.
- The tracker keeps intent. Epics stay where management and other teams read them and link to the specification; the decomposition lives only in the repository.
- Each stage gets the model it needs — the most capable one for design, task decomposition and review; a cheaper one for implementation, testing and documentation updates.
The process is encoded as rules and skills in the workspace, so it holds without any framework's tooling.
What it gives up: the community, documentation and upgrade path a framework brings; portability, since a shape fitted to one context has to be re-derived rather than copied; and independence from the workspace. Design and implementation depend on an agent that can read the whole codebase, and without that the process is considerably more complex to run.
Provenance. The author designed this process as team lead of a six-person team — the lead, two engineers, a PO/BA and two QA engineers — responsible for 80+ microservices on an enterprise platform of 200+. The team adopted it as its way of working, and it has been in use for five months. The largest specification so far spans 35 services and has absorbed four major updates. The linked repository is an example implementation of the files, rules and skills described here, not an extract of the production process.
2. Applicability Criteria
Use this pattern when
- The agent can read the whole system. A meta-repo (RA-003) — every repository of the system cloned into one tree, beside the team's documentation, rules, skills and tool integrations — or a monorepo gives the agent everything at once. Technical design, task links and implementation depend on it; requirements benefit from it (§4.3). This is the decisive precondition.
- The team works agent-first. People set direction, make decisions and review; agents write the artefacts. The process is sized for that division of labour: a specification costs guidance and review, not typing.
- Features routinely cross service boundaries. Specifications earn their cost there; work inside one service mostly lands on the lighter tiers.
- The organisation's tracker is fixed. Reporting, planning and cross-team dependencies run through it and will keep doing so.
- Most changes are small. A process that specifies every change gets abandoned; the tier decision is what keeps specifications affordable for the work that needs them.
- The available frameworks assume a different context — one repository, one developer, their own phase sequence — and fitting the team to a framework would cost more than fitting the ideas to the team.
Do NOT use this pattern when
- The agent is confined to one repository at a time. A cross-service design can't be validated in one session, task links point at files the agent cannot open, and implementation starts without the context it needs. The ideas still apply, but you have to assemble that context by hand at every step, which makes for a different and much heavier process (§7.2).
- A framework's defaults already match the context — one repository, new work, one developer driving one agent. Adopting the framework is cheaper than designing a process.
- Nobody can direct the technical layer. Design and decomposition cap the quality of everything downstream. The most capable model helps, but a person who knows the system still has to direct and review it.
- Requirements must live in a regulated system of record. Where traceability is audited in a requirements-management tool, a Markdown specification cannot be the source of truth.
3. Architecture Overview
3.1 What SDD is made of
Stripped of tooling, SDD comes down to six ideas. The toolkits implement them in different ways; none of the ideas requires a toolkit.
| # | Idea | What it prevents |
|---|---|---|
| 1 | Intent is written and agreed before code. | An agent implementing a conversation instead of a decision |
| 2 | What, how and the steps are separate artefacts — requirements, design, tasks — each reviewed on its own. | Three concerns reviewed as one document, where a confident design hides a wrong requirement |
| 3 | Tasks are small enough for a focused context and trace back to what they implement. | An agent working from an entire specification, drifting out of scope, producing changes nobody can review against intent |
| 4 | The specification is versioned in the repository the agent reads. | Intent kept in a wiki or a ticket the agent never sees, drifting away from the code |
| 5 | Acceptance is checked against the requirements, not against the implementation. | Tests that confirm what was built rather than what was asked for |
| 6 | When behaviour changes, the specification changes first, on the record. | Undocumented deviation — behaviour that exists only in code |
3.2 What a framework adds
A framework packages those ideas together with decisions made for its default context. Depending on the toolkit, the package includes:
- a fixed sequence of phases, each driven by its own command;
- the full artefact set for every feature, whatever its size;
- a project-wide governing document the phases defer to;
- its own folder layout, naming and templates;
- integrations for particular agents;
- an assumed context — typically one repository, with one developer driving the agent through the phases.
GitHub's Spec Kit, one of the open-source toolkits evaluated before this process was designed, is representative: a project constitution, then specify, plan, tasks and implement as successive commands in one repository.
Each of these decisions is reasonable where its assumption holds. Where it doesn't, the team adapts to the framework instead: ceremony on one-line changes, per-repository specifications for features that span many repositories, a second planning system beside the tracker the organisation already mandates. The design question is therefore not which framework, but which parts of the package the context actually needs.
3.3 The premise: agent-written, human-owned
The shape assumes an agent-first way of working. Every artefact in the flow is generated by an AI agent — from the requirements to the implementation, the test scenarios, the review findings and the release notes — and every stage has a person who owns it: works through it with the agent, corrects the draft until it is right, and is accountable for the result. Nothing is typed by hand; nothing is accepted unread.
| Stage | The agent produces | The owner |
|---|---|---|
| Requirements | A draft from the Epic, the business documentation and the owner's direction | The PO/BA corrects and completes it until it states what the business wants |
| Design and tasks | Architecture, contracts, the task breakdown and the task files | The lead validates the approach and signs off the decomposition |
| Implementation | Code, unit tests, documentation updates | The developer reviews the change before it goes into a merge request |
| Review | Findings against the task and the specification | The reviewer decides what is raised and whether the change is approved |
| Testing | Scenarios from the acceptance criteria; automation against the contracts | QA verifies the behaviour on the test environments |
That ownership also catches a change that matches the specification but is still wrong. The PO/BA answers whether the requirements are right; QA checks behaviour against them independently of the implementation; and people verify the result in test environments before release.
Three consequences run through the rest of the architecture:
- A specification costs guidance and review, not writing. The tiers exist to ration people's attention, not text.
- The files are written for agents to read. Links, anchors and explicit scope matter more than a uniform layout. Templates guide rather than constrain: nothing parses the files, so the layout can drift from one specification to the next without affecting the result — provided the content stays where the agents look for it, in the sections and anchors the tasks link to.
- Stages can be separated cheaply — and should be (§4.3).
3.4 Constraints decide the shape
Every part of the process below traces back to one constraint of the context it was built for:
| Constraint | Decision | Where |
|---|---|---|
| Most changes are small | Three tiers of process weight; the tier follows where the engineering happens | §4.1 |
| Features cross many services | The specification lives in the workspace, not in any service repository | §4.2 |
| A stage should check the previous one, not inherit its assumptions | Each stage runs in its own session with its own owner; one engineer can hold every role | §4.3 |
| Agents work best on a bounded context | Tasks are logical units sized for a focused context, linked back to the specification | §4.4 |
| Implementation and test automation must not wait on each other | Contracts are fixed in the design | §4.5 |
| Requirements move, including after release | Long-lived specifications; requirements and design edited in place with a log; completed tasks immutable; defects routed back | §4.6 |
| Several engineers share one workspace repository | A branch per engineer, regular rebases, merge requests; one implementer per specification | §4.7 |
| A feature ships service by service | Release preparation records each part's delivery in the specification | §4.8 |
| The organisation's tracker is mandatory | The tracker keeps intent; the repository keeps the decomposition | §4.9 |
| Stages differ in how much reasoning they need | Each stage is routed to the model it needs | §5 |
| No framework tooling holds the steps in place | The process is encoded as rules and skills in the workspace | §6 |
The table is the transferable part. A team with a different context — one repository, no PO, a tracker it controls — keeps the six ideas of §3.1 and ends up with different rows.
3.5 The flow
Each box in the workspace is a separate agent session owned by a person. The design is validated against real code before anyone implements it, and the review runs apart from the implementation it checks. The dotted paths are asynchronous, and they differ most from a feature-by-feature workflow: status flows back to the tracker the organisation reads, and change after release flows back into the same specification instead of starting a new one.
4. Implementation Details
The example repository implements each part below as files, rules and skills, linked from each section. Its layout and naming are one possible rendering; what carries over is each file's role.
Roles used below: the PO/BA; the lead — the team lead, or the most experienced engineer available for the feature; developers; and QA. A role is the owner of a stage, not a headcount: one engineer directing separate agent sessions can hold all of them.
4.1 Three tiers, decided per unit of work
| Tier | For | Artefact |
|---|---|---|
| Full specification | Multi-service features, complex business logic, UI flows, background event pipelines | A specification folder (§4.2) |
| Standalone task | A change inside one service or library; a dependency update | One page: context, minimal requirements, design notes, a checklist, verification steps |
| Direct fix | A minor bug with an obvious cause, a missing UI detail, a routine configuration change | None — implemented directly with the agent |
The tier follows where the engineering happens, not how far the change propagates. A change to a shared library that eighteen services then pick up through version bumps is engineered in one repository, with one design decision and one verification plan. In production, such a change ran as a standalone task and took about a day for the library and all eighteen services. Scored by reach instead, it would have become a full specification for one design decision and a mechanical rollout.
The tier is proposed at intake and decided by a person. The example repository writes the proposal as a rubric that an intake skill applies; the rubric recommends a direction, and the developer's choice wins, even if it goes against it. A standalone task that starts growing its own requirements and design sections is the signal that it should have been a specification: switch tiers rather than let a one-pager turn into a poor specification.
Examples: triage rubric · one-page task template · the task tier
4.2 A specification is a folder of Markdown files
A specification is a directory of plain Markdown in the workspace, one per feature. It does not live in a service repository, because the feature spans several. It also doesn't live in any agent's configuration directory, because the specification has to outlive the tool.
| File | Holds | Owner (guides the agent) |
|---|---|---|
README.md | Owner, status, target release, delivery notes per environment (§4.8); links to the Epic, the merge requests and the other files | The PO/BA at creation; kept current by the feature owner |
CHANGELOG.md | Every change after the first version: what changed, when, and why | Whoever changes the specification |
requirements.md | Functional requirements, business rules, UX expectations, testable acceptance criteria | The PO/BA |
design.md | Architecture, service interactions, data and event contracts; UI mockups or references | The lead; the UI part may be started by the PO/BA |
| Task index | The task list, the dependencies between tasks, ordering rules | The lead |
| Task files | One per logical unit of work (§4.4) | The lead; executed in a developer's agent session |
| Testing files | Test scenarios, automation scope, test results and the defects they found | QA |
An agent generates every file; its owner directs it, reviews the result, and approves it. The file names follow the requirements → design → tasks convention that agent spec tooling commonly uses, so the files stay usable by an agent's built-in spec features without depending on them. The template is a guide rather than a schema: agents read the files, nothing parses them, and a specification whose layout drifts from the template still works, as long as the sections the tasks link to are where the links say.
Examples: specification template · layout and naming
4.3 Each stage runs in its own session
Each stage — requirements, design and decomposition, implementation, testing, review — runs as a separate agent session with its own owner, and each hand-off is a commit. The owners can be different people or one engineer directing a team of agents; the flow is the same either way.
The separation is there for quality. When one model plans, implements and tests a change in a single session, each step inherits the assumptions of the one before it, and an early mistake gets confirmed rather than caught. A new session starts from the files alone — the specification, not the conversation that produced it — so each stage checks the previous one instead of continuing it. That matches the team's experience, and published results point the same way: models are poor at correcting their own reasoning without external feedback, and rate their own output above equivalent output from others (References).
Requirements. The PO/BA starts a specification from the Epic and works with the agent from the product's business documentation. The agent drafts; the PO/BA reviews and corrects the draft until it is right, and is responsible for it. Requirements describe the desired state of the product, independent of the technology underneath, so they don't need the codebase. Access to the code is still an advantage: obstacles show up while the goal can still change — offering an existing component instead of building a missing one is a common adjustment — and the effort of the whole flow can be estimated at the first stage.
Design. The specification then passes to the lead, who is not necessarily the person who will implement it. With the agent, the lead validates the approach against the service estate and its domain boundaries, then directs the writing of design.md — schemas, payloads, event models — and of the task breakdown. This layer caps the quality of everything downstream, which is why it prioritises experience over availability.
Implementation, testing, review. Each task is implemented in its own session (§4.4), together with the documentation it affects; test scenarios and automation are produced in their own (§4.5); and the review of each change is a separate session again, checked against the task and the specification rather than against the diff alone.
The workspace rules that every session receives carry the project-wide conventions, so each stage's review can concentrate on the feature itself.
Examples: specification phase · author-spec skill · review-mr skill
4.4 Tasks are logical units of work
The design is broken down into tasks. A task is one coherent unit of work that an agent can carry out with a focused context, and it links to the exact sections of requirements.md and design.md it implements.
Decomposition is design work at the scale of an agent's context — sizing the tasks, isolating the dependencies between them, fixing the contracts at their boundaries — and it runs on the same model as the design (§5). What the work is decides the split, not how many repositories it touches. Substantive changes in several services usually become one task per service, because each needs its own context and its own review. A mechanical change repeated across services does not: new logic in a shared library is one task, and bumping that library's version in every service that consumes it is a second — not one task per service.
Three properties follow:
- The agent's context is bounded by the task. It works from one task file and the sections it links to, not from the whole specification, so what it reads matches what it is allowed to change.
- Tasks can run in parallel — one developer working through them in order, or several developers at once. The index records which tasks depend on which.
- Review maps back to the specification. The reviewing session checks each change against its task and the sections the task links to, which is a narrower review than reading a diff to reconstruct its intent.
How much a task spells out depends on what will implement it: the less capable the model, the more of the design the task has to carry (§5).
Examples: task index template · task file template · implement-task skill
4.5 Contracts are fixed in the design
Where a feature has UI or crosses services, design.md fixes the contracts before code is written: request, response and event payloads, and the test identifiers of UI elements. With the contracts fixed, test automation proceeds in parallel with implementation, or ahead of it, and is generated from the requirements rather than from the code that satisfies them. QA's test scenarios come from the acceptance criteria, in a session of their own.
Example: shared test-identifier contract
4.6 Two change regimes
Requirements move: during implementation, when a planned approach turns out not to be available, and after release, when real users meet the feature. The process treats the two halves of a specification differently.
- Requirements and design describe what is true now. They are edited in place, with the change marked where it was made.
CHANGELOG.mdrecords the previous state, the new state and the reason. - Task files describe what was done. A task that has not started is edited along with the design, and the edit is logged. A completed task is never edited: rework arrives as a new delta task that says what changes and why.
The result is a specification that still reads as a description of current behaviour while keeping the full record of how it got there. Specifications are not archived at release. They remain the document someone reads months later to learn why a feature behaves the way it does.
Defects enter through the same two regimes, and they stay traceable to the specification they belong to. A defect found in testing is usually recorded in the specification's own testing files, and sometimes arrives as a comment on the Epic or a message to the PO/BA or the developer. The developer resolves it either as a specification change with new tasks or — when nothing documented changes — as an in-place fix with a note in the specification's metadata. A production defect goes through intake like any other work, usually lands on the standalone-task tier, and links to the specification of the feature it affects. The path from a defect back to the section that owns the behaviour runs through existing links.
Two production cases show the regimes under load. The largest specification spans 35 services; it has been through four major updates and now holds 13 tasks — the original breakdown plus the delta tasks those updates added. On another feature, users started reporting problems on the first working day after release. The team decided to roll back, and within about three hours a backend hotfix was in production, and the affected users' data had been rolled back in the database. The same specification then recorded the change: an edit to the requirements, delta tasks, and a changelog entry.
Example: changelog template
4.7 One implementer per specification
The workspace is a repository, and we work like one. Each engineer works in a branch of their own, rebases onto the main branch regularly — an editor task automates it — and merges back through merge requests, so specification changes become visible to everyone as they land.
Conflicts on specification files are rare in practice, because a specification is normally implemented by one developer. In five months, two developers implemented one specification in parallel to shorten delivery, and it produced no conflicts. A feature that needs several developers working on it at once is usually too large: splitting it into smaller specifications, each easier to implement, test and release, removes the problem at its source.
4.8 Delivery is recorded in the specification
Features ship service by service, so a specification or a standalone task can be partly released: three of its services in production, a fourth still in a test environment. The workspace records that in the artefact itself, at almost no extra cost.
Every commit carries the code of the task or ticket it belongs to — a convention this team followed for years before agents arrived, and one any team working together needs anyway. The release-preparation procedure, run for each promotion through the test environments to production, reads the commits going out per service, maps their codes to specifications and standalone tasks, and adds a delivery note to each: which service, which environment, which date. A commit in service A carrying TW-074 is enough for the specification behind TW-074 to record that its service-A part reached production on that day.
Two things become visible without anyone maintaining them: which parts of a feature are already in production and which aren't, and whether a part is missing from a release — a service whose changes for the feature haven't shipped yet.
Examples: task code in every commit · progress and release rows in the specification README
4.9 The tracker keeps intent
The organisation's tracker remains the system that management, planning and other teams read. Epics and high-level tickets stay there — business milestones, cross-team dependencies, customer requests — and each Epic links to its specification. The decomposition does not: tasks, contracts and change history live only in the workspace. Mirroring them as tickets would create a second copy of the plan that drifts from the first.
The agent bridges the two. With tracker tools available to it over MCP (RA-003), it can read an Epic when a specification starts and report status and progress back, instead of someone keeping tickets in step with the Markdown by hand.
Changes that arrive through the tracker take the same path. When the customer adds adjustments to an Epic as comments and asks the PO/BA to consider them, the PO/BA has the agent read the comments and assess how they affect the existing specification; what is accepted enters the specification as an amendment (§4.6). The Epic records what was asked, and the specification records what was decided.
Example: what the Epic owns and what the specification owns
5. Model Routing by Stage
Separate sessions (§4.3) have a second consequence: each stage can run on a different model. Stages differ in how much reasoning they need, and routing them accordingly controls the cost of running the process.
| Stage | Model | Why |
|---|---|---|
| Requirements | Below the most capable tier | Describes the desired state of the product from business documentation; the owner's review makes it correct |
| Technical design | The most capable available, at its highest reasoning setting | Design caps the quality of everything downstream; cross-service reasoning is where capability pays off |
| Task decomposition and task files | The most capable available | Decomposition is design at the scale of an agent's context: sizing tasks, isolating dependencies, fixing the contracts between them |
| Implementation | Cheaper | The task carries the decisions; the model carries out the task |
| Testing | Cheaper | Scenarios and automation follow from fixed acceptance criteria and contracts |
| Review | The most capable available | Review is where mistakes from the cheaper stages are caught |
| Documentation updates | Cheaper | They describe a change that is already made and reviewed |
One rule keeps the cheaper stages safe: the less capable the model that implements a task, the more detailed the task has to be. Task detail and model capability trade against each other, and the balance is set when the tasks are written.
At the time of writing, the team routes design to Claude Opus at its maximum reasoning effort, task files to Opus one effort level lower, review to Opus, and implementation, testing and documentation updates to Claude Sonnet.
The savings come from the high-volume stages — implementation, testing, documentation — running on a cheaper model, while the most capable model is spent where it decides the outcome: design, task decomposition and review.
6. Encoding the Process in the Workspace
A framework's commands are what keep its steps from being skipped. Without a framework, the workspace carries that job, in two kinds of files. Each is written once as a canonical file and reaches every agent through thin per-agent files — the wrapper mechanism RA-003 describes for skills, applied to rules as well.
- Rules (steering) — short instructions delivered to every agent session: how the tier is chosen, how requirements and acceptance criteria are written, the commit, git and delivery conventions, which environments and repositories belong to the team. A rule earns its place by being something an agent would otherwise get wrong, because everything delivered on every session costs context on every session.
- Skills — step-by-step procedures for the jobs that have to come out the same way every time:
- intake and triage of new work into a tier;
- writing or amending a specification stage by stage, pausing for review after each artefact;
- executing one task end to end — branch, implementation, quality gates, merge request, progress recorded in the specification;
- reviewing a merge request against the specification and the task it implements;
- preparing a release, including the delivery notes it writes into specifications and tasks (§4.8).
The skills stop wherever a person decides: the tier, each artefact of the specification, a commit, a review comment before it is raised. They make the process repeatable; they don't make it autonomous.
Like any process, this one holds only as long as the team follows it. The skills make following it the easy path rather than the only one, which puts the rest on team discipline — as it would with a framework. In this team, it has held for five months, including under pre-release pressure. A team that wants a structural backstop can add cheap CI checks, such as rejecting a merge request whose commits carry no task code, without changing anything else.
Examples: rules · triage rule · requirements rule · skills: plan-task, author-spec, implement-task, review-mr
7. Scalability & Evolution Path
7.1 Refined in use
One engineer designed the process; the people who run it have refined it since. An engineer who hits a gap — a missing step, a rule that no longer fits, a template section nobody fills in — changes the rule, the skill or the template, and the change is reviewed like code. Because the process is a set of files in the workspace, improving it costs the same as any other small change, and every agent picks up the new version in its next session.
7.2 Without a whole-system workspace
The shape assumes the agent sees everything. Where it doesn't — each repository opened on its own — requirements lose only their early feasibility check, since they come from business documentation. Everything after them gets heavier: the design can't be validated across services in one session, the specification needs a home outside every service repository, and each task has to carry the context its links would otherwise provide. The ideas still hold, but the process is much heavier to run. In that situation, the first investment is usually the workspace (RA-003), not more process.
7.3 Keeping it lightweight
A home-grown process can grow into the heavy framework it replaced, one reasonable rule at a time. The tier decision is the guard: most work never reaches a specification, so ceremony added to specifications stays confined to the work that justified it. Remove a rule or template section that stops earning its cost the same way you added it.
8. Trade-off Analysis
| Feature | Benefit | Drawback |
|---|---|---|
| No framework | The process fits the team's context; changing it is an ordinary edit | No community, documentation or upgrade path beyond what the team writes |
| Agent-written, human-owned files | A specification costs guidance and review, not writing; the layout can drift without affecting the result | Less uniform for a person reading across many specifications; each stage is only as good as its owner's review |
| Tiers | Small changes carry no ceremony; the rubric proposes, and a person decides | The tier is a judgement; a misjudged task grows into a poor specification unless someone switches tiers |
| Separate sessions per stage | Each stage checks the previous one instead of inheriting its assumptions; one engineer can run the whole flow | More hand-offs; everything a later stage needs has to be in the files |
| Logical-unit tasks | Bounded agent context, parallel implementation, review against intent | Sizing is a judgement; a large feature still produces many tasks and an ordering to manage |
| Two change regimes | Current behaviour and full history in one place, after release as well as before; defects trace back to the owning section | Specifications grow; they stay navigable only while the changelog and the task index are kept up |
| One implementer per specification | No conflicts on specification files in practice | Parallel work on one specification is the exception; a feature that needs it is split first |
| Delivery recorded in the specification | What is in production, and what is missing, is visible per feature with no manual tracking | Depends on a commit convention the whole team follows |
| Tracker at intent level | No second copy of the plan in tickets; management keeps its view | Two systems; status crosses between them through the agent, not by construction |
| Model routing | The most capable model is spent on design, task decomposition and review | Cheaper implementation needs more detailed tasks |
| Rules and skills | The steps hold across agents without framework tooling | Enforcement is procedural — skills, review and team discipline — not structural |
| Whole-system workspace | Design is validated against real code; requirements get an early feasibility check and a full-flow estimate | A hard dependency for design and implementation; without it the process is much heavier (§7.2) |
| Fitted to one context | Every rule exists for a reason the team can name | Another team re-derives the shape rather than copying it |
9. Conclusion
Spec-driven development is a small set of ideas: agree intent before code; keep what, how and the steps apart; give agents small tasks that trace back to the specification; keep the specification beside the code; verify against what was asked; and change the specification when behaviour changes. A framework is one way to package those ideas, built for its own context.
When that context doesn't match, you can apply the ideas directly and derive the process from the team's own constraints — with agents writing every artefact and a person owning every stage. For one team responsible for 80+ microservices on an AI-native workspace, the process derived that way has run for five months, carried a 35-service feature through four major updates, and handled a post-release rollback within hours.
References
- Example implementation: github.com/ivanbaha/team-workspace — lifecycle overview · rationale and limits · specification template · one-page task template
- Prerequisite architecture: RA-003 · AI-Native Team Workspace
- Field report: Spec-Driven Development Is Simpler Than You Thought
- An SDD toolkit, for comparison: GitHub Spec Kit
- On separating stages: Huang et al., Large Language Models Cannot Self-Correct Reasoning Yet, ICLR 2024 · Panickssery, Bowman, Feng, LLM Evaluators Recognize and Favor Their Own Generations, 2024