Skip to main content

Lightweight Spec-Driven Development (Core Principles Shaped to an AI-Native Meta-Repo)

Metadata

1. Executive Summary​

Spec-driven development (SDD) puts a written, reviewed specification between intent and code, and has an AI agent implement against that document rather than against a conversation. It is commonly adopted as a framework: a toolkit with its own sequence of phases, commands, templates, and governing documents, designed for a default context that is most often a single repository with one developer driving one agent.

This architecture takes the other route. It keeps the ideas SDD rests on, leaves the packaging behind, and derives the process from the context it runs in: an AI-native team workspace (RA-003) that gives agents the whole system in one tree, features that routinely cross service boundaries, and an organisational tracker that stays. Lightweight here means framework-free, not ceremony-free. The process has structure — three tiers, separate stages, two change regimes — but no toolkit, and how much of that structure a change meets depends on the change.

The shape rests on one premise: nothing in the flow is written by hand, and nothing is accepted without its owner's review. Requirements, design, tasks, code, test scenarios and release notes are generated by AI agents; a person owns every stage, works through it with the agent, and answers for the result. This is AI-assisted engineering, not delegation. On that premise:

  • Three tiers of process weight — a full specification, a one-page task, or a direct fix, proposed at intake and decided by a person. The tier follows where the engineering happens, not how far the change propagates.
  • A specification is a folder of Markdown files in the workspace: requirements, design, tasks, a changelog, and an index that links out to the tracker. Agents read and write it, so its layout can drift without affecting the result.
  • Each stage runs in its own session — requirements, design, implementation, testing, review. A full team or one engineer directing a team of agents can run the flow; what protects quality is the separation, not the headcount.
  • Tasks are logical units of work, sized for a focused agent context and linked back to the requirement and design sections they implement.
  • Two change regimes. Requirements and design are edited in place, and every change is logged. A completed task is never edited; rework arrives as a new task. Specifications stay alive after release, and defects route back to them.
  • Delivery is recorded in the specification. Release preparation maps commit codes to specifications, so each one shows which of its parts have reached which environment.
  • The tracker keeps intent. Epics stay where management and other teams read them and link to the specification; the decomposition lives only in the repository.
  • Each stage gets the model it needs — the most capable one for design, task decomposition and review; a cheaper one for implementation, testing and documentation updates.

The process is encoded as rules and skills in the workspace, so it holds without any framework's tooling.

What it gives up: the community, documentation and upgrade path a framework brings; portability, since a shape fitted to one context has to be re-derived rather than copied; and independence from the workspace. Design and implementation depend on an agent that can read the whole codebase, and without that the process is considerably more complex to run.

Provenance. The author designed this process as team lead of a six-person team — the lead, two engineers, a PO/BA and two QA engineers — responsible for 80+ microservices on an enterprise platform of 200+. The team adopted it as its way of working, and it has been in use for five months. The largest specification so far spans 35 services and has absorbed four major updates. The linked repository is an example implementation of the files, rules and skills described here, not an extract of the production process.

2. Applicability Criteria​

Use this pattern when​

  • The agent can read the whole system. A meta-repo (RA-003) — every repository of the system cloned into one tree, beside the team's documentation, rules, skills and tool integrations — or a monorepo gives the agent everything at once. Technical design, task links and implementation depend on it; requirements benefit from it (§4.3). This is the decisive precondition.
  • The team works agent-first. People set direction, make decisions and review; agents write the artefacts. The process is sized for that division of labour: a specification costs guidance and review, not typing.
  • Features routinely cross service boundaries. Specifications earn their cost there; work inside one service mostly lands on the lighter tiers.
  • The organisation's tracker is fixed. Reporting, planning and cross-team dependencies run through it and will keep doing so.
  • Most changes are small. A process that specifies every change gets abandoned; the tier decision is what keeps specifications affordable for the work that needs them.
  • The available frameworks assume a different context — one repository, one developer, their own phase sequence — and fitting the team to a framework would cost more than fitting the ideas to the team.

Do NOT use this pattern when​

  • The agent is confined to one repository at a time. A cross-service design can't be validated in one session, task links point at files the agent cannot open, and implementation starts without the context it needs. The ideas still apply, but you have to assemble that context by hand at every step, which makes for a different and much heavier process (§7.2).
  • A framework's defaults already match the context — one repository, new work, one developer driving one agent. Adopting the framework is cheaper than designing a process.
  • Nobody can direct the technical layer. Design and decomposition cap the quality of everything downstream. The most capable model helps, but a person who knows the system still has to direct and review it.
  • Requirements must live in a regulated system of record. Where traceability is audited in a requirements-management tool, a Markdown specification cannot be the source of truth.

3. Architecture Overview​

3.1 What SDD is made of​

Stripped of tooling, SDD comes down to six ideas. The toolkits implement them in different ways; none of the ideas requires a toolkit.

#IdeaWhat it prevents
1Intent is written and agreed before code.An agent implementing a conversation instead of a decision
2What, how and the steps are separate artefacts — requirements, design, tasks — each reviewed on its own.Three concerns reviewed as one document, where a confident design hides a wrong requirement
3Tasks are small enough for a focused context and trace back to what they implement.An agent working from an entire specification, drifting out of scope, producing changes nobody can review against intent
4The specification is versioned in the repository the agent reads.Intent kept in a wiki or a ticket the agent never sees, drifting away from the code
5Acceptance is checked against the requirements, not against the implementation.Tests that confirm what was built rather than what was asked for
6When behaviour changes, the specification changes first, on the record.Undocumented deviation — behaviour that exists only in code

3.2 What a framework adds​

A framework packages those ideas together with decisions made for its default context. Depending on the toolkit, the package includes:

  • a fixed sequence of phases, each driven by its own command;
  • the full artefact set for every feature, whatever its size;
  • a project-wide governing document the phases defer to;
  • its own folder layout, naming and templates;
  • integrations for particular agents;
  • an assumed context — typically one repository, with one developer driving the agent through the phases.

GitHub's Spec Kit, one of the open-source toolkits evaluated before this process was designed, is representative: a project constitution, then specify, plan, tasks and implement as successive commands in one repository.

Each of these decisions is reasonable where its assumption holds. Where it doesn't, the team adapts to the framework instead: ceremony on one-line changes, per-repository specifications for features that span many repositories, a second planning system beside the tracker the organisation already mandates. The design question is therefore not which framework, but which parts of the package the context actually needs.

3.3 The premise: agent-written, human-owned​

The shape assumes an agent-first way of working. Every artefact in the flow is generated by an AI agent — from the requirements to the implementation, the test scenarios, the review findings and the release notes — and every stage has a person who owns it: works through it with the agent, corrects the draft until it is right, and is accountable for the result. Nothing is typed by hand; nothing is accepted unread.

StageThe agent producesThe owner
RequirementsA draft from the Epic, the business documentation and the owner's directionThe PO/BA corrects and completes it until it states what the business wants
Design and tasksArchitecture, contracts, the task breakdown and the task filesThe lead validates the approach and signs off the decomposition
ImplementationCode, unit tests, documentation updatesThe developer reviews the change before it goes into a merge request
ReviewFindings against the task and the specificationThe reviewer decides what is raised and whether the change is approved
TestingScenarios from the acceptance criteria; automation against the contractsQA verifies the behaviour on the test environments

That ownership also catches a change that matches the specification but is still wrong. The PO/BA answers whether the requirements are right; QA checks behaviour against them independently of the implementation; and people verify the result in test environments before release.

Three consequences run through the rest of the architecture:

  • A specification costs guidance and review, not writing. The tiers exist to ration people's attention, not text.
  • The files are written for agents to read. Links, anchors and explicit scope matter more than a uniform layout. Templates guide rather than constrain: nothing parses the files, so the layout can drift from one specification to the next without affecting the result — provided the content stays where the agents look for it, in the sections and anchors the tasks link to.
  • Stages can be separated cheaply — and should be (§4.3).

3.4 Constraints decide the shape​

Every part of the process below traces back to one constraint of the context it was built for:

ConstraintDecisionWhere
Most changes are smallThree tiers of process weight; the tier follows where the engineering happens§4.1
Features cross many servicesThe specification lives in the workspace, not in any service repository§4.2
A stage should check the previous one, not inherit its assumptionsEach stage runs in its own session with its own owner; one engineer can hold every role§4.3
Agents work best on a bounded contextTasks are logical units sized for a focused context, linked back to the specification§4.4
Implementation and test automation must not wait on each otherContracts are fixed in the design§4.5
Requirements move, including after releaseLong-lived specifications; requirements and design edited in place with a log; completed tasks immutable; defects routed back§4.6
Several engineers share one workspace repositoryA branch per engineer, regular rebases, merge requests; one implementer per specification§4.7
A feature ships service by serviceRelease preparation records each part's delivery in the specification§4.8
The organisation's tracker is mandatoryThe tracker keeps intent; the repository keeps the decomposition§4.9
Stages differ in how much reasoning they needEach stage is routed to the model it needs§5
No framework tooling holds the steps in placeThe process is encoded as rules and skills in the workspace§6

The table is the transferable part. A team with a different context — one repository, no PO, a tracker it controls — keeps the six ideas of §3.1 and ends up with different rows.

3.5 The flow​

Each box in the workspace is a separate agent session owned by a person. The design is validated against real code before anyone implements it, and the review runs apart from the implementation it checks. The dotted paths are asynchronous, and they differ most from a feature-by-feature workflow: status flows back to the tracker the organisation reads, and change after release flows back into the same specification instead of starting a new one.

4. Implementation Details​

The example repository implements each part below as files, rules and skills, linked from each section. Its layout and naming are one possible rendering; what carries over is each file's role.

Roles used below: the PO/BA; the lead — the team lead, or the most experienced engineer available for the feature; developers; and QA. A role is the owner of a stage, not a headcount: one engineer directing separate agent sessions can hold all of them.

4.1 Three tiers, decided per unit of work​

TierForArtefact
Full specificationMulti-service features, complex business logic, UI flows, background event pipelinesA specification folder (§4.2)
Standalone taskA change inside one service or library; a dependency updateOne page: context, minimal requirements, design notes, a checklist, verification steps
Direct fixA minor bug with an obvious cause, a missing UI detail, a routine configuration changeNone — implemented directly with the agent

The tier follows where the engineering happens, not how far the change propagates. A change to a shared library that eighteen services then pick up through version bumps is engineered in one repository, with one design decision and one verification plan. In production, such a change ran as a standalone task and took about a day for the library and all eighteen services. Scored by reach instead, it would have become a full specification for one design decision and a mechanical rollout.

The tier is proposed at intake and decided by a person. The example repository writes the proposal as a rubric that an intake skill applies; the rubric recommends a direction, and the developer's choice wins, even if it goes against it. A standalone task that starts growing its own requirements and design sections is the signal that it should have been a specification: switch tiers rather than let a one-pager turn into a poor specification.

Examples: triage rubric · one-page task template · the task tier

4.2 A specification is a folder of Markdown files​

A specification is a directory of plain Markdown in the workspace, one per feature. It does not live in a service repository, because the feature spans several. It also doesn't live in any agent's configuration directory, because the specification has to outlive the tool.

FileHoldsOwner (guides the agent)
README.mdOwner, status, target release, delivery notes per environment (§4.8); links to the Epic, the merge requests and the other filesThe PO/BA at creation; kept current by the feature owner
CHANGELOG.mdEvery change after the first version: what changed, when, and whyWhoever changes the specification
requirements.mdFunctional requirements, business rules, UX expectations, testable acceptance criteriaThe PO/BA
design.mdArchitecture, service interactions, data and event contracts; UI mockups or referencesThe lead; the UI part may be started by the PO/BA
Task indexThe task list, the dependencies between tasks, ordering rulesThe lead
Task filesOne per logical unit of work (§4.4)The lead; executed in a developer's agent session
Testing filesTest scenarios, automation scope, test results and the defects they foundQA

An agent generates every file; its owner directs it, reviews the result, and approves it. The file names follow the requirements → design → tasks convention that agent spec tooling commonly uses, so the files stay usable by an agent's built-in spec features without depending on them. The template is a guide rather than a schema: agents read the files, nothing parses them, and a specification whose layout drifts from the template still works, as long as the sections the tasks link to are where the links say.

Examples: specification template · layout and naming

4.3 Each stage runs in its own session​

Each stage — requirements, design and decomposition, implementation, testing, review — runs as a separate agent session with its own owner, and each hand-off is a commit. The owners can be different people or one engineer directing a team of agents; the flow is the same either way.

The separation is there for quality. When one model plans, implements and tests a change in a single session, each step inherits the assumptions of the one before it, and an early mistake gets confirmed rather than caught. A new session starts from the files alone — the specification, not the conversation that produced it — so each stage checks the previous one instead of continuing it. That matches the team's experience, and published results point the same way: models are poor at correcting their own reasoning without external feedback, and rate their own output above equivalent output from others (References).

Requirements. The PO/BA starts a specification from the Epic and works with the agent from the product's business documentation. The agent drafts; the PO/BA reviews and corrects the draft until it is right, and is responsible for it. Requirements describe the desired state of the product, independent of the technology underneath, so they don't need the codebase. Access to the code is still an advantage: obstacles show up while the goal can still change — offering an existing component instead of building a missing one is a common adjustment — and the effort of the whole flow can be estimated at the first stage.

Design. The specification then passes to the lead, who is not necessarily the person who will implement it. With the agent, the lead validates the approach against the service estate and its domain boundaries, then directs the writing of design.md — schemas, payloads, event models — and of the task breakdown. This layer caps the quality of everything downstream, which is why it prioritises experience over availability.

Implementation, testing, review. Each task is implemented in its own session (§4.4), together with the documentation it affects; test scenarios and automation are produced in their own (§4.5); and the review of each change is a separate session again, checked against the task and the specification rather than against the diff alone.

The workspace rules that every session receives carry the project-wide conventions, so each stage's review can concentrate on the feature itself.

Examples: specification phase · author-spec skill · review-mr skill

4.4 Tasks are logical units of work​

The design is broken down into tasks. A task is one coherent unit of work that an agent can carry out with a focused context, and it links to the exact sections of requirements.md and design.md it implements.

Decomposition is design work at the scale of an agent's context — sizing the tasks, isolating the dependencies between them, fixing the contracts at their boundaries — and it runs on the same model as the design (§5). What the work is decides the split, not how many repositories it touches. Substantive changes in several services usually become one task per service, because each needs its own context and its own review. A mechanical change repeated across services does not: new logic in a shared library is one task, and bumping that library's version in every service that consumes it is a second — not one task per service.

Three properties follow:

  • The agent's context is bounded by the task. It works from one task file and the sections it links to, not from the whole specification, so what it reads matches what it is allowed to change.
  • Tasks can run in parallel — one developer working through them in order, or several developers at once. The index records which tasks depend on which.
  • Review maps back to the specification. The reviewing session checks each change against its task and the sections the task links to, which is a narrower review than reading a diff to reconstruct its intent.

How much a task spells out depends on what will implement it: the less capable the model, the more of the design the task has to carry (§5).

Examples: task index template · task file template · implement-task skill

4.5 Contracts are fixed in the design​

Where a feature has UI or crosses services, design.md fixes the contracts before code is written: request, response and event payloads, and the test identifiers of UI elements. With the contracts fixed, test automation proceeds in parallel with implementation, or ahead of it, and is generated from the requirements rather than from the code that satisfies them. QA's test scenarios come from the acceptance criteria, in a session of their own.

Example: shared test-identifier contract

4.6 Two change regimes​

Requirements move: during implementation, when a planned approach turns out not to be available, and after release, when real users meet the feature. The process treats the two halves of a specification differently.

  • Requirements and design describe what is true now. They are edited in place, with the change marked where it was made. CHANGELOG.md records the previous state, the new state and the reason.
  • Task files describe what was done. A task that has not started is edited along with the design, and the edit is logged. A completed task is never edited: rework arrives as a new delta task that says what changes and why.

The result is a specification that still reads as a description of current behaviour while keeping the full record of how it got there. Specifications are not archived at release. They remain the document someone reads months later to learn why a feature behaves the way it does.

Defects enter through the same two regimes, and they stay traceable to the specification they belong to. A defect found in testing is usually recorded in the specification's own testing files, and sometimes arrives as a comment on the Epic or a message to the PO/BA or the developer. The developer resolves it either as a specification change with new tasks or — when nothing documented changes — as an in-place fix with a note in the specification's metadata. A production defect goes through intake like any other work, usually lands on the standalone-task tier, and links to the specification of the feature it affects. The path from a defect back to the section that owns the behaviour runs through existing links.

Two production cases show the regimes under load. The largest specification spans 35 services; it has been through four major updates and now holds 13 tasks — the original breakdown plus the delta tasks those updates added. On another feature, users started reporting problems on the first working day after release. The team decided to roll back, and within about three hours a backend hotfix was in production, and the affected users' data had been rolled back in the database. The same specification then recorded the change: an edit to the requirements, delta tasks, and a changelog entry.

Example: changelog template

4.7 One implementer per specification​

The workspace is a repository, and we work like one. Each engineer works in a branch of their own, rebases onto the main branch regularly — an editor task automates it — and merges back through merge requests, so specification changes become visible to everyone as they land.

Conflicts on specification files are rare in practice, because a specification is normally implemented by one developer. In five months, two developers implemented one specification in parallel to shorten delivery, and it produced no conflicts. A feature that needs several developers working on it at once is usually too large: splitting it into smaller specifications, each easier to implement, test and release, removes the problem at its source.

4.8 Delivery is recorded in the specification​

Features ship service by service, so a specification or a standalone task can be partly released: three of its services in production, a fourth still in a test environment. The workspace records that in the artefact itself, at almost no extra cost.

Every commit carries the code of the task or ticket it belongs to — a convention this team followed for years before agents arrived, and one any team working together needs anyway. The release-preparation procedure, run for each promotion through the test environments to production, reads the commits going out per service, maps their codes to specifications and standalone tasks, and adds a delivery note to each: which service, which environment, which date. A commit in service A carrying TW-074 is enough for the specification behind TW-074 to record that its service-A part reached production on that day.

Two things become visible without anyone maintaining them: which parts of a feature are already in production and which aren't, and whether a part is missing from a release — a service whose changes for the feature haven't shipped yet.

Examples: task code in every commit · progress and release rows in the specification README

4.9 The tracker keeps intent​

The organisation's tracker remains the system that management, planning and other teams read. Epics and high-level tickets stay there — business milestones, cross-team dependencies, customer requests — and each Epic links to its specification. The decomposition does not: tasks, contracts and change history live only in the workspace. Mirroring them as tickets would create a second copy of the plan that drifts from the first.

The agent bridges the two. With tracker tools available to it over MCP (RA-003), it can read an Epic when a specification starts and report status and progress back, instead of someone keeping tickets in step with the Markdown by hand.

Changes that arrive through the tracker take the same path. When the customer adds adjustments to an Epic as comments and asks the PO/BA to consider them, the PO/BA has the agent read the comments and assess how they affect the existing specification; what is accepted enters the specification as an amendment (§4.6). The Epic records what was asked, and the specification records what was decided.

Example: what the Epic owns and what the specification owns

5. Model Routing by Stage​

Separate sessions (§4.3) have a second consequence: each stage can run on a different model. Stages differ in how much reasoning they need, and routing them accordingly controls the cost of running the process.

StageModelWhy
RequirementsBelow the most capable tierDescribes the desired state of the product from business documentation; the owner's review makes it correct
Technical designThe most capable available, at its highest reasoning settingDesign caps the quality of everything downstream; cross-service reasoning is where capability pays off
Task decomposition and task filesThe most capable availableDecomposition is design at the scale of an agent's context: sizing tasks, isolating dependencies, fixing the contracts between them
ImplementationCheaperThe task carries the decisions; the model carries out the task
TestingCheaperScenarios and automation follow from fixed acceptance criteria and contracts
ReviewThe most capable availableReview is where mistakes from the cheaper stages are caught
Documentation updatesCheaperThey describe a change that is already made and reviewed

One rule keeps the cheaper stages safe: the less capable the model that implements a task, the more detailed the task has to be. Task detail and model capability trade against each other, and the balance is set when the tasks are written.

At the time of writing, the team routes design to Claude Opus at its maximum reasoning effort, task files to Opus one effort level lower, review to Opus, and implementation, testing and documentation updates to Claude Sonnet.

The savings come from the high-volume stages — implementation, testing, documentation — running on a cheaper model, while the most capable model is spent where it decides the outcome: design, task decomposition and review.

6. Encoding the Process in the Workspace​

A framework's commands are what keep its steps from being skipped. Without a framework, the workspace carries that job, in two kinds of files. Each is written once as a canonical file and reaches every agent through thin per-agent files — the wrapper mechanism RA-003 describes for skills, applied to rules as well.

  • Rules (steering) — short instructions delivered to every agent session: how the tier is chosen, how requirements and acceptance criteria are written, the commit, git and delivery conventions, which environments and repositories belong to the team. A rule earns its place by being something an agent would otherwise get wrong, because everything delivered on every session costs context on every session.
  • Skills — step-by-step procedures for the jobs that have to come out the same way every time:
    • intake and triage of new work into a tier;
    • writing or amending a specification stage by stage, pausing for review after each artefact;
    • executing one task end to end — branch, implementation, quality gates, merge request, progress recorded in the specification;
    • reviewing a merge request against the specification and the task it implements;
    • preparing a release, including the delivery notes it writes into specifications and tasks (§4.8).

The skills stop wherever a person decides: the tier, each artefact of the specification, a commit, a review comment before it is raised. They make the process repeatable; they don't make it autonomous.

Like any process, this one holds only as long as the team follows it. The skills make following it the easy path rather than the only one, which puts the rest on team discipline — as it would with a framework. In this team, it has held for five months, including under pre-release pressure. A team that wants a structural backstop can add cheap CI checks, such as rejecting a merge request whose commits carry no task code, without changing anything else.

Examples: rules · triage rule · requirements rule · skills: plan-task, author-spec, implement-task, review-mr

7. Scalability & Evolution Path​

7.1 Refined in use​

One engineer designed the process; the people who run it have refined it since. An engineer who hits a gap — a missing step, a rule that no longer fits, a template section nobody fills in — changes the rule, the skill or the template, and the change is reviewed like code. Because the process is a set of files in the workspace, improving it costs the same as any other small change, and every agent picks up the new version in its next session.

7.2 Without a whole-system workspace​

The shape assumes the agent sees everything. Where it doesn't — each repository opened on its own — requirements lose only their early feasibility check, since they come from business documentation. Everything after them gets heavier: the design can't be validated across services in one session, the specification needs a home outside every service repository, and each task has to carry the context its links would otherwise provide. The ideas still hold, but the process is much heavier to run. In that situation, the first investment is usually the workspace (RA-003), not more process.

7.3 Keeping it lightweight​

A home-grown process can grow into the heavy framework it replaced, one reasonable rule at a time. The tier decision is the guard: most work never reaches a specification, so ceremony added to specifications stays confined to the work that justified it. Remove a rule or template section that stops earning its cost the same way you added it.

8. Trade-off Analysis​

FeatureBenefitDrawback
No frameworkThe process fits the team's context; changing it is an ordinary editNo community, documentation or upgrade path beyond what the team writes
Agent-written, human-owned filesA specification costs guidance and review, not writing; the layout can drift without affecting the resultLess uniform for a person reading across many specifications; each stage is only as good as its owner's review
TiersSmall changes carry no ceremony; the rubric proposes, and a person decidesThe tier is a judgement; a misjudged task grows into a poor specification unless someone switches tiers
Separate sessions per stageEach stage checks the previous one instead of inheriting its assumptions; one engineer can run the whole flowMore hand-offs; everything a later stage needs has to be in the files
Logical-unit tasksBounded agent context, parallel implementation, review against intentSizing is a judgement; a large feature still produces many tasks and an ordering to manage
Two change regimesCurrent behaviour and full history in one place, after release as well as before; defects trace back to the owning sectionSpecifications grow; they stay navigable only while the changelog and the task index are kept up
One implementer per specificationNo conflicts on specification files in practiceParallel work on one specification is the exception; a feature that needs it is split first
Delivery recorded in the specificationWhat is in production, and what is missing, is visible per feature with no manual trackingDepends on a commit convention the whole team follows
Tracker at intent levelNo second copy of the plan in tickets; management keeps its viewTwo systems; status crosses between them through the agent, not by construction
Model routingThe most capable model is spent on design, task decomposition and reviewCheaper implementation needs more detailed tasks
Rules and skillsThe steps hold across agents without framework toolingEnforcement is procedural — skills, review and team discipline — not structural
Whole-system workspaceDesign is validated against real code; requirements get an early feasibility check and a full-flow estimateA hard dependency for design and implementation; without it the process is much heavier (§7.2)
Fitted to one contextEvery rule exists for a reason the team can nameAnother team re-derives the shape rather than copying it

9. Conclusion​

Spec-driven development is a small set of ideas: agree intent before code; keep what, how and the steps apart; give agents small tasks that trace back to the specification; keep the specification beside the code; verify against what was asked; and change the specification when behaviour changes. A framework is one way to package those ideas, built for its own context.

When that context doesn't match, you can apply the ideas directly and derive the process from the team's own constraints — with agents writing every artefact and a person owning every stage. For one team responsible for 80+ microservices on an AI-native workspace, the process derived that way has run for five months, carried a 35-service feature through four major updates, and handled a post-release rollback within hours.

References​