AI-Native Team Workspace (Meta-Repo Architecture)
1. Executive Summary
An AI coding agent confined to one repository cannot follow a business rule from the database schema through the service layer to the frontend component, because in an enterprise estate those live in three repositories. The usual answer — migrate to a monorepo — is a months-long project with its own risks, and it changes how every team builds and deploys to solve a problem only the agent has.
This architecture solves the agent's problem without touching anyone's repository. A meta-repo clones every repository of the system into one directory tree, keeps their histories, pipelines and deployment lifecycles exactly as they were, and holds no application code of its own. What it does hold is everything an agent needs to work in that tree, in four layers, each usable without the ones below it:
| Layer | What it gives the agent | Where it lives |
|---|---|---|
| Context — what is true | The whole system's source; the team's documentation; an outline tool over it; an opt-in hybrid search index that returns pointers into it, never prose | The tree, docs/, docs_map, docs_search (§4.1) |
| Instructions — what to do | Standing rules and multi-step skills, each written once and delivered to every agent through generated per-agent files, verified in CI | .ai/rules/, .ai/skills/, the generated pointers and wrappers (§4.2) |
| Tools — what it can call | A team MCP server for the always-on capabilities; connectors, skill-local scripts and ordinary CLIs for the rest, under one placement rule | mcp/, .ai/connectors/, <skill>/scripts/ (§4.3) |
| Guards — what stops it | Hooks on tool calls that block what is expensive to undo and report what is easy to forget | scripts/hooks/ (§4.4) |
The first edition of this document described the first layer and the beginnings of the second and third. The second edition, six months on, describes the workspace as it is now and keeps the original thesis unchanged: the context problem is solved by cloning, not by restructuring.
What it costs: automation that runs on folder open, on git pull and on an agent's tool calls, all defined in version-controlled files, which is a supply-chain surface the team has to review like deploy scripts (§5); a workspace large enough that some commands are handed back to the operator rather than run by the agent (§4.5); and a retrieval index that every laptop builds for itself, which is the right default and stops being one at a stated set of signals (§6.2).
Provenance. The author designed and built this workspace as team lead of a seven-person team responsible for 80+ microservices on an enterprise platform of 200+, and it is the team's daily working environment. The production MCP server registers 43 always-on tools, with 26 more attached to a browser-automation server for tasks that need it; tool definitions have not exhausted a session's context. The linked repository is an example implementation, not an extract of the production workspace: it commits its service, library and infrastructure repositories inline so that a fresh clone runs with no remotes to configure, whereas a real deployment clones them from the registry (stated in its README). The same repository carries the reference implementations of RA-004, RA-005 and RA-006; this document covers the workspace and its AI layers, and refers to those for what they own.
2. Applicability Criteria
Use this pattern when
- The system is split across many repositories and will stay that way. Services, frontends, shared libraries and infrastructure manifests each with their own history, CI and deployment cadence. The pattern adds a tree around them; it changes nothing inside them.
- Features routinely cross repository boundaries. An agent that can open the caller, the callee and the shared library in one session is the whole point. Work that stays inside one repository gets little from it.
- The team works agent-first. People direct and review; agents read, write and run. The rules, skills, tools and hooks below are sized for that division of labour, and they are overhead for a team that uses an agent as autocomplete.
- A monorepo migration is off the table — too costly, too risky for an established codebase, or wrong for a polyglot stack whose build systems have nothing to gain from unification.
- The organisation's tracker, CI and repository host are fixed. The workspace integrates with them over MCP; it does not replace any of them.
Do NOT use this pattern when
- The code already lives in one repository. The context problem is solved; take the AI layers piecemeal — rules, skills, the MCP server and the hooks apply unchanged to a monorepo — and skip the cloning machinery.
- One repository, one developer. The cost of the workspace is paid in setup and review discipline; it earns back across a team and a system, not a project.
- Policy forbids cloning the estate onto developer machines. The pattern assumes the agent reads real code on a real laptop. Where data classification or device policy rules that out, the shared-deployment path in §6.2 is a starting point, not an upgrade.
- Nobody will own the automation. Generated files, hooks and the index refresh are cheap to run and cheap to neglect; a workspace whose generators nobody runs drifts back to hand-maintained copies within a quarter.
3. Architecture Overview
3.1 The workspace is a routing layer, not a build system
After yarn setup, the workspace resolves to one tree. Entries marked (cloned) are managed by the sync scripts from a JSON registry; everything else is native to the meta-repo.
team-workspace/
├── AGENTS.md, CLAUDE.md, .cursor/, .kiro/, ← generated per-agent rule pointers (§4.2)
│ .github/copilot-instructions.md, .agents/ and skill wrappers — never hand-edited
├── .mcp.json, .cursor/mcp.json ← project-scoped MCP configs (§4.3)
├── .claude/settings.json ← agent hooks wired for Claude Code (§4.4)
├── .ai/
│ ├── rules/ ← 7 canonical rules, plain Markdown
│ ├── skills/<name>/SKILL.md [+ scripts/] ← 7 canonical skills
│ ├── connectors/grafana/ ← repo-local scripts, run by path
│ └── whoami.example.md ← per-person operator identity (git-ignored copy)
├── mcp/ ← the team MCP server: GitLab, Jira, Grafana,
│ MongoDB, docs_map, docs_search
├── configs/workspace-repos.json ← the registry: remote, path, GitLab project id
├── scripts/ ← setup, sync, daily guard, generators, hooks/
├── .githooks/post-merge, post-rewrite ← sync + change-aware reindex on pull
├── .vscode/settings.json, tasks.json ← detectSubmodules: false; the folderOpen task
├── docs/ ← team knowledge: architecture, guides, sdlc/,
│ SPECs/, tasks/, incidents, spikes
├── backend/*, frontend/*, libs/*, infra/git-ops ← (cloned) the system itself
The meta-repo doesn't try to unify the build systems of what it clones. The example is all JavaScript and uses yarn workspaces for convenience; a polyglot estate keeps per-repository builds, and nothing in the workspace depends on it being otherwise.
3.2 Four layers, each usable without the ones below
The order is the adoption order. A team gets most of the value from the tree and a README at every level, with nothing else installed. The rules and skills add consistency across agents and people. The MCP server adds the external systems. The hooks add a backstop for the three or four mistakes that are cheap to make and expensive to undo. Each layer is independently removable, and the design keeps them that way on purpose: a broken hook must never wedge a session, a disabled search tool must not exist in the tool list, a missing whoami.md must only cost a question.
3.3 The registry drives the sync
One JSON file, configs/workspace-repos.json, lists every repository by category — frontends, backends, libs, infra — with its remote, where it lands on disk, and the GitLab project id the MCP tools address it by:
{
"name": "products-service",
"description": "REST API — Products domain (catalogue, inventory, pricing)",
"git": "git@github.com:your-org/tw-products-service.git",
"path": "./backend",
"localPath": "backend/products-service",
"projectId": 2002
}
localPath and projectId are read by the skills, not by the sync scripts. A skill that pushes a branch or comments on a merge request takes the project id from here and never guesses it; the registry's own comment says why — a wrong one posts comments on, or pushes to, a stranger's repository.
Three scripts cover the lifecycle, all plain Node with no dependencies: setup-workspace.mjs installs the git hooks, clones what is missing and rebuilds the docs index once at the end; update-workspace.mjs pulls what exists and clones what is missing, and deliberately never touches the index; daily-setup-guard.mjs runs on every folder open and gates the real work to once per calendar day, because "the first time someone starts work today" is not an event any tool emits. Two tracked git hooks, post-merge and post-rewrite, sync the nested repos after a pull — the second one because a rebase-by-default team never fires the first, and a post-merge-only setup looks silently broken.
Entries whose remote still says your-org are skipped as placeholders, and a registered directory that is not a git repository of its own is reported as INLINE and left alone — which is what lets the example run green while committing its repositories inline (repo-sync.mjs).
3.4 Constraints decide the shape
| Constraint | Decision | Where |
|---|---|---|
| The estate is many repositories and a migration is not on offer | Clone into one tree; change nothing inside a clone | §3.1 |
| Several agents in use, each reading instructions from its own path | One canonical file per rule and skill; per-agent files generated and checked in CI | §4.2 |
| A prose corpus too large to scan, dense with identifiers | Hybrid retrieval that returns pointers, with a zero-dependency outline tool underneath it | §4.1 |
| Retrieval cannot tell a plan from a description | The corpus is declared, not discovered; specs, tasks, skills and rules are excluded | §4.1 |
| Some capabilities are needed in every session; others in one workflow | MCP for always-on; connectors, skill scripts and CLIs for the rest, under one placement rule | §4.3 |
| Every tool call runs under a developer's own credentials | The MCP server is a local stdio process reading a local .env; attribution is free | §4.3, §6.2 |
| A few mistakes are cheap to make and expensive to undo | Hooks that block those and report everything else; fail open | §4.4 |
| The workspace is large enough to degrade a laptop | Rules on which commands the agent hands back, and a stop rule for a degrading shell | §4.5 |
| Automation that runs unasked is a supply-chain surface | Every automated action is a one-line call into a tracked script; user-level trust settings are the gate | §5 |
4. Implementation Details
4.1 Context: a README hierarchy, an outline, and a hybrid index
The first edition claimed that a README at every directory level lets an agent traverse the system without requiring vector search or RAG indexing. That is still the first layer and still true for navigation. It stopped being enough once the workspace held a documentation corpus — architecture, guides, business flows, runbooks, incident write-ups, plus the README of every service — that an agent could not scan and would otherwise grep its way around, slowly and blind to anything phrased differently from the query. The answer is a ladder of three, each usable without the ones below (design):
| Layer | What the agent gets | Needs |
|---|---|---|
| README hierarchy | Predictable places for written knowledge; the pointer file an agent loads on every session (§4.2) links to the README of every area | nothing |
docs_map (MCP tool) | A table of contents of docs/: every file → its sections → line ranges, with a one-line summary each | nothing |
docs_search (MCP tool, opt-in) | Hybrid semantic + keyword retrieval across the whole curated corpus, ranked | a local embedding model and a local Qdrant |
docs_map is the honest answer to "do we even need RAG?" — ship it first, use it, and build docs_search only for the conceptual queries an outline cannot route.
Both return pointers, never prose. A result is a file path, a heading path, a line range and a snippet. Neither tool answers the question; it locates the section, and the agent reads the real file. That keeps the retrieval layer dumb and debuggable, and it removes the summarise-then-hallucinate failure mode at the design level rather than by prompt.
Retrieval is hybrid because the corpus is identifier-heavy. Role values, environment variable names, header names, project ids. Dense embeddings put role:workspace:team_lead and role:workspace:regional_admin in the same neighbourhood; a BM25 sparse vector keeps them apart. Both live in one Qdrant collection as named vectors and are fused server-side with reciprocal rank fusion. Embeddings run in-process in Node through ONNX; nothing leaves the machine, and no Python is involved (pipeline). At a few thousand chunks, this is a routing problem, not a scale problem, so there is no reranker and no query-rewriting stage: the consumer is a strong model that can judge relevance itself.
The corpus is declared, not discovered (sources.js). A workspace of cloned repositories contains a great deal of Markdown that is not team knowledge, so the index is an explicit list — docs/, the connector docs, service and library READMEs one level deep so a newly cloned service is picked up with no configuration — and what it leaves out is the decision worth understanding before copying the setup:
| Excluded | Why |
|---|---|
| Feature specifications and task one-pagers | They describe intent, not reality. Retrieval has no notion of planned versus shipped, so a spec for an unbuilt feature reads exactly like documentation of a working one — a wrong answer delivered confidently |
| Agent skills and rules | Instructions to an agent, not knowledge about the system, and already delivered by the agent runtime. Indexing them adds a second, worse delivery path in which procedural text competes with reference docs on the domain words it necessarily contains |
| Changelogs and templates | Rewritten by CI, or placeholder prose that matches structural queries |
The excluded files are still read — by path, knowing what they are. The rule the agent carries says so: a docs_search that comes back empty on a feature does not mean nobody has planned it (docs-index.md).
The index looks after itself, and the feature is isolated. It rebuilds when a git pull brings changed documentation, at most once a day otherwise, and self-heals if the Qdrant volume is wiped; rebuilds are blue-green behind an alias swap, so they are safe while someone is searching. docs_search is off by default; when disabled the tool is omitted from the server's tool list, so the model never sees it, and its heavy dependencies are optional and loaded on first use, so if the model download, the container engine or Qdrant is missing, one tool fails and everything else in the server is untouched (safe by construction).
The footprint is the fair objection, and it is small. Enabled and idle, the server holds ~176 MB after the first search, all of it native memory that does not grow with use; a query is sub-second on CPU; the background ingest runs in a child process that frees its memory on exit, at most once a day and on documentation pulls; there is no GPU, no cloud API, and after the first model download no network at all (measured). Known limits are stated in the design doc rather than discovered: no incremental indexing (BM25 statistics are corpus-wide), local edits are not picked up until a rebuild, every developer re-embeds the same corpus, English-only tokenisation. The third of those is the one that eventually argues for moving the index off laptops (§6.2).
4.2 Instructions: rules and skills, written once, generated per agent
Every coding agent reads its instructions from a different conventional path — CLAUDE.md, .github/copilot-instructions.md, .kiro/steering/, .cursor/rules/, .agents/. Maintaining one copy per agent by hand is how they drift, and the drift is invisible: in this repository, before the generator existed, the same skill was named Debug & Report in one wrapper and debug-and-report in another with a different description, so it triggered differently depending on which agent you asked. The fix is structural. Two kinds of canonical files live in .ai/, and everything per-agent is generated from them and verified in CI:
yarn rules:sync # the six rule pointers, from .ai/rules/ + CONTRIBUTING.md
yarn skills:sync # the four skill wrappers per skill, from each SKILL.md's frontmatter
yarn agents:check # fail if anything has drifted
Rules are standing conventions: branch and commit format, what the environments mean and how a version is promoted, which commands the agent hands back to the operator, search before you build, how much process a piece of work earns, how requirements and estimates are written (the seven). They reach an agent through a pointer file at the path it already reads, and every pointer says the same thing: the rules are in CONTRIBUTING.md, which indexes .ai/rules/. The pointer is what an agent reads on every session, and the rules themselves are read on demand, so a rule earns its line in the pointer by being one an agent would otherwise get wrong — which is why there are seven and not thirty. Six identical files are a maintenance smell only when they are maintained by hand (generator).
Skills are procedures for jobs that have to come out the same way every time: triage new work, write or amend a specification, execute one task through to a merge request, review a merge request, promote versions between environments, turn a security scan into verified fixes, debug an environment issue and raise the bug (the register). The canonical SKILL.md is plain Markdown; each agent receives a thin wrapper carrying only the frontmatter that agent needs for discovery and a pointer back — a Claude Code slash command, a transcluding Kiro skill, a Copilot skill, an Antigravity skill (generator). The description in the frontmatter is the trigger, and it is the field that drifted: a skill whose description differs per agent runs in one tool and not in another for no visible reason.
What makes a skill worth writing rather than leaving to judgement, as the register states it: numbered steps with a stated stopping point; the reasoning behind a step and not only the instruction, because a step that says what is followed once and worked around the next time; an explicit Never list; stopping points stated as gates, not advice — author-spec pauses after each artefact and will not draft ahead, implement-task will not commit without approval for that specific commit; and scripts for anything mechanical, prose for anything that is a decision. Skill scripts live in <skill>/scripts/ and write scratch output to a git-ignored <skill>/output/.
Four of the seven skills — plan-task, author-spec, implement-task, review-mr — are the delivery lifecycle, together with the work-triage and requirements-and-estimates rules and the docs/sdlc/, docs/SPECs/ and docs/tasks/ folders. That process is RA-005, which this document does not repeat; here they are simply four more canonical files delivered by the same mechanism.
A small mechanism completes the picture: .ai/whoami.md, per person and git-ignored, tells an agent who is operating so that a skill can attribute work and choose which stage of a staged artefact is theirs to write. Nothing loads it automatically; the skills that need it read it, and fall back to git config user.name plus a question.
4.3 Tools: four places, one placement rule
The first edition presented shell-script connectors as the mechanism and a team MCP server as an alternative. In use, the order is inverted: the MCP server is where the always-on capabilities live, and connectors are one of three other forms, each with a place. The placement rule is written once, in .ai/README.md, and every other file that touches the question links to it:
| Kind | How the agent reaches it | Credentials | Put it here when |
|---|---|---|---|
| MCP tool | Auto-discovered: the agent's client starts the server from a project-scoped config and lists its tools | Root .env | Needed across workflows, or by an agent with no skill running; benefits from being always available |
| Connector | By path: node .ai/connectors/<service>/<script>.mjs, named in a skill | Root .env | A skill-specific integration, a one-off, or the CLI face of an MCP capability that a person also needs from a terminal |
| Skill-local script | From inside that skill: node .ai/skills/<skill>/scripts/<script>.mjs | none today | A mechanical step belonging to exactly one skill |
| Third-party CLI | Its own invocation — git throughout; kubectl/kustomize in the release skill; yarn audit in the security skill | The tool's own auth | The tool already does the job and wrapping it adds nothing; the rules decide which of these the agent runs and which it hands over |
The MCP server is a stdio process in mcp/, started by the agent's client from a config committed to the repository — .mcp.json for Claude Code and VS Code, .cursor/mcp.json for Cursor, absolute-path templates for the clients configured outside the repository (per agent). Nothing is deployed or hosted. It reads credentials from the developer's own root .env, which has a consequence worth stating: every action on the repository host and the tracker — a push, an approval, a ticket — is attributable to the person whose token it ran under, for free (§6.2 is where that stops being free). The example registers 29 tools: GitLab (pipelines, jobs, merge requests, review comments, a safe_push with the protections a raw git push lacks), Jira, Grafana (log search, and grafana_trace_id, which returns the reconstructed call chain of one request in one call — the engine is RA-004's), MongoDB read-only by construction, and the two documentation tools of §4.1. Tools whose configuration is absent are left out of the tool list rather than registered to fail (the server).
What a tool costs. Every active tool's name, description and schema is sent to the model on every turn, used or not. The first edition turned that into a numeric ceiling on tool count, written for the context windows of the time; the second edition withdraws the number. The production server always registers 43 tools, and for browser work a second server adds 26 more, without tool definitions exhausting a session. What survives is the mechanism, and the recommendation it implies: a description is one sentence stating when to use the tool, parameters are described in the schema and not in prose, near-duplicates are one tool with a mode parameter, and a tool a skill depends on is named in that skill at the step that uses it rather than left for the model to find by description (adding a tool).
One capability, two faces. trace-id legitimately exists in two forms. The reconstruction engine exists once, in mcp/src/grafana/trace/; the grafana_trace_id tool calls it, and the connector trace-id.mjs imports it rather than shelling out to the server or duplicating the logic. The agent gets it as a tool, a person gets it as a CLI, and the two cannot disagree about what a trace means. That is the only acceptable shape for a connector that mirrors an MCP tool: a second face on one engine, never a second implementation. The three connectors that had been planned since the first edition — Jira, GitLab, MongoDB — were deleted in the revision that accompanies this document, because every script they listed had an MCP tool doing the job.
4.4 Guards: hooks that block what is hard to undo and report the rest
Rules and skills are procedural; a model that has read never push to main can still shell out and do it. The fourth layer is five small guards that run on tool events (scripts/hooks/), wired for Claude Code in .claude/settings.json and agent-neutral by contract: the event arrives as JSON on stdin and the answer is an exit code — 0 allow, 2 block a pre-tool call or report back to the model after a post-tool call, anything else treated as a broken hook and ignored, because a failing hook must never wedge a session.
| Hook | Event | Does |
|---|---|---|
guard-secrets | before Bash | Blocks a git command that names a credential file or carries a token-shaped string; refuses git add -A only when an untracked .env is actually in the root — a hook that cries wolf gets switched off inside a week |
guard-protected-branch | before Bash | Blocks a raw push to main, a --force without --force-with-lease, a bare git push. gitlab_safe_push already enforces this; the hook closes the gap of the agent shelling out instead |
docs-index-staleness | after Edit/Write | Reports that an indexed file changed and the index is now stale for it — corpus membership evaluated against sources.js itself, not a second copy of the rules |
validate-overlay | after Edit/Write | Builds the affected GitOps overlay after a manifest edit; hand-edited kustomize fails in a way that looks fine in review |
docs-delivery-gate | after a push | Reports that the task's docs/ deliverable is still outstanding. Team docs live in this repository while the code lives in a service repository, so they cannot ride in the service merge request — which is precisely why they get dropped |
Two of the five block; three report. The rule is block only what is genuinely hard to undo — a committed token is rotated and the history rewritten; a pushed branch cannot be un-pushed by refusing it — and report everything else, because a gate that blocks delivery over documentation is also switched off inside a week. Every hook says what to do instead of what it refused, since a refusal without an alternative only produces a retry. The hooks are wired for one agent today; another agent's hook mechanism points at the same scripts, and the two functions that touch the event shape are isolated in one file for exactly that.
4.5 What the agent hands back to the operator
A workspace this size — a dozen or more cloned repositories, a vector store, a local embedding model — has commands that are slow, and a few that reliably break the session. The local-environment rule names them and makes handing them over the normal flow rather than an escalation: the agent never runs the index rebuild, dependency installs or the workspace-wide bootstrap; when a task needs one, it states the exact command and directory, says why in one line, and waits. The same rule sets a stop rule for a degrading shell — garbled echo, arguments going missing from executed commands, wrong exit codes — and treats the first sign as a hard stop, not something to retry around, because a mangled argument to git restore or rm acts on the wrong path. This is a direct cost of the pattern: the tree that gives the agent everything at once is also heavy enough that the agent has to know what not to run.
4.6 Editor and machine settings
Two settings are non-negotiable for a machine running this workspace, and they sit at different trust levels for a reason.
"git.detectSubmodules": false, committed in .vscode/settings.json. Without it, the editor scans every cloned repository in the background, slows noticeably and floods the source-control view with unrelated statuses. It affects nothing about git itself.
"task.allowAutomaticTasks": "on", in the developer's user settings, which a workspace cannot grant itself. It is what lets the folderOpen task run the daily guard, and its absence is the most common reason someone reports that the daily setup never happens — the automation looks broken with no error anywhere. That a repository cannot switch it on is the deliberate boundary §5 relies on; a developer who clones widely leaves it off and runs yarn daily-setup by hand, at the cost of one command a day.
5. Security: version-controlled automation is code execution
Three mechanisms in this workspace run code on a developer's machine without asking — the folderOpen task, the git hooks, and the agent starting the MCP server from a committed config — and a fourth runs on every tool call the agent makes. All of them read their definitions from version-controlled files: .vscode/tasks.json, .githooks/, .mcp.json, .cursor/mcp.json, .claude/settings.json. That is what makes the automation shareable, and it is exactly what makes it a supply-chain surface: anyone who can land a commit can change what runs on your machine (the full argument).
The real threat isn't a movie hack. It is a plausible-looking node scripts/... line added to tasks.json in a 200-file merge request that reviewers skim; opening an untrusted clone with automatic tasks enabled; a compromised dependency with a postinstall script, which needs no help from any of the above; or a local hook edited months ago that no diff will ever show. The failure mode is the same as everywhere else in this design: it is silent. A malicious task that also does the legitimate work looks exactly like a working setup.
The mitigation is not to avoid automation — a stale index that quietly returns wrong answers is the problem this whole layer exists to solve — but to keep it small, tracked, reviewable and boring:
- Every automated action is a one-line call into a tracked script under
scripts/, so the reviewable surface is normal code that goes through normal review, not logic in a JSON string that no linter reads. - Diffs to the five files above are reviewed like diffs to a deploy script; a
CODEOWNERSentry keeps them from changing without a named reviewer, andgit log --oneline -- .githooks .vscode/tasks.json .claude/settings.json .mcp.json .cursor/mcp.jsonis the audit command that shows who changed the automation and when. - The trust gates are per machine and outside the repository's reach: the user-level automatic-tasks setting; editor workspace trust; and, for the MCP config, the client's own approval — Claude Code prompts before using a server from a project
.mcp.jsonand ignores a committed setting that would auto-approve it until the folder is trusted, and that prompt is the gate, not an inconvenience (per client). - Nothing needs elevated privileges and nothing writes outside the workspace and its own
.git/; the hooks fail open, so removing them breaks nothing except freshness. - Credentials live in one git-ignored root
.envshared by the MCP server and the connectors, so there is one file to rotate and one file forguard-secretsto watch; database tools are read-only by construction; and whether production data may be reached at all is decided with the knowledge that whatever a tool returns is sent to a third-party model provider.
6. Scalability & Evolution Path
6.1 More repositories
The pattern's cost grows with the number of clones in three places: editor scanning, which the committed settings handle; the shell, which §4.5 handles by rule; and the retrieval corpus, which picks up a new service's README automatically at depth one and nothing deeper unless it is declared. Adding a repository is one registry entry and yarn setup; removing one is deleting the entry and the directory, and the next ingest drops its content because every rebuild indexes exactly what exists rather than diffing.
6.2 The retrieval index: from laptops to a shared deployment
Every laptop building its own index is the right default, not a stepping stone: no infrastructure, no access control to get wrong, no shared thing to break, and it works offline. It stops being right at a recognisable set of signals, none of which is about corpus size: a full rebuild long enough to notice, on every machine, producing byte-identical indexes; rebuild spikes that interrupt people; two developers getting different answers to the same question because their rebuilds ran at different times; a corpus spanning repositories not everyone clones, so a partial checkout silently produces a partial index; a new joiner's first hour spent on a model download and a full embed. The reason to move is duplicated work and divergence, which is why the fix is centralisation and not a bigger laptop (signals).
Before that, four cheaper levers usually solve the problem that felt like scale: shrink the corpus — most slow indexes are indexing things they should not; widen the daily refresh window; cap the embedding threads so the ingest spike stops being noticeable; and, at a corpus where full rebuilds genuinely hurt, incremental indexing, which is blocked on BM25 statistics being corpus-wide and is not worth building before then.
When the signals persist, the move is two separate steps, and the second does not follow from the first:
Step 1 — a shared index. Qdrant becomes one instance with CI as its only writer: the same ingest pipeline runs on merge to the documentation branch and on a schedule, and laptops never write. Developer clients read through a gateway that validates the team's existing authentication and injects a read-only key; Qdrant itself has no public ingress. The MCP server stays on the laptop, still started by the agent from the workspace config, and every GitLab, Jira, Grafana and MongoDB call still runs under the developer's own credentials — only the docs_search query crosses the network. Queries are still embedded locally, since Qdrant stores and fuses vectors but does not compute them, so a laptop still downloads the model; it just never ingests. Sizing is small — a few hundred megabytes of memory and a gigabyte of volume cover a documentation corpus — and the index needs no backup, because it is reproducible from the Markdown. There is one non-obvious blocker: query time needs the BM25 model that ingest writes, and once ingest runs in CI the laptops do not have it; a vocabulary from one corpus version scoring vectors from another does not error, it silently degrades ranking, so the model has to travel with the collection it was fitted on (the blocker).
Step 2 — a shared MCP server. The server itself moves off the laptop, behind a remote transport, and the last per-machine piece goes with it: no local Node process, no .env on every laptop. It has two costs step 1 does not. An auth layer in front of MCP itself — today the only caller is the agent that spawned the process, so there is nothing to authenticate; a network-reachable server that can push branches, approve merge requests and query databases has to know who is calling. And the credentials model changes: today every action is attributable to a person because it runs under their token; a shared server needs either a service account, under which pushes, approvals and ticket edits all happen as a bot — a governance problem before it is a technical one — or identity pass-through, in which the server exchanges each caller's identity for per-user tokens on every downstream system and acts as them. Attribution survives that, and it is most of the work of step 2.
| Local (today) | Step 1 · shared index | Step 2 · shared MCP server | |
|---|---|---|---|
| Compute | The laptop, on demand; the ingest child frees its memory on exit | A Qdrant service allocated 24/7 regardless of use; the laptop still embeds queries but never ingests | Qdrant and the MCP server, both allocated 24/7 regardless of use |
| Access control | None needed | A gateway and a read-only key | Auth on MCP itself, plus a credential model |
| Operations owner | Nobody — the laptop self-heals | A stale index is a CI failure that needs an owner and an alert | The same, plus the server's uptime |
| Offline | Works | No | No |
| Consistency across machines | Can diverge | One index | One index |
| Onboarding | Model download and a full embed on the first day | Model download only | None |
| Per-user attribution of actions | Yes — every call runs under the developer's own token | Unchanged | Only with identity pass-through |
Of all this, one piece is code today — QDRANT_ENGINE=external, under which the server treats Qdrant as managed elsewhere and never tries to start or repair it. The CI writer, the BM25 model stored beside the collection, the flag that turns off local ingest, the read-only gateway and everything in step 2 are designed, not built, and the repository says so in those words (implemented today, and designed). A team that outgrows the local default usually needs step 1 and nothing more.
The position this architecture takes: for most teams, the local option is the cheapest. It needs no service, no owner and no access control, and its footprint does not load the host. Each step up is a deployment change, not a redesign — the ingest pipeline, the corpus and the query path are the same code at every rung, and what step 2 adds sits around the server rather than inside it — which is why deferring the move is safe.
6.3 More agents, more tools
Supporting a new coding agent is one entry in each generator — a pointer path, a wrapper template — and, where the client supports one, a project-scoped MCP config; the canonical rules and skills do not change. The delivery table in .ai/README.md shows, per agent, its rules pointer, skill wrapper and MCP config side by side, and the honest gaps: clients whose documentation defines no workspace variable get an absolute-path template rather than a committed file that might not start. Growing the tool surface is governed by §4.3's description discipline, not by a count.
6.4 The process layer
What runs on the workspace — how work is triaged into a specification, a one-page task or a direct change; how a specification is authored stage by stage and implemented task by task; how delivery is recorded — is RA-005. It depends on this architecture as its decisive precondition and adds nothing to the mechanisms described here beyond four skills, two rules and three folders.
7. Trade-off Analysis
| Feature | Benefit | Drawback |
|---|---|---|
| Meta-repo over monorepo | Whole-system context for the agent; no migration; every repository's history, CI and deployment unchanged; polyglot by construction | Clones drift if sync is skipped — mitigated by the hooks and the daily guard; onboarding meets an unfamiliar shape, mitigated by one yarn setup |
| Retrieval as pointers | Debuggable; no summarise-then-hallucinate; hybrid ranking keeps identifiers apart; opt-in and isolated | Every laptop embeds the same corpus; no incremental indexing; local edits invisible until a rebuild; curation is a standing job, not a one-time decision |
| Generated per-agent files | One source; drift caught in CI; supporting a new agent is a template | Generated files must never be hand-edited; the generators must know every agent's path, and a file not named there is silently left out |
Local stdio MCP on a local .env | Nothing hosted; every action attributable to a person for free; credentials never leave the laptop | Every tool's schema costs context on every turn; sharing the server later changes the credentials model (§6.2) |
| Hooks | Backstop for the mistakes that cannot be undone; agent-neutral contract; fail open | Wired for one agent today; fail-open means a broken hook is silent; procedural enforcement elsewhere still rests on review and discipline |
| Automation on open, pull and tool call | A fresh index and synced repos without anyone remembering to run anything | A supply-chain surface made of version-controlled files; gated by user-level trust settings the repository cannot grant, which is also why the automation sometimes looks broken |
| Workspace size | Everything in one tree | Editor scanning, shell degradation under load, and a class of commands the agent must hand back rather than run |
| Example committed inline | The reference repository runs from a fresh clone with no remotes | A reader who clones it sees a monorepo; the README and the INLINE sync status exist to correct that |
8. Conclusion
The agent's context problem in a multi-repository estate is solved by cloning, not by restructuring: one tree, every repository unchanged, nothing hosted. What turned that tree into a working environment since the first edition is the four layers on top of it — retrieval that returns pointers into a curated corpus, rules and skills written once and delivered to every agent by generation, a team MCP server for the capabilities every session needs with a placement rule for everything else, and hooks that block the handful of mistakes that cannot be undone. Each layer is removable, and each costs something stated here next to the decision that incurred it: a supply-chain surface that has to be reviewed, a workspace heavy enough that the agent hands some commands back, and an index every laptop builds for itself until a named set of signals says otherwise.
The pattern is in daily use by a team responsible for 80+ microservices on a 200+ service platform, and it is the precondition the team's delivery process (RA-005) is built on. The example repository is small enough to read in an afternoon and, since this revision, says on its first page what it is.
References
- Example implementation: github.com/ivanbaha/team-workspace — what it demonstrates ·
.ai/— rules, skills, connectors, placement rule · MCP server · hybrid RAG design · agent hooks · workspace automation and its security - Extended by: RA-005 · Lightweight Spec-Driven Development
- Related architectures: RA-004 · Log-Based Distributed Tracing · RA-006 · Appropriate Caching