# The ArchTenet Engineering Playbook > Architecture patterns and engineering field reports by Ivan Baha and Vladyslava Prykhodko, drawn from an enterprise platform of 200+ microservices. Two kinds of page: reference architectures, which are prescriptive and implementable, each with applicability criteria, measured costs and a reference implementation; and journal articles, which report what a change did in production, with the trade-off stated rather than buried. Subjects include microservices, distributed caching, observability and tracing, authorization (ABAC), CI/CD and supply-chain security, MongoDB at scale, cloud cost engineering, and AI agents in engineering work. Every page here has a Markdown version at the same path with a `.md` suffix (`/blog/trace-id-in-every-log-line` becomes `/blog/trace-id-in-every-log-line.md`), served as `text/markdown`. A request that sends `Accept: text/markdown` gets that version at the original URL. The links below point at the Markdown; drop the `.md` for the page a person would read. All reference architectures are also available in one fetch at https://archtenet.dev/llms-architectures.txt. ## Reference Architectures > Prescriptive, implementable patterns. Each one states its applicability criteria, what it costs, what it gives up, and where a reference implementation lives. Use this section when the question is how to design or build something. For what a pattern did in production, follow the field report linked inside it. - [RA-001 — Push-Based GitOps for Staging (Test) Environments](https://archtenet.dev/docs/reference-architectures/ra-001-gitops-lite.md): A lightweight CI/CD pattern using Docker Compose and GitHub Actions. (Status: Stable / Production-Proven) - [RA-002 — Logical Database-per-Service (Shared Cluster)](https://archtenet.dev/docs/reference-architectures/ra-002-logical-db-per-service.md): A pragmatic implementation of the "Database per Service" microservices pattern. (Status: Stable / Production-Proven) - [RA-003 — AI-Native Team Workspace (Meta-Repo Architecture)](https://archtenet.dev/docs/reference-architectures/ra-003-ai-native-meta-repo.md): Gives AI coding agents the whole multi-repository system in one tree without a monorepo migration, and layers on top of it what an agent needs to work there — curated retrieval over the team's documentation, rules and skills written once and generated per agent, tools over MCP, and hooks that block what is expensive to undo. (Status: Production-Validated · One Team, 80+ Microservices · In use since late 2025) - [RA-004 — Log-Based Distributed Tracing (Trace-Id Propagation Through Shared Libraries)](https://archtenet.dev/docs/reference-architectures/ra-004-log-based-tracing.md): Correlates a request across every microservice it touches with one HTTP header, one field in the structured log line, and the log pipeline the platform already runs — no tracing backend, no sampling, no process besides the application. (Status: Production-Validated · Enterprise Scale · In use since 2022) - [RA-005 — Lightweight Spec-Driven Development (Core Principles Shaped to an AI-Native Meta-Repo)](https://archtenet.dev/docs/reference-architectures/ra-005-lightweight-sdd.md): Keeps the ideas spec-driven development rests on — intent agreed before code, requirements, design and tasks as separate artefacts, small traceable tasks, specs versioned beside the code — without adopting an SDD framework, and derives the rest of the process from the team's own constraints. (Status: Production-Validated · One Team, 80+ Microservices · In use since 2026) - [RA-006 — Appropriate Caching in a Multi-Service Estate (Balancing Efficiency and Simplicity Across a Shared Store)](https://archtenet.dev/docs/reference-architectures/ra-006-appropriate-caching.md): Eliminates cross-service HTTP hops by letting consumers read owner-managed keys directly from a single shared cache, using server-enforced ACLs, operation-specific failure policies, and a disciplined caching ladder that selects the simplest honest invalidation strategy for each access pattern. (Status: Production-Validated · 40+ Microservices Across Several Teams · In use since 2026) ## Series: CI Pipeline Security > Why standard CI/CD pipelines are vulnerable to zero-alert credential exfiltration, where preventive controls reach their limit, and how to design build systems that contain compromise and win on time. - [The Silent Exfiltration: Why Your CI Pipeline Is an Open Vault](https://archtenet.dev/blog/silent-exfiltration.md): Why your Node.js CI pipeline is an open vault, and how to fix three critical structural vulnerabilities. - [Prevention Has a Ceiling: Designing CI Pipelines That Survive Being Breached](https://archtenet.dev/blog/prevention-has-a-ceiling.md): Why the world's best-defended engineering organisations still had credentials stolen in 2026 CI attacks, where prevention reaches its limit, and how to design pipelines that survive being breached. ## Series: ABAC in Production > One authorization model, three deployment shapes. Where in-process guards stop being enough, what an externalized PDP really costs once you measure it, and how to go node-local without trading away data freshness or latency. - [In-Process Authorization with Guards: The Default That's Enough Until It Isn't](https://archtenet.dev/blog/abac-in-process-guards.md): Part 1 of a series on ABAC in production. The boring, correct starting point — authorization that runs inside your application — and the precise conditions under which you should outgrow it. - [Externalized Authorization Across a Hundred Services: What OPA/OPAL Sidecars Actually Cost](https://archtenet.dev/blog/abac-opa-opal-sidecar-cost.md): Part 2 of a series on ABAC in production. What happens when in-process guards stop being enough, and what the bill looks like — measured, not estimated. - [ABAC in Production: The Migration — DaemonSet, a Node-Local Attribute Cache, and Policy Delivery from Git](https://archtenet.dev/blog/abac-daemonset-migration.md): Part 3 of a series on ABAC in production. How to pull the PDP out of every pod without losing data freshness or hitting a latency wall — DaemonSet, a node-local cache, and why policy and attributes travel different paths. ## Series: Building a Personal AI Agent > Building Azek, a personal AI agent that runs entirely on hardware I own. What it actually takes: the constraints that justify building instead of renting, and the architecture decisions that keep a single-GPU host fast, private, and cheap. - [Why I'm building my own AI agent (and why it runs on my hardware)](https://archtenet.dev/blog/why-im-building-my-own-ai-agent.md): Part 1 of a series on building a personal AI agent. Why build from scratch to run on a laptop when capable cloud assistants are a commodity? The ownership, privacy, and learning constraints behind Azek. - [Two flows: why a second, simpler model runs beside the primary](https://archtenet.dev/blog/two-flows-second-model-and-the-primary.md): Part 2 of a series on building a personal AI agent. Why Azek runs a second, smaller model beside the primary: the KV-cache economics of a single-GPU host, sixteen bounded jobs, and the one audited pipe that keeps them safe. ## Engineering Journal > Field reports and measured experiments: what a change cost, what it saved, and the numbers behind both. Use this section for evidence, trade-offs and failure stories. Use Reference Architectures above for the prescriptive form of the same work. - [Spec-Driven Development Is Simpler Than You Thought](https://archtenet.dev/blog/spec-driven-development-is-simpler-than-you-thought.md): How a team running 80+ microservices adopted spec-driven development without a framework: six core ideas, an AI-native meta-repo where agents see the whole system at once, and a triage step that keeps most work away from specifications altogether. - [A Trace Id in Every Log Line: How Cross-Service Debugging Went From a Four-Engineer Session to One Grafana Query](https://archtenet.dev/blog/trace-id-in-every-log-line.md): Field report from a 200+ microservice enterprise platform. A trace id delivered through the shared libraries one team's 80+ NestJS services already used now covers more than 100 services, and turned hours of multi-engineer debugging into a query an analyst runs in minutes – and later into a tool an AI agent drives. The cost: call topology is inferred, not known. - [One Vulnerability, a Hundred Pipelines, One Board](https://archtenet.dev/blog/ci-deck-hundred-pipelines.md): How a weekend-built, zero-dependency local tool replaced a wall of browser tabs and streamlined security rollouts across a hundred GitLab microservice pipelines. - [The Hidden Complexity of "Easy" Geo-Sharding in MongoDB](https://archtenet.dev/blog/mongodb-geosharding-hidden-cost.md): Geo-sharding looks trivial to adopt — add a shard key, keep data close to users. That simplicity is exactly where the expensive failures hide. A field story about symptoms, root cause, and the bill that arrives monthly. - [Edge-Native RAG — When the Small Stack Is the Right Stack](https://archtenet.dev/blog/rag-at-the-edge.md): Every AI product eventually grows a memory layer, and every memory layer eventually grows a bill. There's a smaller version of that stack that handles a surprising amount of real work — until it doesn't. Here's where the line actually is. - [TDD in the AI Era — How a Luxury Practice Became a Survival Tool](https://archtenet.dev/blog/tdd-ai-era.md): A few years ago TDD was a luxury nobody could afford. AI-assisted development quietly flipped that equation — and now tests are the only thing standing between a confident model and a silent misinterpretation. - [Enterprise AI Compliance: Steerings Over Skills](https://archtenet.dev/blog/enterprise-ai-compliance.md): Why engineering teams must implement a dual-layer architecture of static steerings and dynamic skills to prevent AI agents from silently violating security baselines. - [The Token Tax — How Naive M2M Authentication Quietly Drains Your Cloud Budget](https://archtenet.dev/blog/token-tax-m2m-auth.md): How a naive M2M authentication pattern silently multiplies cloud identity costs — and the lightweight token proxy architecture that cuts operations by 99%. - [The AI-Native Team Workspace: Solving the Multi-Repo Context Crisis](https://archtenet.dev/blog/ai-native-meta-repo.md): Engineering teams are hitting a wall with modern AI coding agents. They lack the system-wide context required to make accurate, architectural-level contributions. - [The "Logical DB-per-Service" Pattern at Scale](https://archtenet.dev/blog/db-per-service-at-scale.md): How we scaled over 200 microservices using a single MongoDB cluster via the Logical Database-per-Service pattern. - [The "GitOps-Lite" Pattern for Small Projects](https://archtenet.dev/blog/gitops-lite-pattern.md): Why we chose Docker Compose over Kubernetes for our Test Environment. ## Optional > Commentary and model reviews, tied to the moment they were written. Consult these only when the question is about them specifically. - [About the authors](https://archtenet.dev/about): Ivan Baha and Vladyslava Prykhodko — background, ORCID, and how to get in touch. HTML only. - [First touch to Claude Fable 5 - A Capricious Vibe-Coding Frontier](https://archtenet.dev/blog/first-touch-to-claude-fable-5.md): With the general availability of Anthropic’s new 'Mythos-class' reasoning model, Claude Fable 5, the developer ecosystem has been flooded with benchmarks. Here is how it holds up under real conditions. - [Claude Fable 5: A Breakthrough in Cybersecurity or a Comedy of Guardrails?](https://archtenet.dev/blog/claude-fable-cybersecurity-guardrails.md): When Anthropic launched Claude Fable 5 as a breakthrough in cybersecurity, I tested it on my own codebases. The results weren't just disappointing—they exposed a structural comedy of guardrails. - [Tokenmaxxing and the Physics Bill — How the AI Compute Mania Is Repriced at Your Grid, Your Tap, and Your Subscription](https://archtenet.dev/blog/tokenmaxxing-physics-bill.md): TL;DR: The software industry has gamified compute at the application layer while the physical layer hits hard limits. Three global chokepoints — hyperscale-captured memory supply, congested grids from Virginia to Dublin to Singapore, and… - [Emergent Creativity: An Architectural View on AI Consciousness and Deception](https://archtenet.dev/blog/ai-emergent-creativity.md): Early artificial intelligence research, which started in the early 1950s, split into two distinct architectural paradigms. The first was a logic-inspired approach that attempted to hard-code intelligence using symbolic expressions and pr…