Skip to main content

14 posts tagged with "architecture"

View All Tags

Spec-Driven Development Is Simpler Than You Thought

· 8 min read
Ivan Baha
Software Team Lead & Architect

Spec-driven development usually arrives as a framework: a sequence of phases, a command for each, templates, a governing document, and the assumption that one developer drives one agent through one repository. Seen that way, it looks like a heavy process to adopt – and for a team with its own tracker, its own roles, and dozens of services, adopting it that way is heavy and often unnecessary.

Under the packaging, SDD is a handful of ideas. My team runs it without a framework. I designed the process around those core ideas; the team adopted it, and five months later it is simply how we work – a team lead, two engineers, a PO/BA and two QA engineers, responsible for 80+ microservices. Agents write every document; people decide, direct each stage, and take responsibility for the result. Two things made it simple: an AI-native meta-repo, where the agent sees the whole system at once, and a triage step that keeps most work away from specifications altogether.

It isn't free. A home-grown process has no community or upgrade path behind it. It fits one team's context rather than every team's, and it leans on the workspace – without one, it becomes considerably heavier.

This piece is the argument, with my team's version as the example. The mechanism, decision by decision, is in RA-005.

A Trace Id in Every Log Line: How Cross-Service Debugging Went From a Four-Engineer Session to One Grafana Query

· 14 min read
Ivan Baha
Software Team Lead & Architect

In 2022, I added a trace id to the two shared libraries every service my team owned already used – the logger and the HTTP client – so that one query in Grafana would return every log line of one request, across every service it touched. The platform is an enterprise system built as microservices from day one; it runs more than 200 of them today across several teams. My team owns 80+ of those. The trace id now covers more than 100 – more than we own – because other teams adopted it.

Before that change, a cross-service bug on a test environment was a meeting: a QA engineer, a business analyst, the team lead, usually a developer or two, sometimes DevOps – several hours of reading logs by timestamp and reconstructing the call chain from memory of the code. After that, the same investigation became a text field on a Grafana board, and the person typing into it was most often the analyst, not an engineer. Minutes rather than hours, and in most cases no engineer at all.

The cost is stated up front because it matters. There are no span IDs. Which service called which is inferred from a forwarded User-Agent and timestamps rather than being known. That was the trade.

This is a field report: what debugging looked like, what changed, and two cases as they happened. The mechanism is documented as a reference architecture, RA-004, and reconstructed as runnable code in team-workspace; this article explains it only as far as the story needs. If you read one section, read the two cases.

Two flows: why a second, simpler model runs beside the primary

· 9 min read
Ivan Baha
Software Team Lead & Architect

The most expensive thing my agent can do is ask its smartest model a trivial question.

That sounds backwards – trivial questions are supposed to be cheap. A two-line summary, a yes/no judgement, a bit of text cleanup: milliseconds of honest work for a competent model. But Azek runs on one machine, and on one machine the flagship model's attention is the scarcest resource in the system. Every small job it handles personally gets paid for twice: once in the GPU it occupies, and once in a currency most people never see billed.

So Azek – the self-hosted agent from the first post – runs two models side by side, on purpose. The big one talks to me. The small one does everything else. This post is about the second lane: why it exists, what drives on it, and why the split turned out to be one of the best decisions in the whole project.

The Hidden Complexity of "Easy" Geo-Sharding in MongoDB

· 8 min read
Vladyslava Prykhodko
Engineering Technical Lead & Architect

Geo-sharding is one of the features that makes MongoDB attractive for a global product. You keep each user's data physically close to them, one logical collection still reads as a single dataset in Compass or Atlas, and the adoption cost looks tiny: add a shard key, and queries route to the right region.

The simplicity is real. So is the failure mode hiding inside it. This is a field story about how "easy" geo-sharding quietly doubled a database bill — and why that was never MongoDB's fault.

ABAC in Production: The Migration — DaemonSet, a Node-Local Attribute Cache, and Policy Delivery from Git

· 8 min read
Vladyslava Prykhodko
Engineering Technical Lead & Architect

Part 2 showed that OPA/OPAL sidecars hold down about a third of the cluster and, on top of that, force you to under-provision headroom. This part is about the migration itself: how to pull the PDP out of every pod without losing data freshness or hitting a latency wall.

The properties of externalized authorization don't change — single policy plane, no drift, tamper-resistant audit log. What changes is where the PDP lives and how policy and data reach it.

Everything below is illustrative and generalized — a reference model, not data or code from any specific system. Substitute your own.

Externalized Authorization Across a Hundred Services: What OPA/OPAL Sidecars Actually Cost

· 11 min read
Vladyslava Prykhodko
Engineering Technical Lead & Architect

Part 2 of a series on ABAC in production. Part 1 covered in-process guards and why they're enough for most teams. This part is about what happens when they stop being enough, and what the bill looks like.

Everything below is illustrative and generalized — a reference model, not data or code from any specific system. Substitute your own. The method is what matters: measure your own numbers the same way, and the proportions will likely hold.

The setup​

Picture a typical modern platform: microservices on something like NestJS, a database, deployed to managed Kubernetes (EKS/GKE/AKS — doesn't matter), roughly a hundred to a hundred and fifty services, zero-trust model. Every user request fans out through 5–15 services, and every service-to-service call is authorized independently — no trust at the perimeter.

Authorization lives outside the applications: an ABAC policy engine (OPA) makes the decisions, and OPAL keeps policies and data in sync. User attributes sit in the database and are editable from a UI. A textbook externalized-authorization setup.

Sooner or later a simple question comes up that you usually have no number for: what does this cost? Not "is OPA expensive in principle," but concretely — how much CPU and how many dollars go into the authorization layer itself. This article is about how to measure that, and why the number turns out to be somewhere other than you'd expect.

In-Process Authorization with Guards: The Default That's Enough Until It Isn't

· 6 min read
Vladyslava Prykhodko
Engineering Technical Lead & Architect

Once you have more than a handful of services, "can this caller do this thing to this resource?" stops being a one-liner. The answer usually depends on attributes — who the subject is, what they own, which tenant they belong to, what action they're attempting, sometimes the time of day. That's attribute-based access control, ABAC: the decision is a function of subject, resource, action, and context, rather than a flat list of roles.

The interesting question isn't whether to do ABAC. It's where the decision gets computed. There's a whole spectrum. This article is the cheap, simple end of it — and most teams should start, and often stay, here. Part 2 measures what the externalized alternative actually costs once you genuinely need it.

Everything below is illustrative and generalized — a reference model, not data or code from any specific system. Substitute your own.

Claude Fable 5: A Breakthrough in Cybersecurity or a Comedy of Guardrails?

· 4 min read
Ivan Baha
Software Team Lead & Architect

Two days ago, Anthropic launched Claude Fable 5, pitching it to the public as a breakthrough in cybersecurity capability. Naturally, I put it to the test on some of my own codebases: a lightweight lib for NestJS apps (nestjs-env-getter) and a highly complex, 40k+ LOC private AI agent orchestrator. The results weren't just disappointing – they exposed a structural comedy of guardrails that completely breaks the promise of AI-driven defence for regular engineers.

Edge-Native RAG — When the Small Stack Is the Right Stack

· 13 min read
Vladyslava Prykhodko
Engineering Technical Lead & Architect

Every AI chatbot has the same problem: it forgets you.

You tell it your name in one conversation, and by the next one it has no idea who you're talking to. You tell it your preferences, and a week later it's making things up. The bigger you build, the worse this gets — and at some point the bill from your LLM provider stops being funny too.

The fix has a name: RAG, short for retrieval-augmented generation. You give the model a memory it can search before it answers. The catch is that the standard way to build it is overcomplicated — a vector database, an embedding service, a chat model API, and a server somewhere to glue them all together. Four services, four monthly bills, four different things that can wake you up at 3 AM. Three months in, you're paying close to $200 a month for what is, honestly, a chatbot that remembers things.