Externalized Authorization Across a Hundred Services: What OPA/OPAL Sidecars Actually Cost
Part 2 of a series on ABAC in production. Part 1 covered in-process guards and why they're enough for most teams. This part is about what happens when they stop being enough, and what the bill looks like.
Everything below is illustrative and generalized — a reference model, not data or code from any specific system. Substitute your own. The method is what matters: measure your own numbers the same way, and the proportions will likely hold.
The setup
Picture a typical modern platform: microservices on something like NestJS, a database, deployed to managed Kubernetes (EKS/GKE/AKS — doesn't matter), roughly a hundred to a hundred and fifty services, zero-trust model. Every user request fans out through 5–15 services, and every service-to-service call is authorized independently — no trust at the perimeter.
Authorization lives outside the applications: an ABAC policy engine (OPA) makes the decisions, and OPAL keeps policies and data in sync. User attributes sit in the database and are editable from a UI. A textbook externalized-authorization setup.
Sooner or later a simple question comes up that you usually have no number for: what does this cost? Not "is OPA expensive in principle," but concretely — how much CPU and how many dollars go into the authorization layer itself. This article is about how to measure that, and why the number turns out to be somewhere other than you'd expect.
First — why not guards
If you have a single framework and your authorization is domain-local, in-process guards are the right answer. The decision is computed where the data already is, there's no network hop, no extra component. Don't over-engineer it.
Externalized authorization — a separate component that makes the decision — pays off when at least one of these is true:
- A polyglot estate. Guards in NestJS are TypeScript. The moment a Go or Python service shows up, you reimplement authorization logic per language. An external PDP behind Envoy ext-authz gives you one policy language across the whole zoo. This is the most concrete trigger.
- Cross-team policy governance. Security owns the policy centrally instead of hoping that every team reimplements tenant isolation, deny-by-default, and break-glass the same way.
- Compliance and audit. You need a single, tamper-resistant decision log of who-accessed-what across the whole fleet. Guards each log differently.
- Changing policy without a redeploy. A permission gets revoked org-wide in minutes, not by shipping code across every repository.
- A no-drift guarantee. A service physically doesn't contain the decision logic, so it can't "accidentally" start deciding differently after the next commit.
To keep the vocabulary straight, the XACML model is handy — it separates four roles that the guard approach fuses into one:
- PEP (Policy Enforcement Point) — intercepts the request and applies the verdict. Usually Envoy in the pod.
- PDP (Policy Decision Point) — computes allow/deny. This is OPA.
- PAP (Policy Administration Point) — where policies are stored and edited. For example, Git → bundles.
- PIP (Policy Information Point) — where the data for a decision comes from. The attribute store.
The whole difference between a guard and OPA is literally where the PDP sits relative to the PEP: in the same process, or pulled out.
The textbook answer: OPA + OPAL sidecars
The canonical layout for externalized authorization in Kubernetes is sidecars. Alongside each application pod you run:
- the PEP — Envoy, which intercepts inbound traffic and makes a gRPC ext-authz call;
- the PDP — OPA with the Envoy plugin;
- the sync agent — OPAL client, which pulls policy and data updates and keeps the local OPA current.
An OPAL server (one per cluster) watches Git and the data source and pushes updates to the clients. Every decision is made locally, in the pod — no network hop to a remote service for a verdict. Latency-wise, this is ideal.
The price of that elegance is a multiplier. The sidecar stack rides along in every replica of every service. And that's where it gets interesting.

A reality check: reserved versus used
In Kubernetes, cost is driven not by what a container uses but by what it reserves (requests) — because the scheduler packs against reservations, and node capacity is held against them. For a component in the synchronous hot path, like the OPA sidecar, the reservation is also set statically: you can't shrink it under load, because an HPA lag in the middle of a request is unacceptable. The reservation is fixed, per-pod, and independent of whether the pod is busy or idle.
Measure your sidecars' actual consumption and you'll almost certainly see this picture. There are two values to keep apart: request — how much CPU the container reserved (this is what you pay for) — and used — what it actually consumes.
First, the thing that trips people up: the reservation is roughly the same on every service even though their load differs wildly. That happens because the sidecar's request isn't tuned per service — it's written once in a shared template (Helm chart / mutating webhook) and injected identically into all ~300 pods. Nobody hand-tunes 300 sidecars. And the default is set with headroom, because OPA sits in the synchronous hot path: if it's starved of CPU during a spike, every request through that pod stalls, and each request triggers 5–15 decisions, so the stall cascades. In other words, you reserve to survive the peak, not for the average moment. The values below are illustrative, but the order of magnitude is typical:
| Container (role) | Reserved (request) | Actually used | Utilization |
|---|---|---|---|
| OPA sidecar (PDP) | 100m | ~6m | ~6% |
| OPAL client (sync) | 100m | ~20m | ~20% |
Read it like this: the reservation is ~100m, while OPA's real consumption is single-digit milliCPU. And crucially, that consumption doesn't move with load — a busy service's PDP burns about the same handful of milliCPU as an idle one's. A single policy decision costs a fraction of a millisecond regardless of how busy the application is; the high-traffic service burns its CPU in business logic (DB calls, serialization), not in its OPA. The cost of authorization is decoupled from application load.
From this follows the conclusion that's the meat of the whole thing: the waste is uniform across the entire fleet. The reservation is the same everywhere, usage is near zero everywhere, which means you're paying for an idle cushion in each of ~300 pods.
To get a precise number rather than a range, sum the CPU reservations across both authz containers fleet-wide in one pass:
kubectl get pods -A -o jsonpath='{range .items[*]}{range .spec.containers[*]}{.name}{"\t"}{.resources.requests.cpu}{"\n"}{end}{end}' \
| grep -E 'opa-envoy|opal-client' \
| awk '{gsub(/m/,"",$2); sum+=$2} END {print sum"m reserved across fleet"}'
For a representative platform of ~150 services at ~2 replicas each (≈300 pods), this easily lands around 30 vCPU reserved fleet-wide for authorization alone.
What that is in dollars
Take a cluster of ~10 general-purpose nodes at 8 vCPU each, roughly $250/node/month on-demand — the whole cluster is about $30k/year.
Authorization: ~30 vCPU ÷ 8 ≈ 3.7 nodes of reservation × $250 ≈ $940/month ≈ $11k/year.
So the authorization sidecars hold down about a third of the production cluster — while the PDP itself spins single-digit milliCPU. The OPA sidecar is reserved at roughly 5–10× what it actually uses. That's the reserved-vs-used gap, now with a concrete figure on it.
(Numbers are on-demand list price as of June 2026; with Savings Plans / RIs the absolute drops, but the proportions hold.)
What moving to a DaemonSet buys
The obvious next move is to pull the stack out of every pod and run one OPA per node (DaemonSet). Pods on a node talk to the node-local PDP. The instance count drops to ~10 instead of ~300.
There's a trade-off to acknowledge: this requires routing traffic via the node's host IP (using the Kubernetes Downward API, status.hostIP), which slightly alters the network boundary compared to standard localhost sidecars — the authz call now traverses the node's internal network instead of staying inside the pod, something to weigh in a strict zero-trust setup.
Because the aggregate decision traffic is tiny (one operation per second at a 5–15 fan-out = single-digit decisions per second across the whole cluster), one OPA per node has enormous headroom. You can raise the per-instance reservation to cover the node's aggregate and still win by an order of magnitude:
| Topology | Σ reservation | ≈ nodes | $/year |
|---|---|---|---|
| Sidecar | ~30 vCPU | ~3.7 | ~$11k |
| DaemonSet (10 × 350m) | ~3.5 vCPU | ~0.4 | ~$1.3k |
| DaemonSet (10 × 500m, with headroom) | ~5 vCPU | ~0.6 | ~$1.9k |
Savings: on the order of $9–10k/year, or −85…88% of authorization cost. About 3 nodes of reservation freed out of 10.
There's a hidden bonus too, invisible in the reservation math: hundreds of OPAL-client copies each run real polling (~20m apiece), which is several vCPU of live work that collapses down to the number of nodes. After the migration, nodes pack tighter than the requests arithmetic alone suggests.
What the move actually buys you
The temptation is to put "we halved the cluster" in the headline. That's not true: it shrinks by about 3 nodes out of 10, roughly a third. But the opposite mistake is just as easy — dismissing a third as small change. So let's be precise, along three axes.
1. The savings are a third of the production cluster, or −85% of authorization cost. As a proportion, that's a large win that justifies the migration on its own. In absolute terms the magnitude depends on the size of your bill: on a compact cluster it's tens of thousands a year, on a large one proportionally more. Don't confuse "a small dollar amount" with "a small fraction" — the fraction here is not small.
2. Resilience goes up, not down — and this may be the main point. In the sidecar model, every extra milliCPU of reservation is multiplied by ~300 pods. So you're boxed in: if you want to give OPA a cushion for spikes, you pay for it 300 times, which forces you to cut the reservation to the bone and balance on the edge of throttling at peak. Headroom there is an unaffordable luxury.
In a DaemonSet, the same reservation is multiplied by ~10. The cushion is thirty times cheaper. Which means you can give each node-local OPA a generous buffer — survive any spike, never hit a CPU wall in the hot path — and still spend several times less than the skinny sidecar did. You simultaneously reclaim a third of the cluster and make the authz layer more resilient than it was. Add statistical multiplexing on top: the load profile across 10 node-aggregates is smoother than across 300 scattered pods, so the same cushion covers a peak more reliably.
That's the real point. The move doesn't change "how much you pay for the same thing" — it changes the trade-off itself: before, you chose between cost and headroom; now you take both.
3. The operational properties come on top, not instead. A single policy plane, no behavioral drift between services, decoupling of policy and attribute distribution from each team's release cycle. The sidecar gave you all that at the cost of a third of the cluster; the DaemonSet gives it for ~5–6% — and with more headroom.
If you run externalized authorization and you haven't measured it, take your own number with the same one-liner. Chances are you're reserving the PDP at many times what it uses across the whole fleet, and cutting headroom thin because every milliCPU of cushion is multiplied per pod.
What's next
Part 3 covers the migration itself: a DaemonSet plus a node-local attribute cache (Valkey/Redis) as the distribution layer, plus a control plane that ships policy bundles from Git, and why attribute-change notifications go over an event bus rather than directly from the owning service. I'll also work through the mechanics of how OPA gets fresh data without holding the entire dataset in every instance's memory.
