Appropriate Caching in a Multi-Service Estate (Balancing Efficiency and Simplicity Across a Shared Store)
1. Executive Summary
In microservice architectures, caching typically polarises into two unsatisfactory archetypes:
- Isolated Per-Service Caches: Each service maintains its own cache tier or in-memory heap. When Service A needs data owned by Service B, it must make a network hop via HTTP or gRPC to Service B, even if Service B has just cached the exact record. Invalidation does not propagate across replicas; memory is multiplied N-fold; and data inconsistencies between pods remain undetectable.
- Coordinated Invalidation Buses: Services share a distributed cache or keep local caches synchronised via a pub/sub message broker (Redis Pub/Sub, Kafka, RabbitMQ). While theoretically reactive, this adds a secondary distributed system whose failure modes — message drops, consumer lag, split-brain states, and out-of-order race conditions — directly compromise cache consistency.
This Reference Architecture presents a pragmatic middle ground, validated in production across more than 40 microservices: a single shared cache server operated under strict single-writer ownership with direct consumer reads.
Rather than coordinating invalidation between peers, the architecture enforces a rigid ownership contract within the key namespace itself: {service}_{entity}_{id}. Exactly one service — the domain owner of the data — writes to and invalidates that key. Consumers read the owner's keys directly from the cache. A cache hit eliminates the inter-service HTTP hop entirely: no outbound request serialisation, no inbound middleware evaluation, no trace fan-out, and no database query. On a cache miss, the consumer makes a normal HTTP call to the owner; the owner's read-through path populates the entry; and subsequent reads by any consumer immediately hit.
To ensure this model remains resilient, maintainable, and cost-effective, the architecture is underpinned by three core design principles:
- Server-Enforced Access Control (ACLs): Consumers are granted read-only permissions on exactly the foreign entities they consume, at the cache server level (Valkey ACLs). A consumer bug can never write or delete another service's cache entries — and because every key is built from the identifier exactly as its loader sees it, no request can poison an owner's entry through the owner's own read path either.
- Granular Operation-Specific Failure Policies: Reads fail open (returning
null, protecting user-facing traffic when the cache is down — or merely silent, because every command has a deadline); queue writes fail closed (throwing explicit errors to prevent silent data loss); and configuration bugs fail loud at startup. - The Caching Ladder (Appropriate Complexity for Each Shape): The team does not apply a blanket caching decorator. Read paths are categorised onto a ladder of rungs — Entity, Bounded List, and Composite Request. Each rung uses the simplest invalidation strategy it can honestly guarantee: exact 1-to-1 eviction for entities, blunt wholesale invalidation for bounded lists, and owner-side eviction plus a short self-healing TTL for composites. On every rung, the entry's TTL is the backstop that bounds staleness when an invalidation is raced or lost.
Finally, the same cache server acts as an operator execution tier: Redis Sets provide deduplicating background work queues, atomic SET NX EX keys provide self-releasing batch locks, and the cache engine's volatile-lru eviction policy ensures that memory pressure evicts ephemeral cached data while never touching queued work — which is also why the queue is capped where work enters it.
Provenance. The author designed this architecture for a team responsible for 80+ microservices on an enterprise platform of 200+. It has since spread beyond that team: more than 40 microservices across several teams now use it in production. The linked repository is a self-contained reference implementation of the library, contracts, services, and GitOps manifests described here, not a direct proprietary extract.
2. Applicability Criteria
Use this pattern when
- Inter-service read amplification is high. Services frequently query peer services for reference or relational data (e.g., resolving user profiles, product catalogs, permissions) across HTTP/gRPC boundaries.
- The team manages multiple microservices sharing an infrastructure tier. Maintaining isolated caching infrastructure (or separate Redis clusters) for each service imposes disproportionate operational and financial overhead.
- A pub/sub invalidation bus is over-engineering. The write volume and latency tolerance allow direct read-through and asynchronous owner deletion without the complexity of message brokers.
- Operations require predictable failure modes. A cache outage must never take down primary business flows, and background jobs must never be acknowledged without being stored. (If accepted jobs must also survive a restart of the cache server itself, keep the queue on durable infrastructure — see §8.2.)
- Background batching and coordination are needed without extra queuing brokers. Work can be deduplicated and debounced using native Redis primitives (Sets and self-releasing locks) without provisioning Kafka, RabbitMQ, or SQS.
Do NOT use this pattern when
- Strong linearizable ACID read consistency is required. An owner's write makes its entries fresh immediately in the common case, but a read racing that write, or an invalidation dropped during a connection blip, can leave a stale value in place for up to one entry TTL. If every consumer read must see transactional updates immediately, queries must hit the authoritative transactional database directly.
- Tenant data must be physically or cryptographically air-gapped. When strict multi-tenant compliance prevents sharing an infrastructure network namespace or cache process across service domains.
- Write frequency matches or exceeds read frequency. High-churn datasets will suffer near-zero hit rates while incurring continuous invalidation overhead.
- Payloads exceed multi-megabyte boundaries. Storing massive JSON blobs or unbounded collections in single keys degrades Redis single-threaded performance and triggers latency spikes.
- One participant's volume would dwarf the rest. ACLs isolate who writes which keys, not how much memory each participant takes. A service whose working set crowds the shared
maxmemoryevicts everyone else's entries; give it its own instance.
3. Architecture Overview
3.1 The Three Foundational Principles
- The Key is the Ownership Contract: Every key is prefixed with the owning service's deployment name:
{service}_{entity}_{id}. The prefix dictates which service writes it, which service invalidates it, and which ACL rules apply. The identifier is used exactly as its loader sees it — validated, never rewritten. - Appropriate Complexity (The Ladder): Avoid "clever" partial invalidation. Use exact invalidation where relationships are 1-to-1, wholesale invalidation where relationships are 1-to-many, and clock-bounded TTLs for whatever crosses service boundaries. On every rung the TTL is also the backstop: it bounds how long a raced or dropped invalidation can leave a stale value.
- Symmetric Behavioral Parity: Development, test, and production environments execute identical serialization and outage semantics. In-memory fallbacks serialize on write and parse on every read, exactly like the server, and delete type-agnostically. The one difference no fallback can emulate — the real client is still connecting when the application boots — is designed around: nothing depends on a check made at startup.
3.2 Constraints Decide the Shape
Every mechanism in this architecture traces back directly to an operational constraint:
| Constraint | Architectural Decision | Section |
|---|---|---|
| Inter-service HTTP calls inflate latency and trace fan-out | Consumers read owner cache keys directly from the shared server | §3.3 |
| A bug in a consumer must not corrupt another service's data | Valkey ACL selectors grant consumers read-only access to exactly the entities they consume | §6.3 |
| A request must not be able to write another identifier's entry | Identifiers are validated, never trimmed or lowercased; non-canonical spellings are served uncached | §4.1 |
| Cache server downtime must not take down user-facing traffic | get operations fail open, falling back to the source of truth | §4.2 |
| A connected but silent server must not hang requests | Every command has a deadline; three timeouts in a row mark the cache down and force a reconnect | §4.2 |
| Enqueued background tasks must never be acknowledged without being stored | addToSet fails closed, throwing CacheUnavailableError (HTTP 503) | §4.2, §4.8 |
| Non-existent entities cause database 404 hammering | Absence is negative-cached using a sentinel value ("not_present") | §4.3 |
| Concurrent misses for the same key cause database stampedes | In-process single-flight promise coalescing joins concurrent requests on each pod | §4.3 |
| Complex list invalidation branches lead to silent stale reads | Any write to an entity wipes all list shapes wholesale (delMany) | §4.5 |
| Multi-service composite responses cannot be invalidated by foreign owners | Composites use canonical keys, owner-side eviction, and short self-healing TTLs | §4.6 |
| Background jobs need deduplication without extra brokers | Redis Sets (SADD/SPOP) and SET NX EX locks provide native queuing | §4.8 |
| Every accepted background job must get a drain | The drain releases its lock and re-checks the queue; a periodic reconcile catches the rest | §4.8 |
| Memory pressure must not evict queued jobs — nor let the queue evict everything else | --maxmemory-policy volatile-lru; the queue carries no TTL and is capped at enqueue | §6.1, §4.8 |
3.3 End-to-End Interaction Flows
Consumer Direct Read & Fallback Flow
The following sequence depicts products-service (consumer) resolving a product's owner via the shared cache, eliminating an HTTP hop to users-service (owner):
Operator Work Queue & Batch Lock Flow
The following sequence illustrates products-sync-service leveraging shared cache primitives for deduplicated background recalculations:
4. Implementation Details
The implementation spans one core library (@tw/cache) and three domain services in the reference workspace (core files shown):
libs/tw-cache/
├── src/
│ ├── cache-keys.ts # Central registry & cacheKey()/cacheKeyOrNull() builders
│ ├── request-key.ts # Canonical requestKey() serializer
│ ├── cache.service.ts # Policy layer, failure handling, timeout breaker, health counters
│ ├── read-through.service.ts # Single-flight, negative caching, no-cache supersession
│ ├── redis-cache.store.ts # ioredis client (offline queue off, per-command timeout)
│ ├── memory-cache.store.ts # JSON-parity memory store for local development
│ ├── cache.module.ts # Module registration & boot-time config validation
│ ├── validate.ts # Input validation (assertTtlSeconds, assertCommandTimeoutMs)
│ ├── errors.ts # Typed errors (CacheUnavailableError, etc.)
│ └── test/cache.fake.ts # Standalone behavioural test double
backend/
├── users-service/ # Owner: user entity
├── products-service/ # Owner: product, productList, catalog composite; Consumer: user
└── products-sync-service/ # Operator: background queue, lock, reconcile loop, categoryStats
infra/git-ops/base/cache/
├── deployment.yaml # Valkey 8: volatile-lru, persistence off, Recreate
└── acl-configmap.yaml # Per-service ACL rules: hashed passwords, per-entity consumer grants
4.1 Strict Key Registry & Identifier Validation
Keys are never concatenated via raw string interpolation ('users-service_user_' + id). Raw concatenation risks casing drift, missing segments, and silent ACL permission denials. Even the operator's queue and lock keys, which carry no id, take their prefix from the registry.
Entity Key Construction
CACHE_NAMESPACES in libs/tw-cache/src/cache-keys.ts acts as the single source of truth for all cached entities:
export const CACHE_NAMESPACES = {
user: { service: 'users-service', entity: 'user' },
categoryStats: { service: 'products-sync-service', entity: 'categoryStats' },
product: { service: 'products-service', entity: 'product' },
productList: { service: 'products-service', entity: 'productList' },
productCatalog: { service: 'products-service', entity: 'req' },
} as const;
The cacheKey(name, id) builder validates identifiers against /^[a-z0-9][a-z0-9._-]*$/ as given — it never trims or lowercases them.
That last point carries more weight than it appears to. A key must be a function of exactly what the loader behind it sees. A builder that normalized identifiers would give ' 1' the key of '1' (and Widgets the key of widgets) while their loaders — matching the raw value against the source of truth — return different answers. Whichever loader filled the key first would then answer for both: a single GET /v1/users/%201, sent with Cache-Control: no-cache, would store "does not exist" under user 1's key, and every consumer reading that key would serve the 404. Normalizing in the builder turns a harmless malformed request into cross-service cache poisoning; validating turns it into an uncached read.
The same rule covers synthetic shapes. The unfiltered product list is cached as the shape all and a category as the shape category.<name>, so no caller-supplied category can ever produce the synthetic shape's key.
Safe Validation for Data (cacheKeyOrNull)
When an identifier comes from data — a path parameter in GET /v1/users/:id, but equally a stored column or connector data — an invalid format (spaces, control characters, a non-canonical spelling) must not throw a 500 error. Instead, cacheKeyOrNull() returns null:
const key = cacheKeyOrNull('user', id);
const user = await (key ? this.readThrough.readThrough({ key, load }) : load());
On a read, the application falls back to the authoritative store, which answers for exactly the identifier it was asked about — an HTTP 404 rather than an internal cache failure. On a write, null means there is nothing to invalidate: an identifier that cannot form a key never had an entry, because every read of it went uncached. The throwing cacheKey() is reserved for identifiers the code itself constructs; a throwing builder on a write path fails a write the database has already committed.
Consumers apply the same rule to the foreign ids they hold. A product's owner id came from a request body, so one that cannot form a key skips the cache read and goes straight to the owner over HTTP-URL-encoded, since it is data, not a path — instead of failing every catalogue request that includes that product.
4.2 Granular Failure Policy: Fail-Open vs. Fail-Closed vs. Fail-Loud
The client library encapsulates the failure matrix, so individual developers cannot inadvertently introduce unsafe behaviours:
// libs/tw-cache/src/cache.service.ts
async get<T>(key: string): Promise<T | null> {
if (this.unreachable) {
this.skipped++;
return null;
}
try {
const value = await this.run(() => this.store.get<T>(key)); // run() counts consecutive timeouts
if (value === null) this.misses++;
else this.hits++;
return value;
} catch (error) {
this.skipped++;
this.logger.warn(`Cache get for ${key} failed, failing open: ${(error as Error).message}`, CONTEXT);
return null;
}
}
async addToSet(key: string, members: string | string[]): Promise<number> {
const items = Array.isArray(members) ? members : [members];
if (this.unreachable) {
throw new CacheUnavailableError(
`Cache unreachable, refusing to add ${items.length} item(s) to ${key}: a set add is a work-queue write, ` +
'and reporting it accepted would lose the work.',
);
}
// ...
}
- Boot-time Validation: Any non-integer or
NaNTTL (Number(process.env.CACHE_TTL)) immediately throwsCacheConfigurationError, crashing the process during bootstrap before accepting traffic. TTLs a service reads for itself — such as the composite TTL — are checked the same way when their module loads (assertTtlSeconds), not once per request. - Offline Queue Disabled: The Redis client initializes with
enableOfflineQueue: false. Commands reject instantly when disconnected, instead of queueing behind a reconnect and applying later against changed state. Commands already in flight when a connection drops are failed rather than re-sent after the reconnect (autoResendUnfulfilledCommands: false) — a re-sentSETis a fill that lands after whatever invalidation happened in between. - A Deadline on Every Command: Connection events report a server that has gone away — an error, or the clean
closea server sends when it shuts down gracefully. A server that is connected but silent — a hung process, or a lost node whose TCP connection nothing has closed yet — emits nothing, and without a deadline every read would wait until the kernel gives up on the socket: minutes. Each command therefore carries a timeout (commandTimeoutMs, 250 ms by default). When it expires, the command fails the way it would against an unreachable server — reads open, queue writes closed — and three timeouts in a row mark the cache unreachable and force a reconnect, so requests stop paying the timeout one by one. - Unknown Outcomes Are Refused, Not Guessed: A queue write that times out may still have reached the server. The endpoint answers 503 anyway, and the retry it asks for is safe: a set add is idempotent.
- A Clean Shutdown, Even Mid-Outage: On shutdown the client sends
QUITand then disconnects outright. While the connection is down,QUITis refused (there is no offline queue to hold it), and a client left to itself would keep reconnecting — holding a stopping pod open until it is killed. A service's own shutdown is not reported as a cache outage, either.
4.3 Read-Through, Negative Caching & In-Process Single-Flight
ReadThroughService eliminates the boilerplate check-load-store pattern and implements three protections:
// libs/tw-cache/src/read-through.service.ts
export class ReadThroughService {
private readonly inflight = new Map<string, Promise<unknown>>();
async readThrough<T>(options: ReadThroughOptions<T>): Promise<T | null> {
assertTtlSeconds(options.ttlSeconds, 'ttlSeconds');
assertTtlSeconds(options.negativeTtlSeconds, 'negativeTtlSeconds');
const { key, load, traceId } = options;
if (!options.noCache) {
const cached = await this.cache.get<unknown>(key);
if (cached !== null) {
const negative = cached === NEGATIVE_CACHE_SENTINEL;
this.logger.debug(JSON.stringify({ cache: negative ? 'negative-hit' : 'hit', key }), CONTEXT, traceId);
return negative ? null : (cached as T);
}
this.logger.debug(JSON.stringify({ cache: 'miss', key }), CONTEXT, traceId);
} else {
this.logger.debug(JSON.stringify({ cache: 'bypass', key }), CONTEXT, traceId);
}
// A no-cache read never joins: the flight under way may have read the source before the write
// the caller is trying to see.
const existing = options.noCache ? undefined : this.inflight.get(key);
if (existing) {
this.logger.debug(JSON.stringify({ cache: 'joined', key }), CONTEXT, traceId);
return existing as Promise<T | null>;
}
// A flight superseded by a later no-cache read still answers its callers, but does not write.
const flight: Promise<T | null> = this.loadAndStore(options, () => this.inflight.get(key) === flight).finally(
() => {
if (this.inflight.get(key) === flight) this.inflight.delete(key);
},
);
this.inflight.set(key, flight);
return flight;
}
}
- Negative Caching: When
load()returnsnull(proven absence in the database), the service writesNEGATIVE_CACHE_SENTINEL("not_present") fornegativeTtlSeconds(default: 30s). Future requests are answered immediately without hitting the database. Loaders that fail or time out must throw, ensuring infrastructure errors are never cached as negative hits. The sentinel is a plain string value, so a namespace whose legitimate values could equal it must not use negative caching. - Single-Flight Coalescing: If 50 concurrent requests miss the same key simultaneously on a single pod, only one database query is dispatched; the other 49 callers join the active Promise. The coalescing is per process, so with N replicas a cold key can still cost N loads — the protection bounds a stampede to one load per pod rather than eliminating it. Nor does it remove the latency spike when a hot key expires: every waiter still waits for that one load. Probabilistic early expiration (Vattani et al.) addresses that part, by recomputing a hot key shortly before it expires; the two techniques are complementary, and this architecture has not yet needed the second.
- No-Cache Supersedes, Never Joins: A
no-cacheread starts its own load and becomes the key's current flight. The flight it superseded still answers its own callers but no longer writes the entry, so an older value from the same pod cannot land on top of the repair. A stale fill from another replica still can; the entry's TTL bounds it (§4.4).
4.4 Owner Write Invalidation: Unconditional, From Every Write Path
Writes must invalidate the cache immediately and unconditionally:
// backend/users-service/src/users/users.service.ts
private invalidateUser(id: string): void {
setImmediate(() => {
// An id that cannot form a key was never cached — every read of it went uncached.
const key = cacheKeyOrNull('user', id);
if (!key) return;
void this.cache.del(key).then((removed) => {
this.logger.debug(`Invalidation removed ${removed} cache entry for user ${id} (key ${key})`);
});
});
}
Key principles of the invalidation contract:
- Off the Response Path: Invalidation is deferred via
setImmediate(), and the key is built there too. The caller's HTTP write request completes without paying the latency of a cache round trip, and nothing about building a cache key can turn a committed write into a 500. - Never Header-Gated:
Cache-Controlheaders govern read freshness; writes always invalidate. - Every Write Path: Invalidation is only as complete as the write paths that remember it. In the reference workspace, registration lives in a separate authentication module, so it creates users through the owner's
UsersService.create(), which invalidates the new id's key. A module, batch job, or migration that writes directly to the store is a write the cache never hears about. - The Delete Count as a Debugging Aid: Redis
DELreturns the integer number of keys actually removed, and the owner logs it on every write. After a read that filled the key,removed 0means the invalidation was built with a different key than the read — a "dead invalidation". On its own,0is also what every write to an uncached key gives, and what a dropped invalidation gives during an outage (the client logs each drop at debug, so the two can be told apart), so the count is a probe rather than an alarm; key drift is prevented by construction (one registry, one builder) and caught by tests.
How stale can an entry get? Usually not at all: the owner's write deletes it and the next read refills it. Three things can still leave an old value in place, and the design names the bound for each rather than pretending they do not happen:
- The deferred
DEL: a client that writes and immediately re-reads can land its GET before thesetImmediatedelete runs. - A dropped
DEL: deletes are best-effort; during a connection blip one is dropped, and if the cache comes back without restarting, the old entry is still there. - A stale fill: a read that loaded the database before the write committed can store its value after the write's
DEL— usually on another replica, where nothing in-process can see it coming. Single-writer ownership removes coordination between services; it does not remove this race inside the owner.
In all three cases, the entry's TTL is the bound: 60 seconds for entities and lists, 30 for negative entries. Clients that must read their own writes send Cache-Control: no-cache, which starts a fresh load instead of joining one already in flight. Closing the stale-fill race outright takes versioned or leased fills — the lease mechanism described in Scaling Memcache at Facebook — which this architecture defers until the TTL stops being a good enough bound.
4.5 List Caching: The Case for Blunt Wholesale Invalidation
List caching commonly fails when engineers attempt "smart" invalidation — tracking category memberships, pagination boundaries, and sorting permutations. This introduces extensive branching logic that fails silently when items move or empty.
This architecture enforces blunt wholesale invalidation:
// backend/products-service/src/products/products.service.ts
private invalidateAfterWrite(productId: string, extraCategories: string[] = []): void {
setImmediate(() => {
const itemKey = cacheKeyOrNull('product', productId);
const keys = [
...this.listKeys(extraCategories), // 'all' + 'category.<name>' per category; un-keyable ones skipped
...(itemKey ? [itemKey] : []),
requestKey('productCatalog', 'v1-products', { expandOwner: true }),
];
void this.cache.delMany(keys).then((removed) => {
this.logger.debug(`Invalidation removed ${removed} cache entries after write to ${productId}`);
});
});
}
- Every Shape Evicted: Any write, update, or deletion of a product invalidates
productList_alland every knownproductList_category.<name>shape via a singledelMany(). - Ghost Category Prevention: When an item is moved to another category or deleted, the old category is passed in
extraCategoriesbecause an empty category can no longer be named by scanning active store records. - Un-keyable Shapes Are Skipped, Not Fatal: A category that cannot form a key (
Home & Garden) is never cached — its reads go straight to the store — so its invalidation is skipped instead of throwing after the write has already committed. - Empty Lists are Valid Values: A query returning no records caches
[]as an authentic value.nullis reserved for proven entity absence. - Free-Text Search is Never Cached:
?search=shoesbypasses the cache and queries the database every time. An unbounded search space cannot be enumerated, making complete invalidation impossible.
4.6 Composite Request Caching: Clock-Bounded Invalidation
Composite endpoints aggregate data across domain boundaries (e.g., GET /v1/products?expandOwner=true, which joins products with user records from users-service).
Composite caching balances local vs. foreign ownership:
-
Canonical Serialization:
requestKey()sorts parameters by name and normalizes the route and parameter names — both written by code — but validates parameter values as given, rejecting anything that is not a plain token. A value is what the loader filters on; rewriting it would alias two different answers onto one key:products-service_req_v1-products_expandowner=true -
Split Invalidation Responsibility:
- Local Product Changes:
products-servicewrites invalidate the composite key directly, ensuring price or title changes are instantly visible. - Upstream User Changes:
users-servicehas no knowledge of downstream composites. Therefore, the clock is the invalidation. The composite carries a tight, self-healing TTL (e.g., 10 seconds).
- Local Product Changes:
-
Freshness Reaches the Parts: A caller's
no-cachebypasses the composite and the owner entries it is built from. Rebuilding a "fresh" composite from a part that outlived its own invalidation would re-cache exactly the value the caller was trying to get past. -
Opt-In Only: Composite caching is explicitly invoked in handler logic. Blanket middleware caching is prohibited to prevent accidental caching of authenticated or user-scoped payloads.
4.7 Active Freshness: End-to-End Traversal of Cache-Control: no-cache
When a client requires guaranteed fresh data, it transmits standard HTTP semantics:
Cache-Control: no-cache
This header executes a bypass-and-refresh:
- The read-through path bypasses the cache read — and never joins a load already in flight, which may predate the caller's write.
- It fetches the authoritative value from the database.
- It overwrites the cache entry with the fresh data, repairing the cache for all subsequent readers.
- Header Propagation: The internal HTTP connector (
@tw/http-connector) automatically forwardsCache-Controlon outbound inter-service calls, ensuring the freshness demand travels end-to-end to the authoritative owner — including through composites, whose parts are bypassed too.
This borrows a request directive that RFC 9111 defines for HTTP caches and applies it to the caches behind the origin, which makes it an amplification lever: browsers send no-cache on every hard reload and whenever developer tools disable caching, and each such request costs the owner a source read — for a composite, one per part. Who may send it is an edge decision: strip or rate-limit it on public traffic at the gateway, and keep it for internal callers, batch jobs, and operators, for whom it is the repair tool.
4.8 The Operator Pattern: Work Queue & Batch Lock
products-sync-service acts as an operator, maintaining asynchronous category aggregations without adding RabbitMQ or Kafka:
// backend/products-sync-service/src/recalculations/recalculations.service.ts
async requestRecalculation(categoryIds: string[]): Promise<{ queued: number; scheduled: boolean }> {
// The queue is never evicted, so it is capped where work enters it
const pending = await this.cache.getSizeOfSet(QUEUE_KEY);
if (pending + categoryIds.length > this.maxQueueSize) throw new QueueFullError(/* ... */);
// SADD deduplicates identical category IDs automatically
const queued = await this.cache.addToSet(QUEUE_KEY, categoryIds);
// Atomic SET NX EX locks the drain timer
const scheduled = await this.scheduleBatch();
return { queued, scheduled };
}
private async scheduleBatch(): Promise<boolean> {
const acquired = await this.cache.acquireLock(LOCK_KEY, this.lockTtlSeconds);
if (!acquired) return false;
setTimeout(() => void this.drain(), this.batchDelayMs);
return true;
}
- Validated, Bounded Input: Category ids are validated with the key builder's own rule — at most 100 per request, 64 characters each — and a request that would take the queue past
MAX_QUEUE_SIZEis refused with503 QUEUE_FULL. The set carries no TTL, so eviction never touches it — which is exactly why it must never grow without bound (§6.1). - Atomic Queue Draining:
SPOP queue batchSizeatomically extracts items. Multiple worker replicas can execute drains concurrently without ever receiving duplicate IDs. - Self-Releasing Lock, Released Early: The lock uses
SET NX EX <ttl>with a TTL longer than the batch delay, so a burst of requests collapses into one drain. When the drain finishes, it deletes the lock and re-checks the queue; without that step, a request arriving after the drain's finalSPOPwould find the lock still held, arm nothing, and wait for some later request. If a worker pod crashes mid-batch, the TTL releases the lock. - Pre-Drain Bulk Invalidation: Before executing expensive upstream HTTP calls, the batch runs
delMany()on all stats keys in the batch. Readers arriving during the calculation window encounter a cache miss and calculate on-demand, preventing stale reads. - Periodic Reconciliation: Every
RECONCILE_INTERVAL_SECONDS(30 by default), the service checksSCARD queueand, if the lock is free, schedules a drain. A one-off check at startup would not work: at bootstrap the cache client is still connecting, an unreachable cache reports the set as empty, and the check never looks again. An interval also catches work stranded while the service keeps running — a lock holder that crashed before arming its timer, a schedule swallowed by a cache outage, an evicted lock.
5. Strategy Selection: Choosing the Appropriate Approach
The fundamental premise of this Reference Architecture is matching the caching mechanism to the access shape. Caching every endpoint identically creates brittle systems.
5.1 The Caching Decision Matrix
| Access Pattern | Target Shape | Strategy / Rung | Invalidation Trigger | Staleness Bound | Failure Policy |
|---|---|---|---|---|---|
Lookup by ID (/users/:id) | Single Entity | Rung 1: Entity Cache | Exact key DEL on owner write | Fresh after the owner's write in the common case; ≤ entity TTL (60s) when an invalidation is raced or dropped | Fail open (null) |
Filtered Query (/products?category=x) | Bounded Hydrated List | Rung 2: Bounded List Cache | Wholesale delMany of all shapes on write | Fresh after the owner's write in the common case; ≤ list TTL (60s) when an invalidation is raced or dropped | Fail open (null) |
Cross-Domain Composite (/products?expandOwner=true) | Multi-Service Joined Object | Rung 3: Composite Request Cache | Local write DEL + Short TTL (10s) | ≤ 10 seconds for foreign updates, on top of the owner entries' own bound | Fail open (null) |
Arbitrary Free-Text (/products?search=text) | Dynamic Filter | Direct Store Query (Uncached) | None (bypasses cache) | None — read from the store | N/A |
| Background Aggregation (Category Stats) | Materialized Aggregate | Operator Worker Pattern | Pre-drain delMany + Background Set Queue | TTL (300s) between recalculations; batch inputs fetched with no-cache | Fail closed (503 on queue write) |
5.2 Rules of Thumb for Balancing Simplicity vs. Efficiency
The decision follows ownership first, then shape:
- Prefer Lower Rungs: If a query can be satisfied by assembling individually cached entities on the client or consumer side, do not introduce a composite request cache.
- Never Cache Identity-Varied Responses: If a response body varies depending on the authenticated caller's identity, permissions, or session attributes, it must never be stored under a request cache key. It is an entity cache requiring an explicit identity segment.
- Measure List Invalidation Frequencies: If write rates to an entity cause list hit rates to approach near-zero, wholesale list caching is counterproductive. Either shorten list TTLs or eliminate list caching entirely, falling back to database indexes.
- Build Keys From What the Loader Sees: An identifier that is not canonical as given is served uncached — never trimmed or lowercased into another identifier's key — and synthetic shapes (
all) get a namespace no caller-supplied value can reach.
6. Server Infrastructure & GitOps Configuration
The cache server is deployed as a single lightweight Valkey (Redis-compatible) instance defined in infra/git-ops/base/cache/deployment.yaml:
apiVersion: apps/v1
kind: Deployment
metadata:
name: cache
spec:
replicas: 1
strategy:
type: Recreate
template:
spec:
containers:
- name: cache
image: docker.io/valkey/valkey:8.1.3
args:
- --maxmemory
- 200mb
- --maxmemory-policy
- volatile-lru
- --save
- ""
- --appendonly
- "no"
- --aclfile
- /etc/valkey/acl/users.acl
resources:
requests:
memory: 384Mi
limits:
memory: 384Mi
Two choices sit outside the server flags:
strategy: Recreate: The default rolling update starts the new pod before stopping the old one. For the overlap, two unrelated caches sit behind one Service: a service that reconnects lands on the new, empty instance while its neighbours still write, invalidate, and enqueue on the old one. A short full outage is the failure this design already handles.- Memory requested in full: A pod using more memory than it requested is among the first the kubelet evicts under node pressure — and evicting this one empties the cache and the queue.
6.1 Eviction Architecture: volatile-lru Protects Work Queues
The choice of --maxmemory-policy volatile-lru over allkeys-lru is deliberate:
- Under memory pressure, the server only evicts keys configured with an explicit TTL (
EX). - Every cached entity, bounded list, composite entry, and batch lock carries an explicit TTL. An evicted lock is merely an early release: the lock is a cost optimization, not the correctness mechanism.
- The work-queue Redis Set carries no TTL. It is the only structure on the server that eviction never touches.
- Therefore, when memory pressure mounts, Valkey sheds ephemeral read caches while guaranteeing that queued background tasks are never evicted.
The same property is the risk. Non-TTL keys cannot be shed, so a queue that outgrew maxmemory would first evict every service's entries and then make the server refuse writes. The queue is therefore bounded where work enters it (§4.8). More generally, one server is one memory budget: ACLs decide who may write which keys, not how much memory one participant's traffic may take from the rest.
6.2 Zero Persistence (--save "" and --appendonly no)
Persistence is disabled completely. A cache restart that resurrects yesterday's state creates ghost entries — resurrecting records that databases deleted or updated during downtime. A clean cold-start guarantees that all keys are rebuilt fresh from the primary databases on demand.
The price is paid by the one structure that is not a cache: a restart of the cache server drops the work queue. The queue survives restarts of the workers, not of the cache. Re-submitting is the recovery; work that must survive a cache restart belongs on durable infrastructure (§8.2).
6.3 Server-Enforced ACL Selectors
Access control is enforced by the server's command authorization, defined in infra/git-ops/base/cache/acl-configmap.yaml:
user default off
user users-service on #<sha256> ~users-service_* +@read +@write +@connection -@dangerous +info -scan -randomkey
user products-service on #<sha256> (~products-service_* +@read +@write +@connection -@dangerous +info -scan -randomkey) (~users-service_user_* +@read +@connection -@dangerous -scan -randomkey)
user products-sync-service on #<sha256> ~products-sync-service_* +@read +@write +@connection -@dangerous +info -scan -randomkey
Using Valkey ACL selectors ((...)), products-service is granted full read-write permissions over its own prefix (~products-service_*), but only +@read — and only on the entity it consumes — over the owner's keys (~users-service_user_*). If a bug in products-service issues DEL users-service_user_1, the server immediately returns NOPERM, and anything else users-service caches later is not readable by default.
A few details matter as much as the patterns:
- Passwords are stored as SHA-256 hashes (
#<sha256>), so the ACL file holds no credential; the cleartext lives only in the Secret the services read. The server reads the ACL file at startup, so a rotation lands with the next restart of the cache pod. - Keyless reads are removed.
SCANandRANDOMKEYtake no key argument, so key patterns cannot scope them; under+@readthey would let any participant list every key on the shared server. INFOis granted explicitly. It sits in@dangerous, but the client's connection-time readiness check sends it.- ACLs authorize; they do not encrypt. Where the pod network is not trusted, connect over TLS (
rediss://).
7. Production Traps Avoided
Every mechanism in this architecture addresses a concrete failure experienced in distributed environments:
| # | Trap / Failure Mode | Observable Symptom | Code Architecture Fix | Repo Evidence |
|---|---|---|---|---|
| 1 | Unit Discrepancies in TTLs | Keys expire in 60ms or live 60,000× too long | Seconds are the only TTL unit; milliseconds are derived explicitly, and named in the one option that uses them (commandTimeoutMs) | assertTtlSeconds() |
| 2 | Header-Gated Invalidation | Writes without no-cache leave stale entries | Writes always invalidate; headers are never inspected on write paths | users.service.ts |
| 3 | Freshness Header Dropped | Cache-Control does nothing across hops | Internal HTTP connectors forward Cache-Control by default | tw-http-connector/src/constants/index.ts |
| 4 | Key Casing & Format Drift | Cache misses on every read; keys never hit | Central declarative registry and one builder for every key | cache-keys.ts |
| 5 | Malformed Identifiers as Keys | Keys named [object Object] | Builders enforce /^[a-z0-9][a-z0-9._-]*$/ on the identifier as given | cache-keys.ts |
| 6 | Fire-and-Forget DEL | Dead invalidations fail silently | del() returns the removed key count; the service logs it on every write | products.service.ts |
| 7 | Iterative Batch Invalidation | Stale reads during recomputation window | Pre-drain bulk invalidation (delMany) before upstream network calls | recalculations.service.ts |
| 8 | Permissive Shared Passwords | Rogue service executes FLUSHALL | ACLs isolate key patterns per service and entity, revoke -@dangerous, and store only password hashes | acl-configmap.yaml |
| 9 | Two-Step SET + EXPIRE | Process crash leaves immortal deadlock | Atomic SET NX EX creates self-releasing locks | cache.service.ts |
| 10 | In-Memory Queue with Timer | Process restart silently loses pending tasks | Redis Sets persist the queue across worker restarts; a periodic reconcile drains orphans | recalculations.service.ts |
| 11 | Fail-Open Queue Acceptance | 202 Accepted sent for tasks never saved | addToSet throws CacheUnavailableError (mapped to HTTP 503) whenever the add is not confirmed | cache.service.ts |
| 12 | Unauthenticated Cron Execution | Background drains fail upstream with 401 | Background operator mints dedicated short-lived service JWTs | recalculations.service.ts |
| 13 | Invisible In-Memory Fallbacks | Pods serve divergent data silently | Memory store logs loud warnings through logger and flags /health (store: "memory") | cache.service.ts |
| 14 | Divergent Test Mock Semantics | Tests pass while production breaks on JSON, or on a caller mutating a value it read | MemoryCacheStore serialises on write and parses on every read, like the server | memory-cache.store.ts |
| 15 | Cache Persistence Ghosts | Cache restart resurrects deleted entities | Persistence disabled (--save "" --appendonly no); clean cold start | deployment.yaml |
| 16 | Default noeviction Exhaustion | Cache becomes read-only when memory fills | --maxmemory-policy volatile-lru evicts TTL data, protecting queue | deployment.yaml |
| 17 | Silent Configuration Failure | Process boots with invalid NaN TTL | Boot-time validation throws loud CacheConfigurationError — for module TTLs and for the TTLs services read themselves | cache.module.ts |
| 18 | Laundered Stale Aggregates | Cached aggregate computed from stale cache | Aggregates are computed from product data, never from another aggregate; the batch fetches it with no-cache | recalculations.service.ts |
| 19 | Parameter-Order Request Duplicate | Same request cached twice under different keys | requestKey() sorts parameters and normalizes code-written names | request-key.ts |
| 20 | Blanket Interceptor Caching | User-specific or error responses cached | Explicit opt-in via handler logic; throwing loaders never store | products.service.ts |
| 21 | "Smart" List Invalidation Bugs | Categories retain deleted ghost items | Wholesale invalidation: any write wipes all list shapes | products.service.ts |
| 22 | Resurrected Negative 404s | Newly created entity serves cached 404 | Every write path — creation included — invalidates the entity key, and every write goes through the owner | users.service.ts |
| 23 | Free-Text Search Keys | Unbounded key explosion exhausts memory | Free-text search parameters are refused by builders; query direct | request-key.ts |
| 24 | Normalizing Builders, Raw Loaders | One request for /v1/users/%201 caches a 404 under user 1's key, served by every consumer | Identifiers validated, never rewritten; non-canonical spellings served uncached | cache-keys.ts |
| 25 | Synthetic Shape Collisions | ?category=all overwrites the unfiltered list with [] | Shapes are all and category.<name>; no category can name the synthetic one | products.service.ts |
| 26 | Throwing Builders After the Commit | A valid write returns 500 and blocks every later invalidation | Invalidation uses cacheKeyOrNull, built after the response | products.service.ts |
| 27 | Fail-Open Without a Deadline | Requests hang on a connected but unresponsive cache while /health looks fine; in-flight writes are re-sent after reconnect | A per-command timeout; three in a row mark the cache down and reconnect; a clean close counts as down; in-flight commands are never re-sent | redis-cache.store.ts |
| 28 | No-Cache Joining a Stale Flight | A read-your-writes request returns the pre-write value | no-cache starts its own load; superseded loads do not write | read-through.service.ts |
| 29 | Lost Operator Wakeups | Work accepted after a drain's last pop — or orphaned before a restart — never runs | The drain releases the lock and re-checks the queue; reconciliation runs on an interval, not once at startup | recalculations.service.ts |
| 30 | Unbounded Un-Evictable Queue | A flood of ids evicts every service's entries, then writes fail | Validated, size-limited enqueue and a MAX_QUEUE_SIZE cap | recalculations.service.ts |
| 31 | Rolling Update of a Single Cache | Two unrelated caches behind one Service during a rollout | strategy: Recreate | deployment.yaml |
| 32 | Reconnect Loop After Shutdown | A pod stopped during a cache outage hangs until it is killed | Shutdown sends QUIT, then disconnects; a refused QUIT no longer leaves the client redialing | redis-cache.store.ts |
8. Scalability & Evolution Path
8.1 From Single Instance to High Availability
- Current State: A single Valkey replica (
replicas: 1,strategy: Recreate) without persistent disks. If the pod restarts, memory is cleared; services fail open for the duration of the restart window, keys are re-populated on demand, and queued background work is lost (§6.2). - Evolution Path (Primary-Replica Failover): When a single instance's availability becomes the limiting factor, the architecture transitions to a primary-replica topology with automated failover (Valkey Sentinel, or a managed service such as AWS ElastiCache / GCP Memorystore). Replication is asynchronous, so a failover can lose the most recent writes — including invalidations, which leaves entries the owner already deleted in place until their TTL expires — and recent queue additions.
- Evolution Path (Sharded Cluster): When throughput or memory saturates, keys are distributed across shards. Without hash tags, Valkey Cluster already hashes the whole key, so a service's entries spread naturally; a hash tag such as
{users-service}_user_1would do the opposite and concentrate every key of a service onto one shard. The real work is elsewhere: multi-key commands likedelManymust be split per hash slot, because the cluster rejects commands that span slots, and the ACL rules must be applied on every node. An application-transparent sharding proxy (e.g., Envoy) avoids the client changes but terminates authentication, so the servers see the proxy's identity rather than each service's — per-service ACL enforcement then has to move into the proxy, or it is lost. Modifying key delimiters also necessitates a coordinated update across the central key registry, client key builders, and GitOps ACL pattern rules (~users-service_user_*).
8.2 Work Queue Evolution
- The Durability Boundaries: The queue survives worker restarts but not cache restarts (persistence is off). And
SPOPextracts IDs before background calculation finishes: if a pod experiences an ungraceful SIGKILL mid-drain, popped IDs are lost until re-triggered. - Evolution Path: For workflows where either loss is unacceptable, move the queue to durable infrastructure while retaining the shared cache for entity reads: Valkey Streams with consumer groups (
XREADGROUP,XACK,XAUTOCLAIM) on a persistent instance — each entry stays pending until acknowledged, and another worker can claim it after a crash — or a dedicated broker (RabbitMQ/Kafka). Patching the set withSMOVEinto a processing set does not scale to batches: it moves one member per call.
9. Trade-off Analysis
| Architectural Decision | Benefit Gained | Concrete Cost Paid |
|---|---|---|
| Shared Cache Tier | Zero inter-service HTTP hops on hits; single cluster to monitor; cluster-wide deduplication | Single shared availability domain and a single memory budget; requires strict ACL governance and per-participant capacity monitoring |
| Single-Writer Ownership | Eliminates distributed invalidation buses and cross-service write conflicts | Consumers cannot repair foreign entries on write; must call owner. The read-fill/invalidate race inside the owner remains: staleness is bounded by the TTL, not zero |
| Validated, Never-Rewritten Keys | No request can write another identifier's entry | Non-canonical spellings of an id are served uncached |
| Wholesale List Invalidation | Eliminates stale list bugs, ghost categories, and complex invalidation code | Writes cause temporary list cache misses across all categories |
| Clock-Bounded Composites | Enables caching expensive cross-service composite views without bus coordination | Upstream updates can be stale for up to the composite TTL duration |
| Fail-Open Read Policy | Cache crashes — and silent caches — never take down user-facing traffic | Owners and their databases must absorb 100% of consumer reads during outages, including any per-item fan-out the cache was hiding |
| Fail-Closed Queue Policy | Accepted background jobs were stored when they were accepted | Enqueue endpoints return HTTP 503 during cache outages; a cache restart still drops the queue |
volatile-lru Eviction | Protects work queues from eviction under load | Non-TTL keys cannot be evicted, so the queue must be capped at enqueue |
| Zero Persistence | Prevents ghost reads after cache restarts; instant recovery | Cache is cold on startup; queued background work is lost with it |
10. Conclusion
High-efficiency caching in a microservice estate does not require complex distributed pub/sub invalidation networks or heavyweight service-mesh tiers. By combining a single shared cache engine with strict namespace ownership, server-level ACL isolation, and a principled ladder of invalidation strategies, teams can eliminate inter-service network hops while keeping their operational footprint radically simple.
The architecture shows that appropriate complexity is the decisive metric: exact invalidation for entities, blunt wholesale invalidation for lists, clock bounds for composites, fail-closed acceptance for background queues — and, on every rung, a TTL that bounds how wrong the cache can be when an invalidation is raced or lost.
References
- Reference Implementation: github.com/ivanbaha/team-workspace
- Core Library:
libs/tw-cache - Key Registry:
libs/tw-cache/src/cache-keys.ts - Failure Policy Layer:
libs/tw-cache/src/cache.service.ts - Single-Flight & Read-Through:
libs/tw-cache/src/read-through.service.ts - GitOps ACLs & Eviction:
infra/git-ops/base/cache/ - Operational Runbook:
docs/guides/debugging-the-cache.md
- Core Library:
- HTTP Caching Specification: RFC 9111 · HTTP Caching
- Valkey Access Control Lists: Valkey ACL Documentation
- Valkey Streams and Consumer Groups: Valkey Streams Documentation
- Cache Stampede Prevention: Vattani, Chierichetti & Lowenstein, Optimal Probabilistic Cache Stampede Prevention, PVLDB 8(8):886 — 897, VLDB 2015, DOI: 10.14778/2757807.2757813
- Leases and Stale Sets: Nishtala et al., Scaling Memcache at Facebook, 10th USENIX Symposium on Networked Systems Design and Implementation (NSDI '13), pp. 385 — 398, USENIX