Skip to main content

Appropriate Caching in a Multi-Service Estate (Balancing Efficiency and Simplicity Across a Shared Store)

Metadata
  • ID: RA-006-APPROPRIATE-CACHING
  • Status: Production-Validated · 40+ Microservices Across Several Teams · In use since 2026
  • Author: Ivan Baha (ORCID: 0009-0005-7024-7724)
  • Published Date: 2026-09-28
  • Repository: github.com/ivanbaha/team-workspace — reference implementation of @tw/cache, participating services, and GitOps manifests
  • Tags: #caching #redis #valkey #microservices #architecture #fault-tolerance

1. Executive Summary​

In microservice architectures, caching typically polarises into two unsatisfactory archetypes:

  1. Isolated Per-Service Caches: Each service maintains its own cache tier or in-memory heap. When Service A needs data owned by Service B, it must make a network hop via HTTP or gRPC to Service B, even if Service B has just cached the exact record. Invalidation does not propagate across replicas; memory is multiplied N-fold; and data inconsistencies between pods remain undetectable.
  2. Coordinated Invalidation Buses: Services share a distributed cache or keep local caches synchronised via a pub/sub message broker (Redis Pub/Sub, Kafka, RabbitMQ). While theoretically reactive, this adds a secondary distributed system whose failure modes — message drops, consumer lag, split-brain states, and out-of-order race conditions — directly compromise cache consistency.

This Reference Architecture presents a pragmatic middle ground, validated in production across more than 40 microservices: a single shared cache server operated under strict single-writer ownership with direct consumer reads.

Rather than coordinating invalidation between peers, the architecture enforces a rigid ownership contract within the key namespace itself: {service}_{entity}_{id}. Exactly one service — the domain owner of the data — writes to and invalidates that key. Consumers read the owner's keys directly from the cache. A cache hit eliminates the inter-service HTTP hop entirely: no outbound request serialisation, no inbound middleware evaluation, no trace fan-out, and no database query. On a cache miss, the consumer makes a normal HTTP call to the owner; the owner's read-through path populates the entry; and subsequent reads by any consumer immediately hit.

To ensure this model remains resilient, maintainable, and cost-effective, the architecture is underpinned by three core design principles:

  • Server-Enforced Access Control (ACLs): Consumers are granted read-only permissions on exactly the foreign entities they consume, at the cache server level (Valkey ACLs). A consumer bug can never write or delete another service's cache entries — and because every key is built from the identifier exactly as its loader sees it, no request can poison an owner's entry through the owner's own read path either.
  • Granular Operation-Specific Failure Policies: Reads fail open (returning null, protecting user-facing traffic when the cache is down — or merely silent, because every command has a deadline); queue writes fail closed (throwing explicit errors to prevent silent data loss); and configuration bugs fail loud at startup.
  • The Caching Ladder (Appropriate Complexity for Each Shape): The team does not apply a blanket caching decorator. Read paths are categorised onto a ladder of rungs — Entity, Bounded List, and Composite Request. Each rung uses the simplest invalidation strategy it can honestly guarantee: exact 1-to-1 eviction for entities, blunt wholesale invalidation for bounded lists, and owner-side eviction plus a short self-healing TTL for composites. On every rung, the entry's TTL is the backstop that bounds staleness when an invalidation is raced or lost.

Finally, the same cache server acts as an operator execution tier: Redis Sets provide deduplicating background work queues, atomic SET NX EX keys provide self-releasing batch locks, and the cache engine's volatile-lru eviction policy ensures that memory pressure evicts ephemeral cached data while never touching queued work — which is also why the queue is capped where work enters it.

Provenance. The author designed this architecture for a team responsible for 80+ microservices on an enterprise platform of 200+. It has since spread beyond that team: more than 40 microservices across several teams now use it in production. The linked repository is a self-contained reference implementation of the library, contracts, services, and GitOps manifests described here, not a direct proprietary extract.


2. Applicability Criteria​

Use this pattern when​

  • Inter-service read amplification is high. Services frequently query peer services for reference or relational data (e.g., resolving user profiles, product catalogs, permissions) across HTTP/gRPC boundaries.
  • The team manages multiple microservices sharing an infrastructure tier. Maintaining isolated caching infrastructure (or separate Redis clusters) for each service imposes disproportionate operational and financial overhead.
  • A pub/sub invalidation bus is over-engineering. The write volume and latency tolerance allow direct read-through and asynchronous owner deletion without the complexity of message brokers.
  • Operations require predictable failure modes. A cache outage must never take down primary business flows, and background jobs must never be acknowledged without being stored. (If accepted jobs must also survive a restart of the cache server itself, keep the queue on durable infrastructure — see §8.2.)
  • Background batching and coordination are needed without extra queuing brokers. Work can be deduplicated and debounced using native Redis primitives (Sets and self-releasing locks) without provisioning Kafka, RabbitMQ, or SQS.

Do NOT use this pattern when​

  • Strong linearizable ACID read consistency is required. An owner's write makes its entries fresh immediately in the common case, but a read racing that write, or an invalidation dropped during a connection blip, can leave a stale value in place for up to one entry TTL. If every consumer read must see transactional updates immediately, queries must hit the authoritative transactional database directly.
  • Tenant data must be physically or cryptographically air-gapped. When strict multi-tenant compliance prevents sharing an infrastructure network namespace or cache process across service domains.
  • Write frequency matches or exceeds read frequency. High-churn datasets will suffer near-zero hit rates while incurring continuous invalidation overhead.
  • Payloads exceed multi-megabyte boundaries. Storing massive JSON blobs or unbounded collections in single keys degrades Redis single-threaded performance and triggers latency spikes.
  • One participant's volume would dwarf the rest. ACLs isolate who writes which keys, not how much memory each participant takes. A service whose working set crowds the shared maxmemory evicts everyone else's entries; give it its own instance.

3. Architecture Overview​

3.1 The Three Foundational Principles​

  1. The Key is the Ownership Contract: Every key is prefixed with the owning service's deployment name: {service}_{entity}_{id}. The prefix dictates which service writes it, which service invalidates it, and which ACL rules apply. The identifier is used exactly as its loader sees it — validated, never rewritten.
  2. Appropriate Complexity (The Ladder): Avoid "clever" partial invalidation. Use exact invalidation where relationships are 1-to-1, wholesale invalidation where relationships are 1-to-many, and clock-bounded TTLs for whatever crosses service boundaries. On every rung the TTL is also the backstop: it bounds how long a raced or dropped invalidation can leave a stale value.
  3. Symmetric Behavioral Parity: Development, test, and production environments execute identical serialization and outage semantics. In-memory fallbacks serialize on write and parse on every read, exactly like the server, and delete type-agnostically. The one difference no fallback can emulate — the real client is still connecting when the application boots — is designed around: nothing depends on a check made at startup.

3.2 Constraints Decide the Shape​

Every mechanism in this architecture traces back directly to an operational constraint:

ConstraintArchitectural DecisionSection
Inter-service HTTP calls inflate latency and trace fan-outConsumers read owner cache keys directly from the shared server§3.3
A bug in a consumer must not corrupt another service's dataValkey ACL selectors grant consumers read-only access to exactly the entities they consume§6.3
A request must not be able to write another identifier's entryIdentifiers are validated, never trimmed or lowercased; non-canonical spellings are served uncached§4.1
Cache server downtime must not take down user-facing trafficget operations fail open, falling back to the source of truth§4.2
A connected but silent server must not hang requestsEvery command has a deadline; three timeouts in a row mark the cache down and force a reconnect§4.2
Enqueued background tasks must never be acknowledged without being storedaddToSet fails closed, throwing CacheUnavailableError (HTTP 503)§4.2, §4.8
Non-existent entities cause database 404 hammeringAbsence is negative-cached using a sentinel value ("not_present")§4.3
Concurrent misses for the same key cause database stampedesIn-process single-flight promise coalescing joins concurrent requests on each pod§4.3
Complex list invalidation branches lead to silent stale readsAny write to an entity wipes all list shapes wholesale (delMany)§4.5
Multi-service composite responses cannot be invalidated by foreign ownersComposites use canonical keys, owner-side eviction, and short self-healing TTLs§4.6
Background jobs need deduplication without extra brokersRedis Sets (SADD/SPOP) and SET NX EX locks provide native queuing§4.8
Every accepted background job must get a drainThe drain releases its lock and re-checks the queue; a periodic reconcile catches the rest§4.8
Memory pressure must not evict queued jobs — nor let the queue evict everything else--maxmemory-policy volatile-lru; the queue carries no TTL and is capped at enqueue§6.1, §4.8

3.3 End-to-End Interaction Flows​

Consumer Direct Read & Fallback Flow​

The following sequence depicts products-service (consumer) resolving a product's owner via the shared cache, eliminating an HTTP hop to users-service (owner):

Operator Work Queue & Batch Lock Flow​

The following sequence illustrates products-sync-service leveraging shared cache primitives for deduplicated background recalculations:


4. Implementation Details​

The implementation spans one core library (@tw/cache) and three domain services in the reference workspace (core files shown):

libs/tw-cache/
├── src/
│ ├── cache-keys.ts # Central registry & cacheKey()/cacheKeyOrNull() builders
│ ├── request-key.ts # Canonical requestKey() serializer
│ ├── cache.service.ts # Policy layer, failure handling, timeout breaker, health counters
│ ├── read-through.service.ts # Single-flight, negative caching, no-cache supersession
│ ├── redis-cache.store.ts # ioredis client (offline queue off, per-command timeout)
│ ├── memory-cache.store.ts # JSON-parity memory store for local development
│ ├── cache.module.ts # Module registration & boot-time config validation
│ ├── validate.ts # Input validation (assertTtlSeconds, assertCommandTimeoutMs)
│ ├── errors.ts # Typed errors (CacheUnavailableError, etc.)
│ └── test/cache.fake.ts # Standalone behavioural test double
backend/
├── users-service/ # Owner: user entity
├── products-service/ # Owner: product, productList, catalog composite; Consumer: user
└── products-sync-service/ # Operator: background queue, lock, reconcile loop, categoryStats
infra/git-ops/base/cache/
├── deployment.yaml # Valkey 8: volatile-lru, persistence off, Recreate
└── acl-configmap.yaml # Per-service ACL rules: hashed passwords, per-entity consumer grants

4.1 Strict Key Registry & Identifier Validation​

Keys are never concatenated via raw string interpolation ('users-service_user_' + id). Raw concatenation risks casing drift, missing segments, and silent ACL permission denials. Even the operator's queue and lock keys, which carry no id, take their prefix from the registry.

Entity Key Construction​

CACHE_NAMESPACES in libs/tw-cache/src/cache-keys.ts acts as the single source of truth for all cached entities:

export const CACHE_NAMESPACES = {
user: { service: 'users-service', entity: 'user' },
categoryStats: { service: 'products-sync-service', entity: 'categoryStats' },
product: { service: 'products-service', entity: 'product' },
productList: { service: 'products-service', entity: 'productList' },
productCatalog: { service: 'products-service', entity: 'req' },
} as const;

The cacheKey(name, id) builder validates identifiers against /^[a-z0-9][a-z0-9._-]*$/ as given — it never trims or lowercases them.

That last point carries more weight than it appears to. A key must be a function of exactly what the loader behind it sees. A builder that normalized identifiers would give ' 1' the key of '1' (and Widgets the key of widgets) while their loaders — matching the raw value against the source of truth — return different answers. Whichever loader filled the key first would then answer for both: a single GET /v1/users/%201, sent with Cache-Control: no-cache, would store "does not exist" under user 1's key, and every consumer reading that key would serve the 404. Normalizing in the builder turns a harmless malformed request into cross-service cache poisoning; validating turns it into an uncached read.

The same rule covers synthetic shapes. The unfiltered product list is cached as the shape all and a category as the shape category.<name>, so no caller-supplied category can ever produce the synthetic shape's key.

Safe Validation for Data (cacheKeyOrNull)​

When an identifier comes from data — a path parameter in GET /v1/users/:id, but equally a stored column or connector data — an invalid format (spaces, control characters, a non-canonical spelling) must not throw a 500 error. Instead, cacheKeyOrNull() returns null:

const key = cacheKeyOrNull('user', id);
const user = await (key ? this.readThrough.readThrough({ key, load }) : load());

On a read, the application falls back to the authoritative store, which answers for exactly the identifier it was asked about — an HTTP 404 rather than an internal cache failure. On a write, null means there is nothing to invalidate: an identifier that cannot form a key never had an entry, because every read of it went uncached. The throwing cacheKey() is reserved for identifiers the code itself constructs; a throwing builder on a write path fails a write the database has already committed.

Consumers apply the same rule to the foreign ids they hold. A product's owner id came from a request body, so one that cannot form a key skips the cache read and goes straight to the owner over HTTP-URL-encoded, since it is data, not a path — instead of failing every catalogue request that includes that product.


4.2 Granular Failure Policy: Fail-Open vs. Fail-Closed vs. Fail-Loud​

The client library encapsulates the failure matrix, so individual developers cannot inadvertently introduce unsafe behaviours:

// libs/tw-cache/src/cache.service.ts
async get<T>(key: string): Promise<T | null> {
if (this.unreachable) {
this.skipped++;
return null;
}
try {
const value = await this.run(() => this.store.get<T>(key)); // run() counts consecutive timeouts
if (value === null) this.misses++;
else this.hits++;
return value;
} catch (error) {
this.skipped++;
this.logger.warn(`Cache get for ${key} failed, failing open: ${(error as Error).message}`, CONTEXT);
return null;
}
}

async addToSet(key: string, members: string | string[]): Promise<number> {
const items = Array.isArray(members) ? members : [members];
if (this.unreachable) {
throw new CacheUnavailableError(
`Cache unreachable, refusing to add ${items.length} item(s) to ${key}: a set add is a work-queue write, ` +
'and reporting it accepted would lose the work.',
);
}
// ...
}
  • Boot-time Validation: Any non-integer or NaN TTL (Number(process.env.CACHE_TTL)) immediately throws CacheConfigurationError, crashing the process during bootstrap before accepting traffic. TTLs a service reads for itself — such as the composite TTL — are checked the same way when their module loads (assertTtlSeconds), not once per request.
  • Offline Queue Disabled: The Redis client initializes with enableOfflineQueue: false. Commands reject instantly when disconnected, instead of queueing behind a reconnect and applying later against changed state. Commands already in flight when a connection drops are failed rather than re-sent after the reconnect (autoResendUnfulfilledCommands: false) — a re-sent SET is a fill that lands after whatever invalidation happened in between.
  • A Deadline on Every Command: Connection events report a server that has gone away — an error, or the clean close a server sends when it shuts down gracefully. A server that is connected but silent — a hung process, or a lost node whose TCP connection nothing has closed yet — emits nothing, and without a deadline every read would wait until the kernel gives up on the socket: minutes. Each command therefore carries a timeout (commandTimeoutMs, 250 ms by default). When it expires, the command fails the way it would against an unreachable server — reads open, queue writes closed — and three timeouts in a row mark the cache unreachable and force a reconnect, so requests stop paying the timeout one by one.
  • Unknown Outcomes Are Refused, Not Guessed: A queue write that times out may still have reached the server. The endpoint answers 503 anyway, and the retry it asks for is safe: a set add is idempotent.
  • A Clean Shutdown, Even Mid-Outage: On shutdown the client sends QUIT and then disconnects outright. While the connection is down, QUIT is refused (there is no offline queue to hold it), and a client left to itself would keep reconnecting — holding a stopping pod open until it is killed. A service's own shutdown is not reported as a cache outage, either.

4.3 Read-Through, Negative Caching & In-Process Single-Flight​

ReadThroughService eliminates the boilerplate check-load-store pattern and implements three protections:

// libs/tw-cache/src/read-through.service.ts
export class ReadThroughService {
private readonly inflight = new Map<string, Promise<unknown>>();

async readThrough<T>(options: ReadThroughOptions<T>): Promise<T | null> {
assertTtlSeconds(options.ttlSeconds, 'ttlSeconds');
assertTtlSeconds(options.negativeTtlSeconds, 'negativeTtlSeconds');

const { key, load, traceId } = options;

if (!options.noCache) {
const cached = await this.cache.get<unknown>(key);
if (cached !== null) {
const negative = cached === NEGATIVE_CACHE_SENTINEL;
this.logger.debug(JSON.stringify({ cache: negative ? 'negative-hit' : 'hit', key }), CONTEXT, traceId);
return negative ? null : (cached as T);
}
this.logger.debug(JSON.stringify({ cache: 'miss', key }), CONTEXT, traceId);
} else {
this.logger.debug(JSON.stringify({ cache: 'bypass', key }), CONTEXT, traceId);
}

// A no-cache read never joins: the flight under way may have read the source before the write
// the caller is trying to see.
const existing = options.noCache ? undefined : this.inflight.get(key);
if (existing) {
this.logger.debug(JSON.stringify({ cache: 'joined', key }), CONTEXT, traceId);
return existing as Promise<T | null>;
}

// A flight superseded by a later no-cache read still answers its callers, but does not write.
const flight: Promise<T | null> = this.loadAndStore(options, () => this.inflight.get(key) === flight).finally(
() => {
if (this.inflight.get(key) === flight) this.inflight.delete(key);
},
);
this.inflight.set(key, flight);
return flight;
}
}
  • Negative Caching: When load() returns null (proven absence in the database), the service writes NEGATIVE_CACHE_SENTINEL ("not_present") for negativeTtlSeconds (default: 30s). Future requests are answered immediately without hitting the database. Loaders that fail or time out must throw, ensuring infrastructure errors are never cached as negative hits. The sentinel is a plain string value, so a namespace whose legitimate values could equal it must not use negative caching.
  • Single-Flight Coalescing: If 50 concurrent requests miss the same key simultaneously on a single pod, only one database query is dispatched; the other 49 callers join the active Promise. The coalescing is per process, so with N replicas a cold key can still cost N loads — the protection bounds a stampede to one load per pod rather than eliminating it. Nor does it remove the latency spike when a hot key expires: every waiter still waits for that one load. Probabilistic early expiration (Vattani et al.) addresses that part, by recomputing a hot key shortly before it expires; the two techniques are complementary, and this architecture has not yet needed the second.
  • No-Cache Supersedes, Never Joins: A no-cache read starts its own load and becomes the key's current flight. The flight it superseded still answers its own callers but no longer writes the entry, so an older value from the same pod cannot land on top of the repair. A stale fill from another replica still can; the entry's TTL bounds it (§4.4).

4.4 Owner Write Invalidation: Unconditional, From Every Write Path​

Writes must invalidate the cache immediately and unconditionally:

// backend/users-service/src/users/users.service.ts
private invalidateUser(id: string): void {
setImmediate(() => {
// An id that cannot form a key was never cached — every read of it went uncached.
const key = cacheKeyOrNull('user', id);
if (!key) return;
void this.cache.del(key).then((removed) => {
this.logger.debug(`Invalidation removed ${removed} cache entry for user ${id} (key ${key})`);
});
});
}

Key principles of the invalidation contract:

  1. Off the Response Path: Invalidation is deferred via setImmediate(), and the key is built there too. The caller's HTTP write request completes without paying the latency of a cache round trip, and nothing about building a cache key can turn a committed write into a 500.
  2. Never Header-Gated: Cache-Control headers govern read freshness; writes always invalidate.
  3. Every Write Path: Invalidation is only as complete as the write paths that remember it. In the reference workspace, registration lives in a separate authentication module, so it creates users through the owner's UsersService.create(), which invalidates the new id's key. A module, batch job, or migration that writes directly to the store is a write the cache never hears about.
  4. The Delete Count as a Debugging Aid: Redis DEL returns the integer number of keys actually removed, and the owner logs it on every write. After a read that filled the key, removed 0 means the invalidation was built with a different key than the read — a "dead invalidation". On its own, 0 is also what every write to an uncached key gives, and what a dropped invalidation gives during an outage (the client logs each drop at debug, so the two can be told apart), so the count is a probe rather than an alarm; key drift is prevented by construction (one registry, one builder) and caught by tests.

How stale can an entry get? Usually not at all: the owner's write deletes it and the next read refills it. Three things can still leave an old value in place, and the design names the bound for each rather than pretending they do not happen:

  • The deferred DEL: a client that writes and immediately re-reads can land its GET before the setImmediate delete runs.
  • A dropped DEL: deletes are best-effort; during a connection blip one is dropped, and if the cache comes back without restarting, the old entry is still there.
  • A stale fill: a read that loaded the database before the write committed can store its value after the write's DEL — usually on another replica, where nothing in-process can see it coming. Single-writer ownership removes coordination between services; it does not remove this race inside the owner.

In all three cases, the entry's TTL is the bound: 60 seconds for entities and lists, 30 for negative entries. Clients that must read their own writes send Cache-Control: no-cache, which starts a fresh load instead of joining one already in flight. Closing the stale-fill race outright takes versioned or leased fills — the lease mechanism described in Scaling Memcache at Facebook — which this architecture defers until the TTL stops being a good enough bound.


4.5 List Caching: The Case for Blunt Wholesale Invalidation​

List caching commonly fails when engineers attempt "smart" invalidation — tracking category memberships, pagination boundaries, and sorting permutations. This introduces extensive branching logic that fails silently when items move or empty.

This architecture enforces blunt wholesale invalidation:

// backend/products-service/src/products/products.service.ts
private invalidateAfterWrite(productId: string, extraCategories: string[] = []): void {
setImmediate(() => {
const itemKey = cacheKeyOrNull('product', productId);
const keys = [
...this.listKeys(extraCategories), // 'all' + 'category.<name>' per category; un-keyable ones skipped
...(itemKey ? [itemKey] : []),
requestKey('productCatalog', 'v1-products', { expandOwner: true }),
];
void this.cache.delMany(keys).then((removed) => {
this.logger.debug(`Invalidation removed ${removed} cache entries after write to ${productId}`);
});
});
}
  • Every Shape Evicted: Any write, update, or deletion of a product invalidates productList_all and every known productList_category.<name> shape via a single delMany().
  • Ghost Category Prevention: When an item is moved to another category or deleted, the old category is passed in extraCategories because an empty category can no longer be named by scanning active store records.
  • Un-keyable Shapes Are Skipped, Not Fatal: A category that cannot form a key (Home & Garden) is never cached — its reads go straight to the store — so its invalidation is skipped instead of throwing after the write has already committed.
  • Empty Lists are Valid Values: A query returning no records caches [] as an authentic value. null is reserved for proven entity absence.
  • Free-Text Search is Never Cached: ?search=shoes bypasses the cache and queries the database every time. An unbounded search space cannot be enumerated, making complete invalidation impossible.

4.6 Composite Request Caching: Clock-Bounded Invalidation​

Composite endpoints aggregate data across domain boundaries (e.g., GET /v1/products?expandOwner=true, which joins products with user records from users-service).

Composite caching balances local vs. foreign ownership:

  1. Canonical Serialization: requestKey() sorts parameters by name and normalizes the route and parameter names — both written by code — but validates parameter values as given, rejecting anything that is not a plain token. A value is what the loader filters on; rewriting it would alias two different answers onto one key:

    products-service_req_v1-products_expandowner=true
  2. Split Invalidation Responsibility:

    • Local Product Changes: products-service writes invalidate the composite key directly, ensuring price or title changes are instantly visible.
    • Upstream User Changes: users-service has no knowledge of downstream composites. Therefore, the clock is the invalidation. The composite carries a tight, self-healing TTL (e.g., 10 seconds).
  3. Freshness Reaches the Parts: A caller's no-cache bypasses the composite and the owner entries it is built from. Rebuilding a "fresh" composite from a part that outlived its own invalidation would re-cache exactly the value the caller was trying to get past.

  4. Opt-In Only: Composite caching is explicitly invoked in handler logic. Blanket middleware caching is prohibited to prevent accidental caching of authenticated or user-scoped payloads.


4.7 Active Freshness: End-to-End Traversal of Cache-Control: no-cache​

When a client requires guaranteed fresh data, it transmits standard HTTP semantics:

Cache-Control: no-cache

This header executes a bypass-and-refresh:

  1. The read-through path bypasses the cache read — and never joins a load already in flight, which may predate the caller's write.
  2. It fetches the authoritative value from the database.
  3. It overwrites the cache entry with the fresh data, repairing the cache for all subsequent readers.
  4. Header Propagation: The internal HTTP connector (@tw/http-connector) automatically forwards Cache-Control on outbound inter-service calls, ensuring the freshness demand travels end-to-end to the authoritative owner — including through composites, whose parts are bypassed too.

This borrows a request directive that RFC 9111 defines for HTTP caches and applies it to the caches behind the origin, which makes it an amplification lever: browsers send no-cache on every hard reload and whenever developer tools disable caching, and each such request costs the owner a source read — for a composite, one per part. Who may send it is an edge decision: strip or rate-limit it on public traffic at the gateway, and keep it for internal callers, batch jobs, and operators, for whom it is the repair tool.


4.8 The Operator Pattern: Work Queue & Batch Lock​

products-sync-service acts as an operator, maintaining asynchronous category aggregations without adding RabbitMQ or Kafka:

// backend/products-sync-service/src/recalculations/recalculations.service.ts
async requestRecalculation(categoryIds: string[]): Promise<{ queued: number; scheduled: boolean }> {
// The queue is never evicted, so it is capped where work enters it
const pending = await this.cache.getSizeOfSet(QUEUE_KEY);
if (pending + categoryIds.length > this.maxQueueSize) throw new QueueFullError(/* ... */);
// SADD deduplicates identical category IDs automatically
const queued = await this.cache.addToSet(QUEUE_KEY, categoryIds);
// Atomic SET NX EX locks the drain timer
const scheduled = await this.scheduleBatch();
return { queued, scheduled };
}

private async scheduleBatch(): Promise<boolean> {
const acquired = await this.cache.acquireLock(LOCK_KEY, this.lockTtlSeconds);
if (!acquired) return false;

setTimeout(() => void this.drain(), this.batchDelayMs);
return true;
}
  • Validated, Bounded Input: Category ids are validated with the key builder's own rule — at most 100 per request, 64 characters each — and a request that would take the queue past MAX_QUEUE_SIZE is refused with 503 QUEUE_FULL. The set carries no TTL, so eviction never touches it — which is exactly why it must never grow without bound (§6.1).
  • Atomic Queue Draining: SPOP queue batchSize atomically extracts items. Multiple worker replicas can execute drains concurrently without ever receiving duplicate IDs.
  • Self-Releasing Lock, Released Early: The lock uses SET NX EX <ttl> with a TTL longer than the batch delay, so a burst of requests collapses into one drain. When the drain finishes, it deletes the lock and re-checks the queue; without that step, a request arriving after the drain's final SPOP would find the lock still held, arm nothing, and wait for some later request. If a worker pod crashes mid-batch, the TTL releases the lock.
  • Pre-Drain Bulk Invalidation: Before executing expensive upstream HTTP calls, the batch runs delMany() on all stats keys in the batch. Readers arriving during the calculation window encounter a cache miss and calculate on-demand, preventing stale reads.
  • Periodic Reconciliation: Every RECONCILE_INTERVAL_SECONDS (30 by default), the service checks SCARD queue and, if the lock is free, schedules a drain. A one-off check at startup would not work: at bootstrap the cache client is still connecting, an unreachable cache reports the set as empty, and the check never looks again. An interval also catches work stranded while the service keeps running — a lock holder that crashed before arming its timer, a schedule swallowed by a cache outage, an evicted lock.

5. Strategy Selection: Choosing the Appropriate Approach​

The fundamental premise of this Reference Architecture is matching the caching mechanism to the access shape. Caching every endpoint identically creates brittle systems.

5.1 The Caching Decision Matrix​

Access PatternTarget ShapeStrategy / RungInvalidation TriggerStaleness BoundFailure Policy
Lookup by ID (/users/:id)Single EntityRung 1: Entity CacheExact key DEL on owner writeFresh after the owner's write in the common case; ≤ entity TTL (60s) when an invalidation is raced or droppedFail open (null)
Filtered Query (/products?category=x)Bounded Hydrated ListRung 2: Bounded List CacheWholesale delMany of all shapes on writeFresh after the owner's write in the common case; ≤ list TTL (60s) when an invalidation is raced or droppedFail open (null)
Cross-Domain Composite (/products?expandOwner=true)Multi-Service Joined ObjectRung 3: Composite Request CacheLocal write DEL + Short TTL (10s)≤ 10 seconds for foreign updates, on top of the owner entries' own boundFail open (null)
Arbitrary Free-Text (/products?search=text)Dynamic FilterDirect Store Query (Uncached)None (bypasses cache)None — read from the storeN/A
Background Aggregation (Category Stats)Materialized AggregateOperator Worker PatternPre-drain delMany + Background Set QueueTTL (300s) between recalculations; batch inputs fetched with no-cacheFail closed (503 on queue write)

5.2 Rules of Thumb for Balancing Simplicity vs. Efficiency​

The decision follows ownership first, then shape:

  1. Prefer Lower Rungs: If a query can be satisfied by assembling individually cached entities on the client or consumer side, do not introduce a composite request cache.
  2. Never Cache Identity-Varied Responses: If a response body varies depending on the authenticated caller's identity, permissions, or session attributes, it must never be stored under a request cache key. It is an entity cache requiring an explicit identity segment.
  3. Measure List Invalidation Frequencies: If write rates to an entity cause list hit rates to approach near-zero, wholesale list caching is counterproductive. Either shorten list TTLs or eliminate list caching entirely, falling back to database indexes.
  4. Build Keys From What the Loader Sees: An identifier that is not canonical as given is served uncached — never trimmed or lowercased into another identifier's key — and synthetic shapes (all) get a namespace no caller-supplied value can reach.

6. Server Infrastructure & GitOps Configuration​

The cache server is deployed as a single lightweight Valkey (Redis-compatible) instance defined in infra/git-ops/base/cache/deployment.yaml:

apiVersion: apps/v1
kind: Deployment
metadata:
name: cache
spec:
replicas: 1
strategy:
type: Recreate
template:
spec:
containers:
- name: cache
image: docker.io/valkey/valkey:8.1.3
args:
- --maxmemory
- 200mb
- --maxmemory-policy
- volatile-lru
- --save
- ""
- --appendonly
- "no"
- --aclfile
- /etc/valkey/acl/users.acl
resources:
requests:
memory: 384Mi
limits:
memory: 384Mi

Two choices sit outside the server flags:

  • strategy: Recreate: The default rolling update starts the new pod before stopping the old one. For the overlap, two unrelated caches sit behind one Service: a service that reconnects lands on the new, empty instance while its neighbours still write, invalidate, and enqueue on the old one. A short full outage is the failure this design already handles.
  • Memory requested in full: A pod using more memory than it requested is among the first the kubelet evicts under node pressure — and evicting this one empties the cache and the queue.

6.1 Eviction Architecture: volatile-lru Protects Work Queues​

The choice of --maxmemory-policy volatile-lru over allkeys-lru is deliberate:

  • Under memory pressure, the server only evicts keys configured with an explicit TTL (EX).
  • Every cached entity, bounded list, composite entry, and batch lock carries an explicit TTL. An evicted lock is merely an early release: the lock is a cost optimization, not the correctness mechanism.
  • The work-queue Redis Set carries no TTL. It is the only structure on the server that eviction never touches.
  • Therefore, when memory pressure mounts, Valkey sheds ephemeral read caches while guaranteeing that queued background tasks are never evicted.

The same property is the risk. Non-TTL keys cannot be shed, so a queue that outgrew maxmemory would first evict every service's entries and then make the server refuse writes. The queue is therefore bounded where work enters it (§4.8). More generally, one server is one memory budget: ACLs decide who may write which keys, not how much memory one participant's traffic may take from the rest.

6.2 Zero Persistence (--save "" and --appendonly no)​

Persistence is disabled completely. A cache restart that resurrects yesterday's state creates ghost entries — resurrecting records that databases deleted or updated during downtime. A clean cold-start guarantees that all keys are rebuilt fresh from the primary databases on demand.

The price is paid by the one structure that is not a cache: a restart of the cache server drops the work queue. The queue survives restarts of the workers, not of the cache. Re-submitting is the recovery; work that must survive a cache restart belongs on durable infrastructure (§8.2).

6.3 Server-Enforced ACL Selectors​

Access control is enforced by the server's command authorization, defined in infra/git-ops/base/cache/acl-configmap.yaml:

user default off
user users-service on #<sha256> ~users-service_* +@read +@write +@connection -@dangerous +info -scan -randomkey
user products-service on #<sha256> (~products-service_* +@read +@write +@connection -@dangerous +info -scan -randomkey) (~users-service_user_* +@read +@connection -@dangerous -scan -randomkey)
user products-sync-service on #<sha256> ~products-sync-service_* +@read +@write +@connection -@dangerous +info -scan -randomkey

Using Valkey ACL selectors ((...)), products-service is granted full read-write permissions over its own prefix (~products-service_*), but only +@read — and only on the entity it consumes — over the owner's keys (~users-service_user_*). If a bug in products-service issues DEL users-service_user_1, the server immediately returns NOPERM, and anything else users-service caches later is not readable by default.

A few details matter as much as the patterns:

  • Passwords are stored as SHA-256 hashes (#<sha256>), so the ACL file holds no credential; the cleartext lives only in the Secret the services read. The server reads the ACL file at startup, so a rotation lands with the next restart of the cache pod.
  • Keyless reads are removed. SCAN and RANDOMKEY take no key argument, so key patterns cannot scope them; under +@read they would let any participant list every key on the shared server.
  • INFO is granted explicitly. It sits in @dangerous, but the client's connection-time readiness check sends it.
  • ACLs authorize; they do not encrypt. Where the pod network is not trusted, connect over TLS (rediss://).

7. Production Traps Avoided​

Every mechanism in this architecture addresses a concrete failure experienced in distributed environments:

#Trap / Failure ModeObservable SymptomCode Architecture FixRepo Evidence
1Unit Discrepancies in TTLsKeys expire in 60ms or live 60,000× too longSeconds are the only TTL unit; milliseconds are derived explicitly, and named in the one option that uses them (commandTimeoutMs)assertTtlSeconds()
2Header-Gated InvalidationWrites without no-cache leave stale entriesWrites always invalidate; headers are never inspected on write pathsusers.service.ts
3Freshness Header DroppedCache-Control does nothing across hopsInternal HTTP connectors forward Cache-Control by defaulttw-http-connector/src/constants/index.ts
4Key Casing & Format DriftCache misses on every read; keys never hitCentral declarative registry and one builder for every keycache-keys.ts
5Malformed Identifiers as KeysKeys named [object Object]Builders enforce /^[a-z0-9][a-z0-9._-]*$/ on the identifier as givencache-keys.ts
6Fire-and-Forget DELDead invalidations fail silentlydel() returns the removed key count; the service logs it on every writeproducts.service.ts
7Iterative Batch InvalidationStale reads during recomputation windowPre-drain bulk invalidation (delMany) before upstream network callsrecalculations.service.ts
8Permissive Shared PasswordsRogue service executes FLUSHALLACLs isolate key patterns per service and entity, revoke -@dangerous, and store only password hashesacl-configmap.yaml
9Two-Step SET + EXPIREProcess crash leaves immortal deadlockAtomic SET NX EX creates self-releasing lockscache.service.ts
10In-Memory Queue with TimerProcess restart silently loses pending tasksRedis Sets persist the queue across worker restarts; a periodic reconcile drains orphansrecalculations.service.ts
11Fail-Open Queue Acceptance202 Accepted sent for tasks never savedaddToSet throws CacheUnavailableError (mapped to HTTP 503) whenever the add is not confirmedcache.service.ts
12Unauthenticated Cron ExecutionBackground drains fail upstream with 401Background operator mints dedicated short-lived service JWTsrecalculations.service.ts
13Invisible In-Memory FallbacksPods serve divergent data silentlyMemory store logs loud warnings through logger and flags /health (store: "memory")cache.service.ts
14Divergent Test Mock SemanticsTests pass while production breaks on JSON, or on a caller mutating a value it readMemoryCacheStore serialises on write and parses on every read, like the servermemory-cache.store.ts
15Cache Persistence GhostsCache restart resurrects deleted entitiesPersistence disabled (--save "" --appendonly no); clean cold startdeployment.yaml
16Default noeviction ExhaustionCache becomes read-only when memory fills--maxmemory-policy volatile-lru evicts TTL data, protecting queuedeployment.yaml
17Silent Configuration FailureProcess boots with invalid NaN TTLBoot-time validation throws loud CacheConfigurationError — for module TTLs and for the TTLs services read themselvescache.module.ts
18Laundered Stale AggregatesCached aggregate computed from stale cacheAggregates are computed from product data, never from another aggregate; the batch fetches it with no-cacherecalculations.service.ts
19Parameter-Order Request DuplicateSame request cached twice under different keysrequestKey() sorts parameters and normalizes code-written namesrequest-key.ts
20Blanket Interceptor CachingUser-specific or error responses cachedExplicit opt-in via handler logic; throwing loaders never storeproducts.service.ts
21"Smart" List Invalidation BugsCategories retain deleted ghost itemsWholesale invalidation: any write wipes all list shapesproducts.service.ts
22Resurrected Negative 404sNewly created entity serves cached 404Every write path — creation included — invalidates the entity key, and every write goes through the ownerusers.service.ts
23Free-Text Search KeysUnbounded key explosion exhausts memoryFree-text search parameters are refused by builders; query directrequest-key.ts
24Normalizing Builders, Raw LoadersOne request for /v1/users/%201 caches a 404 under user 1's key, served by every consumerIdentifiers validated, never rewritten; non-canonical spellings served uncachedcache-keys.ts
25Synthetic Shape Collisions?category=all overwrites the unfiltered list with []Shapes are all and category.<name>; no category can name the synthetic oneproducts.service.ts
26Throwing Builders After the CommitA valid write returns 500 and blocks every later invalidationInvalidation uses cacheKeyOrNull, built after the responseproducts.service.ts
27Fail-Open Without a DeadlineRequests hang on a connected but unresponsive cache while /health looks fine; in-flight writes are re-sent after reconnectA per-command timeout; three in a row mark the cache down and reconnect; a clean close counts as down; in-flight commands are never re-sentredis-cache.store.ts
28No-Cache Joining a Stale FlightA read-your-writes request returns the pre-write valueno-cache starts its own load; superseded loads do not writeread-through.service.ts
29Lost Operator WakeupsWork accepted after a drain's last pop — or orphaned before a restart — never runsThe drain releases the lock and re-checks the queue; reconciliation runs on an interval, not once at startuprecalculations.service.ts
30Unbounded Un-Evictable QueueA flood of ids evicts every service's entries, then writes failValidated, size-limited enqueue and a MAX_QUEUE_SIZE caprecalculations.service.ts
31Rolling Update of a Single CacheTwo unrelated caches behind one Service during a rolloutstrategy: Recreatedeployment.yaml
32Reconnect Loop After ShutdownA pod stopped during a cache outage hangs until it is killedShutdown sends QUIT, then disconnects; a refused QUIT no longer leaves the client redialingredis-cache.store.ts

8. Scalability & Evolution Path​

8.1 From Single Instance to High Availability​

  • Current State: A single Valkey replica (replicas: 1, strategy: Recreate) without persistent disks. If the pod restarts, memory is cleared; services fail open for the duration of the restart window, keys are re-populated on demand, and queued background work is lost (§6.2).
  • Evolution Path (Primary-Replica Failover): When a single instance's availability becomes the limiting factor, the architecture transitions to a primary-replica topology with automated failover (Valkey Sentinel, or a managed service such as AWS ElastiCache / GCP Memorystore). Replication is asynchronous, so a failover can lose the most recent writes — including invalidations, which leaves entries the owner already deleted in place until their TTL expires — and recent queue additions.
  • Evolution Path (Sharded Cluster): When throughput or memory saturates, keys are distributed across shards. Without hash tags, Valkey Cluster already hashes the whole key, so a service's entries spread naturally; a hash tag such as {users-service}_user_1 would do the opposite and concentrate every key of a service onto one shard. The real work is elsewhere: multi-key commands like delMany must be split per hash slot, because the cluster rejects commands that span slots, and the ACL rules must be applied on every node. An application-transparent sharding proxy (e.g., Envoy) avoids the client changes but terminates authentication, so the servers see the proxy's identity rather than each service's — per-service ACL enforcement then has to move into the proxy, or it is lost. Modifying key delimiters also necessitates a coordinated update across the central key registry, client key builders, and GitOps ACL pattern rules (~users-service_user_*).

8.2 Work Queue Evolution​

  • The Durability Boundaries: The queue survives worker restarts but not cache restarts (persistence is off). And SPOP extracts IDs before background calculation finishes: if a pod experiences an ungraceful SIGKILL mid-drain, popped IDs are lost until re-triggered.
  • Evolution Path: For workflows where either loss is unacceptable, move the queue to durable infrastructure while retaining the shared cache for entity reads: Valkey Streams with consumer groups (XREADGROUP, XACK, XAUTOCLAIM) on a persistent instance — each entry stays pending until acknowledged, and another worker can claim it after a crash — or a dedicated broker (RabbitMQ/Kafka). Patching the set with SMOVE into a processing set does not scale to batches: it moves one member per call.

9. Trade-off Analysis​

Architectural DecisionBenefit GainedConcrete Cost Paid
Shared Cache TierZero inter-service HTTP hops on hits; single cluster to monitor; cluster-wide deduplicationSingle shared availability domain and a single memory budget; requires strict ACL governance and per-participant capacity monitoring
Single-Writer OwnershipEliminates distributed invalidation buses and cross-service write conflictsConsumers cannot repair foreign entries on write; must call owner. The read-fill/invalidate race inside the owner remains: staleness is bounded by the TTL, not zero
Validated, Never-Rewritten KeysNo request can write another identifier's entryNon-canonical spellings of an id are served uncached
Wholesale List InvalidationEliminates stale list bugs, ghost categories, and complex invalidation codeWrites cause temporary list cache misses across all categories
Clock-Bounded CompositesEnables caching expensive cross-service composite views without bus coordinationUpstream updates can be stale for up to the composite TTL duration
Fail-Open Read PolicyCache crashes — and silent caches — never take down user-facing trafficOwners and their databases must absorb 100% of consumer reads during outages, including any per-item fan-out the cache was hiding
Fail-Closed Queue PolicyAccepted background jobs were stored when they were acceptedEnqueue endpoints return HTTP 503 during cache outages; a cache restart still drops the queue
volatile-lru EvictionProtects work queues from eviction under loadNon-TTL keys cannot be evicted, so the queue must be capped at enqueue
Zero PersistencePrevents ghost reads after cache restarts; instant recoveryCache is cold on startup; queued background work is lost with it

10. Conclusion​

High-efficiency caching in a microservice estate does not require complex distributed pub/sub invalidation networks or heavyweight service-mesh tiers. By combining a single shared cache engine with strict namespace ownership, server-level ACL isolation, and a principled ladder of invalidation strategies, teams can eliminate inter-service network hops while keeping their operational footprint radically simple.

The architecture shows that appropriate complexity is the decisive metric: exact invalidation for entities, blunt wholesale invalidation for lists, clock bounds for composites, fail-closed acceptance for background queues — and, on every rung, a TTL that bounds how wrong the cache can be when an invalidation is raced or lost.


References​