Perspective

The case for mechanical sympathy in caching

Allocation discipline, data-oriented layout, and single-threaded replicated state machines — and why they decide whether your cache has a predictable tail or a hopeful one.

The Aeron Cache team · October 10, 2026

Your average latency is a comforting lie

Everyone quotes their average. Almost nobody is paged because of their average. You get paged because the 99.9th-percentile request landed on a stop-the-world pause, or waited on a lock that a background thread was holding, or triggered an allocation storm that tipped the young generation over right when traffic spiked. The tail is where caches actually hurt, and the tail is almost entirely a function of decisions made far below the API: how memory is laid out, how many objects you create per request, and how many threads are allowed to touch the same state.

“Mechanical sympathy” — is the idea that you get the best from a system by working with the grain of the hardware rather than against it. For a cache, that’s not an aesthetic preference. It’s the difference between a tail you can reason about and one you can only pray over. This is the worldview Aeron Cache is built on, and here’s what it actually means in the code.

Allocation is the tax you pay in jitter

On the JVM, the cost of an object is rarely the allocation itself — modern allocators are fast. The cost is later, when the garbage collector has to find and reclaim it, usually at a moment you don’t control. Every object you create per request is a small deposit into a fund the GC will eventually withdraw, with interest, as a pause. Reduce the allocations on the hot path and you reduce the frequency and severity of those withdrawals.

So Aeron Cache is built to be low-GC — and we’re careful with that word: it is low-GC and GC-light by design, not zero-GC, because nothing honest in a running system is strictly garbage-free. The techniques are the ones the Agrona and Aeron ecosystem has refined for years:

  • Off-heap and reusable buffers. State and messages live in Agrona buffers — ExpandableArrayBuffer, MutableDirectBuffer — that are written in place and reused, rather than a churn of short-lived byte arrays and wrapper objects.
  • Object pooling and reuse. Hot objects are pooled and recycled (DequeReusableObjectPool, a Reusable / ReusableLong contract) so steady-state request handling approaches a flat allocation profile instead of a sawtooth.
  • Flyweights over object graphs. Structured data is accessed through flyweights — for example a TimerDetailsFlyweight over a TTL timer record — which read fields directly out of a buffer instead of inflating them into a tree of Java objects that the GC then has to trace.

None of this is free to write. It’s more work than new-ing an object and moving on. The payoff is that the GC has less to chase, so the pauses that define your tail get rarer and shorter.

The wire format is part of the memory story

How data crosses the wire determines how much garbage you make the instant it arrives. A reflection-based JSON codec allocates an object graph for every message; that’s convenient at the edge and expensive in the hot path. Aeron Cache’s native transport uses SBE — Simple Binary Encoding — generated, zero-copy codecs: little-endian, schema-defined (schema id 7), with variable-length string encoding. You read a field straight out of the inbound buffer at a known offset. No intermediate object, no reflection, no per-message allocation to decode.

That’s why the transport and the memory discipline are the same conversation, not two. SBE over Aeron — on UDP or, for a co-located client, IPC — keeps the zero-copy property end to end, and the cache’s internal buffers keep it from there. We pull this thread all the way through in Aeron + SBE transport.

Data-oriented layout: respect the cache line

Hardware is fast when the data it needs is already in the CPU cache and laid out predictably. It stalls when it has to chase pointers across the heap to assemble one logical record. Data-oriented design means arranging state so the common operation walks contiguous, compact memory instead of a scattered object graph — which is exactly what buffer-backed records and flyweights give you. The win isn’t theoretical: predictable access patterns mean predictable timing, and predictable timing is the whole game for a tail-sensitive cache.

One thread, one truth: the single-threaded replicated state machine

Here’s the decision that ties it together. The clustered cache is a ClusteredService — AbstractCacheClusterService — driven by a single service agent thread. A single writer sounds like a bottleneck until you account for everything it removes: no locks, no lock contention, no cache-line ping-pong between cores fighting over the same state, no memory barriers on every access, and no concurrency bugs in the state transitions. One thread applying operations in order is both faster in the common case and far easier to reason about.

It also happens to be exactly what RAFT wants. Consensus replicates an ordered log of commands; a deterministic state machine that applies that log in order on one thread will arrive at identical state on every node. That determinism is what makes snapshots trustworthy and recovery boring — a new leader rebuilds from the latest snapshot plus the tail of the log and gets the same answer. You tune the agent’s responsiveness-versus-CPU tradeoff with Aeron’s configurable IdleStrategy rather than by adding threads. The consensus mechanics live in RAFT consensus; the full treatment of the low-GC design is in low-GC design.

The takeaway

Mechanical sympathy in a cache isn’t a benchmark you post; it’s a set of constraints you accept so that the worst case stays near the average case. Allocate less, so the GC interrupts you less. Decode in place, so arrival is cheap. Lay data out for the cache line, so access is predictable. Put one thread in charge of the truth, so the state machine is deterministic and the cluster agrees. We don’t publish latency numbers — you should measure your own workload, and we’d rather earn your trust than spend it on a graph — but these are the design choices that give a cache a tail you can actually plan around.