Perspective

Why we built Aeron Cache

The gap between the low-latency messaging stack we trust in microseconds and the caches we tolerate everywhere else — and why we decided to close it.

The Aeron Cache team · October 10, 2026

There are two tiers of infrastructure, and only one of them got the memo

If you build trading systems, telemetry pipelines, or anything where the tail latency is the product, you already live in two worlds. In the hot path, you reach for infrastructure that was designed by people who count cache lines: Aeron for transport, Agrona for off-heap buffers and lock-free data structures, SBE for wire formats that don’t allocate. That stack treats the JVM and the hardware with respect, and it rewards you with latency you can actually reason about.

Then you need a cache. And somewhere between the message bus and the cache, the standards quietly drop. You accept a hop through a service that allocates on every request, serializes through a reflection-heavy codec, coordinates across threads with locks, and invalidates by telling every reader to go fetch the whole value again. It works. It is also the part of your architecture where predictability goes to die. We kept noticing that the cache was the least disciplined component in otherwise disciplined systems, and we got tired of apologizing for it.

Aeron Cache is our answer: a cache built on the same foundations we already trust for the hot path. This post is the honest version of why.

The gap: messaging got mechanical sympathy, caching didn’t

The Aeron ecosystem is the result of a particular worldview — that you get predictable performance by understanding the machine, not by throwing hardware at the problem. Off-heap buffers so the garbage collector has less to chase. Flyweights and codecs that read structured data in place instead of inflating it into object graphs. Single consumers on a ring buffer instead of lock contention. This is the lineage behind the “mechanical sympathy” movement, and it produces systems whose worst case is close to their average case.

Caches, meanwhile, mostly inherited a different tradition: convenience first, and let the GC and the network sort out the rest. For a lot of workloads that’s genuinely fine. But for the systems we care about — market data fan-out, pricing distribution, control-plane state that has to survive a node dying mid-flight — the cache’s jitter shows up directly in your p99. You can’t buy your way out of a GC pause that lands at the wrong moment, and you can’t un-send a full-value refresh to ten thousand subscribers.

So the gap is not “caches are slow.” The gap is that caches rarely bring the same engineering discipline to the problem that the rest of a low-latency stack takes for granted. We thought that discipline was portable.

The bet

Aeron Cache is a key-value and counter store built on Aeron, Agrona, and SBE. The bet behind it has three parts.

First: build on Aeron, not beside it. The native transport is an Aeron SBE gateway — schema-defined, little-endian, zero-copy codecs — running over UDP or IPC. If you’re already an Aeron shop, the cache speaks your language on the wire, and the IPC path means a co-located client can talk to the cache without touching the network stack at all. (We also expose HTTP, WebSocket — including a bidirectional endpoint — and SSE, because not every client is in the hot path, and not every client is Java.)

Second: make high availability a first-class property, not a bolt-on. The clustered deployment runs on Aeron Cluster, which gives us RAFT consensus. The cache itself is a ClusteredService — a deterministic replicated state machine driven by the replicated log, with snapshots and cluster-managed timers for TTL. That means your cached state survives node loss the same way a well-built control plane does: committed writes are replicated before they’re acknowledged, and a new leader rebuilds from the log and the latest snapshot. We’ll go deep on the mechanics in RAFT consensus. When you don’t need all that, there’s an ephemeral single-node mode and a monolith that runs the whole thing — cluster, HTTP, WS, SSE — in one process for local work.

Third: stream deltas, not invalidations. This is the part we’re proudest of. Aeron Cache supports JSON Merge Patch (RFC 7386) for partial updates, and it carries that idea all the way through to the subscription layer with patch-mode subscriptions: subscribers can receive only the PATCH_ITEM deltas — the fields that actually changed — instead of the whole value or a “go refetch” nudge. Paired with the client-side EmbeddedObjectCache, which deep-merges those deltas into a local JSON object, you get a cache that distributes change the way change actually happens: incrementally. That reshapes how you build config distribution and pricing fan-out, and it’s worth its own article — see patch-native streaming.

What this is good at

We designed for a specific shape of problem, and it’s worth being concrete about it:

  • Real-time pricing and market-data fan-out — streaming subscriptions, with patch-mode for tick deltas so you ship the one field that moved.
  • Feature flags and dynamic config — JSON merge patch plus patch-mode streaming plus the embedded object cache, so every service holds a locally-consistent copy that updates field-by-field.
  • Session and presence state with TTL — timed entries, with cancellable removal timers so a heartbeat can extend a lease instead of rewriting it.
  • Leaderboards, rate-limiting, and metrics — native counter caches with atomic increment/decrement and bulk operations.
  • HA control-plane state — the RAFT cluster and snapshots, for state that simply must outlive a node.

If your workload is none of these, a conventional cache may serve you better, and we’d rather tell you that now.

Honest about where we are

Aeron Cache is pre-1.0. The repository versions are 0.0.x-SNAPSHOT, and the official v1.0 release is coming soon. The architecture is in place — the RAFT clustered service, the SBE gateway, the streaming and patch surface, the polyglot embedded clients (Java, TypeScript, Python, Rust), the Rust CLI, the Helm charts, and the local Homebrew monolith are all real and in the source — but we are still hardening, and we are not going to pretend otherwise. We also deliberately publish no latency numbers yet. We describe the design for predictable tail behavior — low-GC, data-oriented, single-threaded state machines — and we’ll let you measure your own workload rather than wave a benchmark at you. It’s a low-GC design, not a zero-GC one, and we’ll always say it that way.

What we’re confident about is the premise: the discipline that makes low-latency messaging predictable belongs in the cache too. If that resonates, start at Getting Started, see where we land against the alternatives in comparison, and then come argue with us.