Deep Dive

RAFT consensus on Aeron Cluster

How Aeron Cluster drives the cache as a deterministic replicated state machine, and why determinism is the whole game.

Why it matters

A cache that forgets everything when a node dies is a convenience. A cache that survives node loss without losing a single acknowledged write is infrastructure. Aeron Cache is the second kind, and it gets there by refusing to treat replication as an afterthought. The store is not a map with a backup job — it is a deterministic replicated state machine driven by a consensus-ordered log. Every node that replays the same log arrives at byte-identical state. That single property is what makes high availability, snapshots, and even TTL expiry correct.

Aeron Cluster gives you RAFT

Aeron Cluster implements the RAFT consensus algorithm over Aeron’s reliable messaging (see Aeron + SBE transport). A cluster elects a leader; the leader sequences client commands into a replicated log and waits for a majority of nodes to persist each entry before it is committed. Followers replay the committed log in the same order. If the leader fails, the surviving majority elects a new one and continues — a three-node cluster tolerates the loss of any one node.

Your code never implements elections, log replication, or quorum. You implement a ClusteredService. In Aeron Cache that contract is satisfied by AbstractCacheClusterService, which Aeron Cluster calls back as the log is replayed:

public class AbstractCacheClusterService<I extends Reusable, K extends Reusable, V extends Reusable>
        implements ClusteredService {

    public void onSessionMessage(final ClientSession session, final long timestamp,
                                 final DirectBuffer buffer, final int offset,
                                 final int length, final Header header) {
        final int templateId = (buffer.getShort(offset + 2, java.nio.ByteOrder.LITTLE_ENDIAN) & 0xFFFF);
        if (templateId == schemaDetails.getCreateCacheId()) {
            handleCreateCache(session, buffer, offset, decoder, encoder, createCacheRequestDetails, cacheManager);
        } else if (templateId == schemaDetails.getAddCacheEntryId()) {
            handleAddCacheEntry(/* ... */);
        }
        // ... one branch per command template id
    }
}

The crucial detail: onSessionMessage does not run when a client sends a message. It runs when that message has already been committed to the log and replayed to this node. Every replica executes the exact same sequence of handleAddCacheEntry, handlePatchValue, handleRemoveCacheEntry calls, in the exact same order. Leader and follower converge because they run the same deterministic function over the same input tape.

Determinism is non-negotiable

If any handler branched on something a follower couldn’t reproduce — System.currentTimeMillis(), a random number, a thread race, iteration over an unordered collection — replicas would diverge, and a failover would silently corrupt state. Aeron Cache is written to the discipline that keeps replicas identical.

The sharpest example is TTL. A naive cache would stamp each entry with wall-clock now + ttl and sweep it on a background thread. In a replicated state machine that is a bug: every node’s wall clock differs, and the sweeper thread is non-deterministic. Aeron Cache instead reads time from the cluster and schedules a cluster timer:

if (ttl > 0 && addCacheEntryResult.getStatus() == CacheOperationStatus.SUCCESS) {
    long now = cluster.time();          // consensus time, identical on every node
    long deadline = now + ttl;
    timerService.scheduleItemRemoval(cacheId, key, cache, deadline);
}

cluster.time() is the cluster’s logical timestamp, agreed via the log — not the local system clock. The timer itself is registered with Aeron Cluster, and its firing arrives as another ordered log event:

public void onTimerEvent(final long correlationId, final long timestamp) {
    cacheTimerService.onTimerEvent(correlationId, timestamp);
    cacheCountersTimerService.onTimerEvent(correlationId, timestamp);
}

So expiry is just another committed command. Every replica removes the entry at the same logical instant, and the removal fans out to subscribers (see streaming subscriptions) exactly once. Because the timer lives in the state machine, it is also cancellable and introspectable — cancelItemRemoval cancels the scheduled timer, and handleGetAllTimers streams the pending set back to a client.

Snapshots: bounding replay

A log that only grows would make restart and new-node catch-up unbounded. Aeron Cluster periodically asks the service to snapshot, and Aeron Cache serializes its full state — caches, counter caches, and both timer wheels — into the snapshot publication:

public void onTakeSnapshot(final ExclusivePublication snapshotPublication) {
    int cumulativeLength = cacheTimerService.onTakeSnapshot(snapshotPublication, snapshotBuffer);
    snapshotPublication.offer(snapshotBuffer, 0, cumulativeLength);
    cacheManager.takeSnapshot(snapshotPublication, cluster);
    countersCacheManager.takeSnapshot(snapshotPublication, cluster);
}

On restart, onStart loads the latest snapshot and then replays only the log entries recorded after it — so recovery cost is bounded by the snapshot cadence, not by uptime. A recovering or freshly joined node reaches a consistent state the same way: load snapshot, replay tail, done.

Three deployment shapes

The same ClusteredService runs in three topologies, and the trade-off is purely operational:

  • Clustered (default). A Kubernetes StatefulSet of 3 replicas (replicaCount: 3) running the full RAFT cluster. This is the HA control-plane story: survive node loss, no acknowledged-write loss. Choose this for state that must outlive a pod.
  • Ephemeral. EphemeralCacheApplication — a single non-RAFT node. No quorum, no replication, lower overhead. Ideal for dev, CI, and genuinely ephemeral workloads where losing the cache on restart is acceptable.
  • Monolith. cache-monolith launches a single-node cluster plus the HTTP, WebSocket, and SSE gateways in one process. This is what brew install aeron-cache runs for local desktop use — the full API surface, zero cluster ceremony.

The through-line: the business logic is identical across all three. You develop against ephemeral or monolith, deploy to the 3-node cluster, and the determinism that makes the single node correct is exactly what makes the cluster consistent.

Takeaway: consensus gives you ordering; determinism gives you agreement. Aeron Cache earns both by making every mutation — including the passage of time — a committed, replayable log event.