Deep Dive

Designing for low GC (mechanical sympathy)

Agrona buffers, object pools, flyweights, single-threaded agents and configurable idle strategies — performance built in, not tuned in.

Why it matters

Garbage collection is the quiet killer of tail latency. A cache can be algorithmically perfect and still stutter every time the JVM pauses to collect the debris of a million short-lived request objects. The usual answer is to tune the collector after the fact. Aeron Cache takes the harder, more durable route: it is written so the hot path barely produces garbage in the first place. This is mechanical sympathy — code shaped to how CPUs, caches, and the memory system actually work.

A precise claim, because precision matters: this is a low-GC / GC-light design, not a zero-GC one. The service still allocates during startup, snapshot handling, and cold paths, and the JVM will still collect. What the design eliminates is steady-state, per-request allocation on the request/response path — the allocation pattern that drives collection frequency and pause jitter under load.

Off-heap buffers with Agrona

Aeron Cache builds on Agrona, the low-level buffer and data-structure library beneath Aeron. Instead of Java byte[] churn and ByteBuffer ceremony, the service works through Agrona’s MutableDirectBuffer abstractions. AbstractCacheClusterService holds its egress and snapshot buffers as long-lived ExpandableArrayBuffer instances, created once and reused for the life of the service:

private final MutableDirectBuffer egressBuffer = new ExpandableArrayBuffer();
private final MutableDirectBuffer snapshotBuffer = new ExpandableArrayBuffer();

Every response is encoded into that same egressBuffer and offered to the client session. An ExpandableArrayBuffer grows if a message needs more room, then stays grown — so after warm-up the buffer no longer resizes and encoding allocates nothing.

Object pools instead of new

Where per-operation objects are genuinely needed, Aeron Cache recycles them. The DequeReusableObjectPool is a single-threaded pool (no synchronization, because the service agent is single-threaded) pre-populated with instances; acquire() hands one back, clear()-ing and returning it on release():

private static final int TIMER_DETAILS_POOL_INITIAL_SIZE = 128;

this.timerDetailsPool = new DequeReusableObjectPool<>(
        () -> new TimerDetails<>(timerIndexSupplier.get(), timerKeySupplier.get()),
        TIMER_DETAILS_POOL_INITIAL_SIZE, true);

When a client asks for all pending TTL timers, the service acquires TimerDetails from the pool, fills them, streams them out, and returns every one:

private void returnTimerDetailsToPool() {
    var timers = allTimersResult.getTimers();
    for (TimerDetails<I, K> timer : timers) {
        timerDetailsPool.release(timer);
    }
}

The same philosophy runs through the request/result objects. Rather than decoding each command into a fresh object graph, the service holds one pre-allocated request details and result instance per message type — addCacheEntryRequestDetails, patchValueRequestDetails, incrementCounterResult, and dozens more — all constructed once in the constructor and reused on every message. The common contract is Reusable:

public interface Reusable<T> {
    void clear();
    void copyFrom(T source);
    T value();
}

Decode populates a reusable; process reads it; the next message clear()s and refills it. No allocation per request.

Flyweights over the wire and over state

A flyweight is an object that holds no data of its own — it is a typed window onto bytes that live elsewhere. The SBE codecs (see Aeron + SBE transport) are flyweights over the message buffer: fields are read at fixed offsets with zero copying. Aeron Cache applies the same idea to its own state. TimerDetailsFlyweight is a thin view used to drive timer expiry without materializing intermediate objects:

public class TimerDetailsFlyweight<I extends Reusable, K extends Reusable> {
    I cache;
    K key;
    long correlationId;
}

When a timer fires, the service reads cache id, key, and correlation id straight off the flyweight and performs the removal — no garbage produced to express “this entry expired.”

Single-threaded agents

The entire cluster service runs on one thread — the Aeron Cluster agent. This is the quiet superpower behind everything above: because there is exactly one thread touching the state, the pools need no locks, the reusable request/result objects are safe to share, and the data access pattern is naturally cache-friendly (one thread, one working set, good locality). It also makes behavior deterministic, which is precisely what RAFT replication requires (see RAFT consensus). Concurrency bugs that plague lock-based caches simply cannot occur, because there is no concurrency to get wrong.

Configurable idle strategies

A single-threaded agent spins in a loop asking “any work?”. How it waits when the answer is “no” is a direct latency-vs-CPU trade-off, and Aeron Cache makes it configurable rather than hard-coded. IdleStrategyConfig resolves an Agrona IdleStrategy from an environment spec, validated eagerly at startup:

switch (alias) {
    case BusySpinIdleStrategy.ALIAS:     return new BusySpinIdleStrategy();
    case BackoffIdleStrategy.ALIAS:      return new BackoffIdleStrategy();
    case YieldingIdleStrategy.ALIAS:     return new YieldingIdleStrategy();
    case NoOpIdleStrategy.ALIAS:         return new NoOpIdleStrategy();
    case SleepingIdleStrategy.ALIAS:     return new SleepingIdleStrategy(/* period */);
    // sleep-ms, or a fully-qualified IdleStrategy class name
}

Pick spin (busy-spin) to burn a core for the lowest possible wake-up latency; pick sleep-ms to give the CPU back on an idle dev box; backoff splits the difference. The idle strategy is threaded right into the response path — after offering a message the service calls idleStrategy.idle() — and the sendMessage back-pressure loop idles rather than busy-failing when a session’s publication is full:

void sendMessage(final ClientSession session, MutableDirectBuffer msgBuffer, int len) {
    long offered;
    while (session != null && (offered = session.offer(msgBuffer, 0, len)) < 0) {
        publicationFailureHandler.handleOfferFailure(offered);
        idleStrategy.idle();
    }
}

Performance by design

None of these choices is a bolt-on. Off-heap buffers, reusable request/result objects, pooled timer details, flyweight state, a single-threaded lock-free agent, and a tunable idle strategy compound into a system whose steady-state allocation is low by construction. That is why there are no magic GC flags to discover here — the garbage was designed out before the collector ever ran.

Takeaway: low-GC is an architecture, not a setting. Allocate at startup, reuse forever, and keep one thread honest — then the collector has almost nothing to do.