Skip to content
LinuxDeep Dive Published Updated 6 min readViews unavailable

Linux BPF Ring Buffer: Ordering, Backpressure, and Event Delivery

Design BPF event pipelines with the shared ring buffer, including reservation ordering, verifier constraints, backpressure, and user-space consumption.

The BPF ring buffer is a kernel map type for transferring variable-sized records from BPF programs to user space. Its defining design choice is a shared multi-producer, single-consumer buffer: producers running on different CPUs reserve space in the same ring. Compared with a per-CPU perf-event buffer, this can use memory more efficiently and preserve the order in which reservations are made across CPUs.

That ordering is useful for correlated process lifecycle events, but it is not a durable log and it is not automatically ordered by an application’s wall-clock timestamp. A ring buffer is a bounded telemetry channel. If its consumer falls behind, producers must respond to reservation failure rather than blocking the kernel.

Choose reserve/submit or output deliberately

The reserve API returns a pointer directly into ring-buffer memory and avoids an extra copy. Its size must be known to the verifier at the reservation site. Every successful reservation must be committed or discarded; the verifier tracks the reference and rejects paths that leak it. The output API copies data from a buffer prepared elsewhere, which adds a copy but can handle record sizes that are not known at verification time.

#include "vmlinux.h"
#include <bpf/bpf_helpers.h>

struct event {
    __u64 timestamp_ns;
    __u32 pid;
    __u16 version;
    __u16 size;
    char comm[16];
};

struct {
    __uint(type, BPF_MAP_TYPE_RINGBUF);
    __uint(max_entries, 1 << 20); /* power-of-two byte capacity */
} events SEC(".maps");

static __always_inline int emit_event(__u32 pid)
{
    struct event *record;

    record = bpf_ringbuf_reserve(&events, sizeof(*record), 0);
    if (!record)
        return 0; /* record loss must be observable in the application */

    record->timestamp_ns = bpf_ktime_get_ns();
    record->pid = pid;
    record->version = 1;
    record->size = sizeof(*record);
    if (bpf_get_current_comm(record->comm, sizeof(record->comm)) < 0) {
        bpf_ringbuf_discard(record, 0);
        return 0;
    }
    bpf_ringbuf_submit(record, 0);
    return 0;
}

SEC("tracepoint/syscalls/sys_enter_openat")
int on_openat(void *ctx)
{
    (void)ctx;
    return emit_event((__u32)(bpf_get_current_pid_tgid() >> 32));
}

char LICENSE[] SEC("license") = "GPL";

The example is a complete BPF-side fragment for a libbpf build: generate vmlinux.h from the target BTF source and compile it with the libbpf headers. The map’s max_entries is its capacity in bytes and must be a power of two. The sample intentionally drops an event if reservation fails; a production program should count such losses through a separate counter map or another explicit health signal. It must not retry indefinitely or sleep in an attempt to make space.

The version and size fields make the record contract explicit. Keep the matching fixed-width layout in a shared header or mirrored, testable ABI definitions on the BPF and user-space sides. Do not add compiler-specific packed layout to save a few bytes without checking alignment and verifier access. If a schema change is not backward compatible, use a new version and make the consumer reject or route it deliberately; silently interpreting an old record with a new struct can produce plausible but corrupt telemetry.

Validate records at the consumer boundary

The libbpf callback receives a pointer into the mapped ring buffer. Validate the record before reading it, and copy it before returning if downstream work outlives the callback. The following is a consumer-side pattern; struct event is the matching user-space definition of the wire record, using uint64_t, uint32_t, and uint16_t from <stdint.h>:

static int on_sample(void *ctx, void *data, size_t len)
{
    struct event sample;

    (void)ctx;
    if (len != sizeof(sample)) {
        record_bad_sample_length(len);
        return 0; /* keep draining later records */
    }

    memcpy(&sample, data, sizeof(sample));
    if (sample.version != 1 || sample.size != sizeof(sample)) {
        record_unsupported_sample(sample.version, sample.size);
        return 0;
    }

    enqueue_or_process(&sample);
    return 0;
}

record_bad_sample_length, record_unsupported_sample, and enqueue_or_process stand for application-owned telemetry and processing. They must be bounded: a slow callback delays draining, so expensive parsing, disk I/O, or network export belongs in a bounded worker queue. If that queue fills, the application needs its own explicit drop/backpressure policy; the kernel ring being drained does not mean downstream processing kept up. For an asynchronous queue, copy the sample into queue-owned storage before callback return rather than retaining data.

Create the user-space manager with ring_buffer__new() using the BPF map file descriptor and this callback, then poll it with a finite timeout so the process can also handle shutdown and health reporting. Treat a negative libbpf result as an error to classify and log; free the manager during orderly shutdown. The timeout does not provide durability: process termination loses unread telemetry, and no callback can recover events that failed reservation in the kernel.

Reservation order creates a head-of-line constraint

Each producer reserves a record before filling it. A later record can be committed first, but the consumer cannot pass an earlier reservation that is still busy. A slow path between reserve and submit can therefore delay visible delivery of records behind it. Keep the reservation window short: collect expensive context before reserving when safe, populate the fixed record promptly, then submit or discard on every path.

The ordering guarantee is about reservation order, not a global claim that one CPU’s event physically occurred before another CPU’s. Add a monotonic timestamp and the identifiers needed to reconstruct application-level causality. If sharding into multiple ring buffers, document that each buffer has its own ordering domain.

Size for bursts and measure loss

Capacity is a burst budget, not a rate limit. As a first sizing estimate, multiply peak event rate by aligned record size and the longest consumer pause you intend to absorb, then add headroom for burstiness and other producers. This is not a guarantee: workload shape, record alignment, and callback stalls must be measured under representative load. Ring-buffer reservation fails without blocking when there is insufficient free space. Track failed reservations, bytes produced, records consumed, and consumer latency; an apparently healthy BPF program may otherwise be silently shedding the events needed for an incident investigation. A separate per-CPU counter map is one option for low-contention loss accounting; make the counter’s reset, aggregation, and exporter-failure behavior explicit.

User-space consumers can use libbpf’s ring-buffer manager and poll or wait for records. The callback should validate record length and version before decoding it, avoid unbounded work, and hand expensive processing to an application queue. A slow callback is part of the backpressure path. Memory mapping and epoll notification can reduce overhead, but neither removes the need for capacity planning or loss accounting.

When perf buffers may still be the right fit

The shared ring buffer is a good fit when cross-CPU reservation order or shared memory efficiency matters. Per-CPU perf buffers can remain preferable when independent per-CPU streams simplify contention or loss isolation. Compare the semantics your consumer needs, not just throughput in a microbenchmark. Migration from perf buffers also changes how records are ordered and how per-CPU loss is observed.

Treat the ring buffer as an explicitly lossy, bounded handoff. Define the loss signal, keep reserve-to-submit intervals short, version event records, and test what happens when the consumer pauses or exits.

Related:

Sources:

Comments