Skip to content

bpftrace Explanation

What this page covers

How a bpftrace script becomes kernel-executing BPF programs: the compile pipeline, the probe types and the kernel mechanisms behind them, where type information comes from (BTF, DWARF, C headers), the per-CPU map and aggregation engine, the sync/async split, how the language has evolved since 0.22, and the security model. For flags, defaults, probe tables and version facts see Reference. For recipes see How-to Guides. For the summary see the bpftrace hub.

Architecture

bpftrace is a compiler with an attached runtime, not an interpreter. The upstream README states the contract: LLVM is the compiler backend, and libbpf handles interaction with the Linux BPF subsystem. bcc is still a build dependency for a few residual paths, but loading and attaching go through libbpf (vendored as a git submodule since 0.25).

Script-to-Kernel Pipeline

The diagram shows the path from script text to attached programs and back to terminal output, using the real stage names from the source tree (src/parser.cpp, src/ast/passes/*, src/runtime).

flowchart TB
    subgraph Input["Script input"]
        BT["script.bt / -e one-liner<br/>config block, imports, macros, probes"]
        IMP["import targets<br/>.bt / .h / .bpf.c"]
    end

    subgraph Frontend["bpftrace compiler (user space)"]
        PARSE["bpftrace parser<br/>(own grammar, not Clang)"]
        CLANG["libclang passes<br/>C definitions, #include, .bpf.c"]
        TYPES["Semantic passes<br/>macro expansion, BTF/DWARF type resolution,<br/>type checker, probe expansion"]
        CG["codegen_llvm<br/>LLVM IR"]
        LLVM["LLVM BPF backend<br/>(LLVM 18-23)"]
    end

    subgraph Runtime["bpftrace runtime"]
        LIBBPF["libbpf<br/>load maps + programs, create links"]
        OUT["Output handlers<br/>text / JSON, map printing, symbolization (blazesym)"]
    end

    subgraph Kernel["Linux kernel (6.1+)"]
        VER["BPF verifier + JIT"]
        HOOKS["Attach points<br/>kprobe / fentry / tracepoint / uprobe / usdt / perf events"]
        MAPS[("BPF maps<br/>per-CPU hash, LRU hash")]
        RB[("BPF ring buffer<br/>async events")]
    end

    BT --> PARSE --> TYPES
    IMP --> CLANG --> TYPES
    TYPES --> CG --> LLVM -->|"BPF ELF object"| LIBBPF
    LIBBPF -->|"bpf() syscall"| VER --> HOOKS
    HOOKS -->|"update"| MAPS
    HOOKS -->|"printf / print / exit"| RB
    RB --> OUT
    MAPS -->|"read on print or exit"| OUT

Design consequences:

  • Own language, C interop on the side. bpftrace has its own parser. libclang is only used for C definitions, #included headers and (since 0.26 stable imports) .bpf.c files that are compiled and linked into the program.
  • Compiled once per invocation. Programs are real BPF objects checked by the verifier and JIT-compiled. There is no per-event interpretation.
  • In-kernel aggregation. count(), sum(), hist() and friends update per-CPU map slots on every hit. Only reduced results cross into user space.
  • Ring buffer for events. Since 0.24.0 BPF_MAP_TYPE_RINGBUF is mandatory. printf(), print() and exit() are events on that ring buffer. A perf buffer is still used for skboutput.
  • -d STAGE exposes every step. ast, types, codegen, codegen-opt, libbpf and verifier stages print the intermediate forms, which is how verifier rejections are debugged.

Session Lifecycle

A bpftrace run is a short-lived session. The sequence shows the order in which begin, attachment, event delivery and teardown happen.

sequenceDiagram
    participant U as Operator shell
    participant B as bpftrace process
    participant L as libbpf
    participant K as Kernel (verifier, hooks)
    U->>B: bpftrace script.bt
    B->>B: parse, resolve types, codegen, LLVM compile
    B->>L: load BPF object (maps + programs)
    L->>K: bpf(BPF_PROG_LOAD), verifier + JIT
    B->>B: run begin probes
    L->>K: create links for every attach point
    Note over B,K: attach failure is fatal unless missing_probes is warn or ignore
    B-->>U: Attached N probes
    K->>K: probe fires, update per-CPU maps
    K-->>B: ring buffer events (printf, print, exit)
    B-->>U: formatted output
    U->>B: Ctrl-C (or exit() or -c child ends)
    B->>L: detach links
    B->>B: run end probes
    B-->>U: print non-empty maps (print_maps_on_exit)

begin runs before other probes are attached and end after they are detached. Since 0.25 a script can have several begin and end probes, which run in declaration order. A script that only has begin/end probes exits automatically (since 0.24).

Probe Providers

Probe types map onto distinct kernel mechanisms. That is why the required kernel config list fans out and why some probe types need BTF. The full attach syntax table is in Reference.

Provider Kernel mechanism Type information Trade-off
tracepoint Static tracepoints compiled into the kernel BTF for args (since 0.25) Most stable interface across kernel versions
rawtracepoint Same tracepoints, raw arguments BTF (since 0.24) Lower overhead, but arguments differ from tracepoint
kprobe / kretprobe Dynamic kernel instrumentation; wildcards use kprobe_multi links; matching entry/return pairs collapse into a kprobe session program when the kernel supports it None: argN registers, manual casts Any traceable function, but function names are not a stable ABI
fentry / fexit BPF trampolines BTF gives typed args.<name> Near-zero overhead, arguments also visible on return
uprobe / uretprobe User-space breakpoints on file offsets; uprobe_multi links for wildcards DWARF for args (only DWARF use left after 0.24) Works on any binary with symbols
usdt User static probes; semaphores activated via kernel uprobe refcounting Probe notes in the ELF Stable application-defined interface
profile / interval / software / hardware perf_event_open timers, software counters, PMCs None Sampling rather than tracing every event
iter BPF iterators over tasks, files, VMAs BTF (ctx) Experimental; cannot mix with other probes
watchpoint Hardware breakpoints None Limited number of debug registers

Discovery is part of the design. bpftrace -l lists attach points that exist on this host, and bpftrace -lv prints argument types and struct layouts from BTF or DWARF, so a script can be written against what the running kernel actually has.

Type Information: BTF, DWARF, and Headers

bpftrace needs types to read struct fields safely. It has three sources, and the project has steadily moved toward BTF:

  1. Kernel BTF (/sys/kernel/btf/vmlinux, module BTF under /sys/kernel/btf/). Upstream calls it the preferred source because it cannot drift from the running kernel. fentry, rawtracepoint, iter, and since 0.25 tracepoint args depend on it. BPFTRACE_HAVE_BTF is defined so scripts can branch on it.
  2. DWARF for user-space binaries. 0.24.0 dropped most DWARF support and kept uprobe argument parsing. 0.26.0 added DWARF-based statement probes (uprobe:binary@file:line) and dw_ustack() stack unwinding for x86_64 builds with LLVM 21+.
  3. C headers and definitions parsed by libclang. Mixing header definitions with BTF can cause redefinition conflicts, so upstream recommends using BTF exclusively where possible.

Why the kernel floor moved to 6.1

The dependency policy supports stable kernels plus the four most recent LTS kernels. With BTF, ring buffers, kprobe_multi, uprobe_multi and session probes all in use, 0.26.0 raised the documented minimum from 5.15 to 6.1. Older kernels should use an older bpftrace release.

Map and Aggregation Engine

The stdlib distinguishes two cost classes: cheap writes and expensive synchronous reads.

Efficient writes. count(), sum(), avg(), min(), max() and stats() use per-CPU map storage. Each CPU updates its own slot without locks, which avoids cross-CPU contention on hot paths. A plain @x++ on a shared map is not atomic (bpftrace does not generate atomic instructions for ++), so concurrent writers can lose updates.

The sequence shows why async reads are cheap and sync reads are not.

sequenceDiagram
    participant CPUs as Kernel CPUs (hit path)
    participant PM as Per-CPU map slots
    participant R as interval consumer
    participant U as bpftrace user space
    CPUs->>PM: lock-free increment of the local CPU slot
    Note over PM: no cross-CPU coordination while tracing
    R->>U: print(@) is an async event on the ring buffer
    U->>PM: read and sum all CPU slots in user space
    R->>PM: if (@ > 10) is a sync read in BPF, iterates every CPU slot

The idiomatic shape pairs cheap producers on hot probes with a slower interval:s:N consumer doing async print(@) then clear(@). Casting a per-CPU value ((int64)@c) or comparing it forces the expensive synchronous read.

Declared maps (let @m = lruhash(1000);, stable since 0.25) choose the BPF map type and size explicitly. Undeclared maps get max_map_keys entries (default 4096). A full map drops updates, which is why upstream recommends declaring maps up front.

Cost Model by Operation

Documented economics, worth treating as design rules when composing scripts:

Operation Class Documented behavior
@ = count(); / @ = sum(x); on hit path write per-CPU update, contention-free
@x++ on a scalar map write plain read-modify-write, may lose updates under concurrency
print(@) from an interval clause read, async event to user space, which sums the per-CPU slots
if (@ > 10) or (int64)@ read, sync explicitly flagged expensive, iterates all CPUs in BPF
printf() / print() / exit() output, async event on the ring buffer, handled later in user space
clear(@) / zero(@) maintenance, async resets aggregation windows between drains

Async Event Delivery

Helpers are sync (run in the BPF program), async (sent as an event and handled later by the bpftrace process), or compile-time (resolved before load). Async delivery keeps probe execution cheap but has two documented side effects:

  1. Event loss. If the ring buffer fills faster than user space drains it, printf() output is dropped. perf_rb_pages sizes the buffer.
  2. Delayed exit. exit() is handled asynchronously, so probes can keep firing briefly after it is called.

Maps are printed by reference, not by value. print(@) in a tight loop can show the final value repeatedly because user space reads the map after the loop finished.

This flowchart shows the two common patterns: producers writing maps continuously, and a consumer interval draining them.

flowchart LR
    P1["kretprobe:vfs_read<br/>@bytes = hist(retval)"] --> M[("@bytes per-CPU map")]
    P2["tracepoint:raw_syscalls:sys_enter<br/>@syscalls = count()"] --> M2[("@syscalls per-CPU map")]
    M2 --> C["interval:1s<br/>print(@syscalls), clear(@syscalls)"]
    C --> RB[("ring buffer")]
    RB --> S["stdout, one line per second"]
    M --> O["exit: print_maps_on_exit"] --> S2["stdout histogram"]

Language Evolution Since 0.22

The language has changed more between 0.22 (January 2025) and 0.27 (September 2026) than in the previous five years. The direction is toward a typed, modular language that can host larger tools, while keeping one-liners short.

Theme Change Version
Variables let declarations and block scoping for $vars 0.22
Control flow for ($i : 0..N) ranges, break, continue; while deprecated 0.24, 0.25
Literals and types true/false, duration literals (1s, 100ms), records, tuple indexing, no implicit integer promotion, default literal type int8 0.24 to 0.26
Code reuse Hygienic macros (macro name(...) { ... }), stable in 0.25; macro type substitution with typeof in 0.27 0.24 to 0.27
Modules import "file.bt", .h, .bpf.c imports, stable in 0.26 0.25 to 0.26
Pointers . auto-dereference is the preferred field access (replaces ->); & address-of stable in 0.26 0.25, 0.26
Maps Map declarations with explicit BPF map types 0.24 (stable 0.25)
Tooling --fmt formatter, test: and bench: probes, --probe-filter 0.25
Call syntax Builtins callable as functions (comm(), pid()); stdlib merged into one "helpers" namespace 0.24

Most macros in the stdlib (assert, ppid, is_err, container_of, str_concat) are themselves written as bpftrace macros. This is why new helpers now appear in almost every release.

Distribution Modes

bpftrace reaches hosts in three ways:

  1. Distro packages (apt, dnf, apk, pacman, zypper, emerge, nixpkgs). They match the distro's kernels and LLVM, but lag upstream. Ubuntu and Debian stable releases typically carry a version several minors behind, so scripts written against the current docs may not parse.
  2. Static AppImage release assets for x86_64 and (since 0.27.0) arm64. They bundle LLVM and libbpf, so they give current upstream behavior on any distro with a supported kernel.
  3. Source builds: Nix flake (recommended upstream, also used by CI) or a distro build with CMake. The builder owns the LLVM/libbpf/bcc matrix.

An ahead-of-time path also exists in the source tree: an undocumented --aot FILE option writes an artifact that runs under a separate bpftrace-aotrt runtime. The AOT header stores a hash of the bpftrace version string, and the runtime refuses artifacts from any other version. The artifact is therefore tied to one bpftrace build, not only to a kernel. The flag is absent from the man page and --help, so treat it as experimental.

Relationship to bcc and libbpf Tools

Upstream positions bpftrace and bcc as complementary. Exploration starts with bpftrace one-liners, moves to ad-hoc bpftrace scripts, and only moves to bcc or libbpf C/Rust tools when a tool needs rich argument parsing, a long-lived daemon, or custom user-space logic. The docs cite xfsdist as an example: 22 lines in bpftrace versus 131 lines in bcc's Python version. Imports, macros, getopt() and .bpf.c linking (0.24 to 0.26) push that boundary out, but a persistent fleet agent is still better served by a compiled libbpf program. See the eBPF learning-path comparison.

Benchmarks

The upstream docs publish no throughput or overhead numbers, and none are invented here. The documented performance characteristics are: per-CPU aggregation avoids write contention, synchronous map reads are expensive, async output keeps formatting off the probed code path, and fentry trampolines have lower overhead than kprobes. Since 0.25, bench: probes (bpftrace --bench) measure the average nanoseconds of a code block, which is the upstream way to compare idioms (for example count() vs lhist()) on your own hardware. Absolute overhead depends on event rate and action size, so measure on the target host.

Security Model

Security posture of operating bpftrace: privileges, what tracing exposes, third-party script risk, and kernel hardening interactions.

Capability Model

bpftrace loads and attaches real BPF programs, so every probe type needs elevated privileges. Since 0.25.0 bpftrace checks for CAP_BPF, CAP_PERFMON, CAP_DAC_READ_SEARCH and CAP_DAC_OVERRIDE instead of requiring uid 0. The full table is in Reference. A separate --unsafe flag gates destructive helpers (system(), signal(), override(), write_user()) that change system state rather than observe it.

Container runtimes need explicit capability grants (--privileged or targeted capabilities plus access to /sys/kernel/tracing and BTF). Namespace-confined root without BPF and perf capabilities fails at load, not at parse. Since 0.23, pid, tid and ustack report values from bpftrace's own PID namespace; pid(init) and tid(init) (0.24+) give the initial-namespace view.

Data Exposure While Tracing

A tracing language is a data-exfiltration primitive by construction:

  • str(args.filename) prints filesystem paths, including tokens passed to open calls.
  • uprobe/USDT probes capture user-space function arguments: credential buffers, request bodies, serialization inputs. sslsnoop.bt and bashreadline.bt in the bundled tools show how easy this is.
  • kretprobe histograms (hist(retval)) leak distributions that are themselves sensitive on multi-tenant hosts.

Operate bpftrace only on systems you own or are authorized to instrument. Govern incident usage with the same approvals as packet capture: same blast radius, different layer.

Third-Party Script Risk

.bt scripts look inert but compile to programs with kernel read access under your privileges. Treat a borrowed script like a root shell:

  1. Read every probe clause before running and check that the targets match the claimed purpose.
  2. Watch for broad wildcards (kprobe:*-class matches instrument far more than needed) and for --unsafe requirements.
  3. Check import statements: imported .bt and .bpf.c files change what runs. bpftrace refuses imports from world-writable directories.
  4. Prefer vendoring community scripts into a reviewed repo over piping from gists. Treat updates as code review events.

bpftrace -l script.bt lists the probes a script would attach, and --dry-run loads and attaches without running. Both help audit a script before a real session.

Kernel Hardening Interactions

Hardened hosts intentionally restrict what bpftrace needs. Expect friction and document exceptions rather than loosening globally:

  • Kernel lockdown (often enabled with Secure Boot) blocks bpftrace. The upstream developer guide lists disabling Secure Boot, mokutil --disable-validation, or temporarily lifting lockdown as the options. bpftrace has a dedicated lockdown check (src/lockdown.cpp) to report this clearly.
  • Vendor kernels (cloud images, trimmed builds) may ship without CONFIG_KPROBE_EVENTS, CONFIG_UPROBE_EVENTS or BTF. The readiness check lives in How-to Guides.
  • kernel.unprivileged_bpf_disabled does not affect bpftrace, which assumes privileged operation.
  • auditd/seccomp profiles will see unusual bpf() and perf_event_open() syscall traffic when bpftrace runs. Allowlist deliberately.

For permanent fleet deployment prefer purpose-built daemons compiled against libbpf over ad-hoc bpftrace sessions: smaller privilege surface, reviewed binaries, and no compiler present on hosts.

Overhead as an Availability Concern

Aggressive probing is a self-inflicted outage vector on busy systems. Mitigations grounded in documented mechanics:

  • Keep hit-path actions minimal. Aggregate in per-CPU maps instead of emitting a printf() per event.
  • Drain maps asynchronously on slow intervals. Sync per-CPU reads inside hot clauses are flagged expensive.
  • Prefer fentry over kprobe and tracepoints over kprobes where both exist.
  • Sample (profile:hz:99, software:faults:100) rather than record when the question tolerates estimation.
  • Respect max_probes (default 1024). Raising it is documented as able to cause high overhead or even freeze the system.

Lifecycle Hygiene

Ctrl-C detaches programs. bpftrace does not pin objects, except iter:...:pin probes that pin to /sys/fs/bpf on purpose. After an abnormal exit or a killed wrapper, check for leftovers:

sudo bpftool prog show && sudo bpftool map show

Sources