bpftrace Explanation¶
What this page covers
How a bpftrace script becomes kernel-executing BPF programs: the compile pipeline, the probe types and the kernel mechanisms behind them, where type information comes from (BTF, DWARF, C headers), the per-CPU map and aggregation engine, the sync/async split, how the language has evolved since 0.22, and the security model. For flags, defaults, probe tables and version facts see Reference. For recipes see How-to Guides. For the summary see the bpftrace hub.
Architecture¶
bpftrace is a compiler with an attached runtime, not an interpreter. The upstream README states the contract: LLVM is the compiler backend, and libbpf handles interaction with the Linux BPF subsystem. bcc is still a build dependency for a few residual paths, but loading and attaching go through libbpf (vendored as a git submodule since 0.25).
Script-to-Kernel Pipeline¶
The diagram shows the path from script text to attached programs and back to terminal output, using the real stage names from the source tree (src/parser.cpp, src/ast/passes/*, src/runtime).
flowchart TB
subgraph Input["Script input"]
BT["script.bt / -e one-liner<br/>config block, imports, macros, probes"]
IMP["import targets<br/>.bt / .h / .bpf.c"]
end
subgraph Frontend["bpftrace compiler (user space)"]
PARSE["bpftrace parser<br/>(own grammar, not Clang)"]
CLANG["libclang passes<br/>C definitions, #include, .bpf.c"]
TYPES["Semantic passes<br/>macro expansion, BTF/DWARF type resolution,<br/>type checker, probe expansion"]
CG["codegen_llvm<br/>LLVM IR"]
LLVM["LLVM BPF backend<br/>(LLVM 18-23)"]
end
subgraph Runtime["bpftrace runtime"]
LIBBPF["libbpf<br/>load maps + programs, create links"]
OUT["Output handlers<br/>text / JSON, map printing, symbolization (blazesym)"]
end
subgraph Kernel["Linux kernel (6.1+)"]
VER["BPF verifier + JIT"]
HOOKS["Attach points<br/>kprobe / fentry / tracepoint / uprobe / usdt / perf events"]
MAPS[("BPF maps<br/>per-CPU hash, LRU hash")]
RB[("BPF ring buffer<br/>async events")]
end
BT --> PARSE --> TYPES
IMP --> CLANG --> TYPES
TYPES --> CG --> LLVM -->|"BPF ELF object"| LIBBPF
LIBBPF -->|"bpf() syscall"| VER --> HOOKS
HOOKS -->|"update"| MAPS
HOOKS -->|"printf / print / exit"| RB
RB --> OUT
MAPS -->|"read on print or exit"| OUT
Design consequences:
- Own language, C interop on the side. bpftrace has its own parser. libclang is only used for C definitions,
#included headers and (since 0.26 stable imports).bpf.cfiles that are compiled and linked into the program. - Compiled once per invocation. Programs are real BPF objects checked by the verifier and JIT-compiled. There is no per-event interpretation.
- In-kernel aggregation.
count(),sum(),hist()and friends update per-CPU map slots on every hit. Only reduced results cross into user space. - Ring buffer for events. Since 0.24.0
BPF_MAP_TYPE_RINGBUFis mandatory.printf(),print()andexit()are events on that ring buffer. A perf buffer is still used forskboutput. -d STAGEexposes every step.ast,types,codegen,codegen-opt,libbpfandverifierstages print the intermediate forms, which is how verifier rejections are debugged.
Session Lifecycle¶
A bpftrace run is a short-lived session. The sequence shows the order in which begin, attachment, event delivery and teardown happen.
sequenceDiagram
participant U as Operator shell
participant B as bpftrace process
participant L as libbpf
participant K as Kernel (verifier, hooks)
U->>B: bpftrace script.bt
B->>B: parse, resolve types, codegen, LLVM compile
B->>L: load BPF object (maps + programs)
L->>K: bpf(BPF_PROG_LOAD), verifier + JIT
B->>B: run begin probes
L->>K: create links for every attach point
Note over B,K: attach failure is fatal unless missing_probes is warn or ignore
B-->>U: Attached N probes
K->>K: probe fires, update per-CPU maps
K-->>B: ring buffer events (printf, print, exit)
B-->>U: formatted output
U->>B: Ctrl-C (or exit() or -c child ends)
B->>L: detach links
B->>B: run end probes
B-->>U: print non-empty maps (print_maps_on_exit)
begin runs before other probes are attached and end after they are detached. Since 0.25 a script can have several begin and end probes, which run in declaration order. A script that only has begin/end probes exits automatically (since 0.24).
Probe Providers¶
Probe types map onto distinct kernel mechanisms. That is why the required kernel config list fans out and why some probe types need BTF. The full attach syntax table is in Reference.
| Provider | Kernel mechanism | Type information | Trade-off |
|---|---|---|---|
tracepoint |
Static tracepoints compiled into the kernel | BTF for args (since 0.25) |
Most stable interface across kernel versions |
rawtracepoint |
Same tracepoints, raw arguments | BTF (since 0.24) | Lower overhead, but arguments differ from tracepoint |
kprobe / kretprobe |
Dynamic kernel instrumentation; wildcards use kprobe_multi links; matching entry/return pairs collapse into a kprobe session program when the kernel supports it |
None: argN registers, manual casts |
Any traceable function, but function names are not a stable ABI |
fentry / fexit |
BPF trampolines | BTF gives typed args.<name> |
Near-zero overhead, arguments also visible on return |
uprobe / uretprobe |
User-space breakpoints on file offsets; uprobe_multi links for wildcards |
DWARF for args (only DWARF use left after 0.24) |
Works on any binary with symbols |
usdt |
User static probes; semaphores activated via kernel uprobe refcounting | Probe notes in the ELF | Stable application-defined interface |
profile / interval / software / hardware |
perf_event_open timers, software counters, PMCs |
None | Sampling rather than tracing every event |
iter |
BPF iterators over tasks, files, VMAs | BTF (ctx) |
Experimental; cannot mix with other probes |
watchpoint |
Hardware breakpoints | None | Limited number of debug registers |
Discovery is part of the design. bpftrace -l lists attach points that exist on this host, and bpftrace -lv prints argument types and struct layouts from BTF or DWARF, so a script can be written against what the running kernel actually has.
Type Information: BTF, DWARF, and Headers¶
bpftrace needs types to read struct fields safely. It has three sources, and the project has steadily moved toward BTF:
- Kernel BTF (
/sys/kernel/btf/vmlinux, module BTF under/sys/kernel/btf/). Upstream calls it the preferred source because it cannot drift from the running kernel.fentry,rawtracepoint,iter, and since 0.25 tracepointargsdepend on it.BPFTRACE_HAVE_BTFis defined so scripts can branch on it. - DWARF for user-space binaries. 0.24.0 dropped most DWARF support and kept uprobe argument parsing. 0.26.0 added DWARF-based statement probes (
uprobe:binary@file:line) anddw_ustack()stack unwinding for x86_64 builds with LLVM 21+. - C headers and definitions parsed by libclang. Mixing header definitions with BTF can cause redefinition conflicts, so upstream recommends using BTF exclusively where possible.
Why the kernel floor moved to 6.1
The dependency policy supports stable kernels plus the four most recent LTS kernels. With BTF, ring buffers, kprobe_multi, uprobe_multi and session probes all in use, 0.26.0 raised the documented minimum from 5.15 to 6.1. Older kernels should use an older bpftrace release.
Map and Aggregation Engine¶
The stdlib distinguishes two cost classes: cheap writes and expensive synchronous reads.
Efficient writes. count(), sum(), avg(), min(), max() and stats() use per-CPU map storage. Each CPU updates its own slot without locks, which avoids cross-CPU contention on hot paths. A plain @x++ on a shared map is not atomic (bpftrace does not generate atomic instructions for ++), so concurrent writers can lose updates.
The sequence shows why async reads are cheap and sync reads are not.
sequenceDiagram
participant CPUs as Kernel CPUs (hit path)
participant PM as Per-CPU map slots
participant R as interval consumer
participant U as bpftrace user space
CPUs->>PM: lock-free increment of the local CPU slot
Note over PM: no cross-CPU coordination while tracing
R->>U: print(@) is an async event on the ring buffer
U->>PM: read and sum all CPU slots in user space
R->>PM: if (@ > 10) is a sync read in BPF, iterates every CPU slot
The idiomatic shape pairs cheap producers on hot probes with a slower interval:s:N consumer doing async print(@) then clear(@). Casting a per-CPU value ((int64)@c) or comparing it forces the expensive synchronous read.
Declared maps (let @m = lruhash(1000);, stable since 0.25) choose the BPF map type and size explicitly. Undeclared maps get max_map_keys entries (default 4096). A full map drops updates, which is why upstream recommends declaring maps up front.
Cost Model by Operation¶
Documented economics, worth treating as design rules when composing scripts:
| Operation | Class | Documented behavior |
|---|---|---|
@ = count(); / @ = sum(x); on hit path |
write | per-CPU update, contention-free |
@x++ on a scalar map |
write | plain read-modify-write, may lose updates under concurrency |
print(@) from an interval clause |
read, async | event to user space, which sums the per-CPU slots |
if (@ > 10) or (int64)@ |
read, sync | explicitly flagged expensive, iterates all CPUs in BPF |
printf() / print() / exit() |
output, async | event on the ring buffer, handled later in user space |
clear(@) / zero(@) |
maintenance, async | resets aggregation windows between drains |
Async Event Delivery¶
Helpers are sync (run in the BPF program), async (sent as an event and handled later by the bpftrace process), or compile-time (resolved before load). Async delivery keeps probe execution cheap but has two documented side effects:
- Event loss. If the ring buffer fills faster than user space drains it,
printf()output is dropped.perf_rb_pagessizes the buffer. - Delayed exit.
exit()is handled asynchronously, so probes can keep firing briefly after it is called.
Maps are printed by reference, not by value. print(@) in a tight loop can show the final value repeatedly because user space reads the map after the loop finished.
This flowchart shows the two common patterns: producers writing maps continuously, and a consumer interval draining them.
flowchart LR
P1["kretprobe:vfs_read<br/>@bytes = hist(retval)"] --> M[("@bytes per-CPU map")]
P2["tracepoint:raw_syscalls:sys_enter<br/>@syscalls = count()"] --> M2[("@syscalls per-CPU map")]
M2 --> C["interval:1s<br/>print(@syscalls), clear(@syscalls)"]
C --> RB[("ring buffer")]
RB --> S["stdout, one line per second"]
M --> O["exit: print_maps_on_exit"] --> S2["stdout histogram"]
Language Evolution Since 0.22¶
The language has changed more between 0.22 (January 2025) and 0.27 (September 2026) than in the previous five years. The direction is toward a typed, modular language that can host larger tools, while keeping one-liners short.
| Theme | Change | Version |
|---|---|---|
| Variables | let declarations and block scoping for $vars |
0.22 |
| Control flow | for ($i : 0..N) ranges, break, continue; while deprecated |
0.24, 0.25 |
| Literals and types | true/false, duration literals (1s, 100ms), records, tuple indexing, no implicit integer promotion, default literal type int8 |
0.24 to 0.26 |
| Code reuse | Hygienic macros (macro name(...) { ... }), stable in 0.25; macro type substitution with typeof in 0.27 |
0.24 to 0.27 |
| Modules | import "file.bt", .h, .bpf.c imports, stable in 0.26 |
0.25 to 0.26 |
| Pointers | . auto-dereference is the preferred field access (replaces ->); & address-of stable in 0.26 |
0.25, 0.26 |
| Maps | Map declarations with explicit BPF map types | 0.24 (stable 0.25) |
| Tooling | --fmt formatter, test: and bench: probes, --probe-filter |
0.25 |
| Call syntax | Builtins callable as functions (comm(), pid()); stdlib merged into one "helpers" namespace |
0.24 |
Most macros in the stdlib (assert, ppid, is_err, container_of, str_concat) are themselves written as bpftrace macros. This is why new helpers now appear in almost every release.
Distribution Modes¶
bpftrace reaches hosts in three ways:
- Distro packages (
apt,dnf,apk,pacman,zypper,emerge, nixpkgs). They match the distro's kernels and LLVM, but lag upstream. Ubuntu and Debian stable releases typically carry a version several minors behind, so scripts written against the current docs may not parse. - Static AppImage release assets for x86_64 and (since 0.27.0) arm64. They bundle LLVM and libbpf, so they give current upstream behavior on any distro with a supported kernel.
- Source builds: Nix flake (recommended upstream, also used by CI) or a distro build with CMake. The builder owns the LLVM/libbpf/bcc matrix.
An ahead-of-time path also exists in the source tree: an undocumented --aot FILE option writes an artifact that runs under a separate bpftrace-aotrt runtime. The AOT header stores a hash of the bpftrace version string, and the runtime refuses artifacts from any other version. The artifact is therefore tied to one bpftrace build, not only to a kernel. The flag is absent from the man page and --help, so treat it as experimental.
Relationship to bcc and libbpf Tools¶
Upstream positions bpftrace and bcc as complementary. Exploration starts with bpftrace one-liners, moves to ad-hoc bpftrace scripts, and only moves to bcc or libbpf C/Rust tools when a tool needs rich argument parsing, a long-lived daemon, or custom user-space logic. The docs cite xfsdist as an example: 22 lines in bpftrace versus 131 lines in bcc's Python version. Imports, macros, getopt() and .bpf.c linking (0.24 to 0.26) push that boundary out, but a persistent fleet agent is still better served by a compiled libbpf program. See the eBPF learning-path comparison.
Benchmarks¶
The upstream docs publish no throughput or overhead numbers, and none are invented here. The documented performance characteristics are: per-CPU aggregation avoids write contention, synchronous map reads are expensive, async output keeps formatting off the probed code path, and fentry trampolines have lower overhead than kprobes. Since 0.25, bench: probes (bpftrace --bench) measure the average nanoseconds of a code block, which is the upstream way to compare idioms (for example count() vs lhist()) on your own hardware. Absolute overhead depends on event rate and action size, so measure on the target host.
Security Model¶
Security posture of operating bpftrace: privileges, what tracing exposes, third-party script risk, and kernel hardening interactions.
Capability Model¶
bpftrace loads and attaches real BPF programs, so every probe type needs elevated privileges. Since 0.25.0 bpftrace checks for CAP_BPF, CAP_PERFMON, CAP_DAC_READ_SEARCH and CAP_DAC_OVERRIDE instead of requiring uid 0. The full table is in Reference. A separate --unsafe flag gates destructive helpers (system(), signal(), override(), write_user()) that change system state rather than observe it.
Container runtimes need explicit capability grants (--privileged or targeted capabilities plus access to /sys/kernel/tracing and BTF). Namespace-confined root without BPF and perf capabilities fails at load, not at parse. Since 0.23, pid, tid and ustack report values from bpftrace's own PID namespace; pid(init) and tid(init) (0.24+) give the initial-namespace view.
Data Exposure While Tracing¶
A tracing language is a data-exfiltration primitive by construction:
str(args.filename)prints filesystem paths, including tokens passed to open calls.- uprobe/USDT probes capture user-space function arguments: credential buffers, request bodies, serialization inputs.
sslsnoop.btandbashreadline.btin the bundled tools show how easy this is. - kretprobe histograms (
hist(retval)) leak distributions that are themselves sensitive on multi-tenant hosts.
Operate bpftrace only on systems you own or are authorized to instrument. Govern incident usage with the same approvals as packet capture: same blast radius, different layer.
Third-Party Script Risk¶
.bt scripts look inert but compile to programs with kernel read access under your privileges. Treat a borrowed script like a root shell:
- Read every probe clause before running and check that the targets match the claimed purpose.
- Watch for broad wildcards (
kprobe:*-class matches instrument far more than needed) and for--unsaferequirements. - Check
importstatements: imported.btand.bpf.cfiles change what runs. bpftrace refuses imports from world-writable directories. - Prefer vendoring community scripts into a reviewed repo over piping from gists. Treat updates as code review events.
bpftrace -l script.bt lists the probes a script would attach, and --dry-run loads and attaches without running. Both help audit a script before a real session.
Kernel Hardening Interactions¶
Hardened hosts intentionally restrict what bpftrace needs. Expect friction and document exceptions rather than loosening globally:
- Kernel lockdown (often enabled with Secure Boot) blocks bpftrace. The upstream developer guide lists disabling Secure Boot,
mokutil --disable-validation, or temporarily lifting lockdown as the options. bpftrace has a dedicated lockdown check (src/lockdown.cpp) to report this clearly. - Vendor kernels (cloud images, trimmed builds) may ship without
CONFIG_KPROBE_EVENTS,CONFIG_UPROBE_EVENTSor BTF. The readiness check lives in How-to Guides. kernel.unprivileged_bpf_disableddoes not affect bpftrace, which assumes privileged operation.- auditd/seccomp profiles will see unusual
bpf()andperf_event_open()syscall traffic when bpftrace runs. Allowlist deliberately.
For permanent fleet deployment prefer purpose-built daemons compiled against libbpf over ad-hoc bpftrace sessions: smaller privilege surface, reviewed binaries, and no compiler present on hosts.
Overhead as an Availability Concern¶
Aggressive probing is a self-inflicted outage vector on busy systems. Mitigations grounded in documented mechanics:
- Keep hit-path actions minimal. Aggregate in per-CPU maps instead of emitting a
printf()per event. - Drain maps asynchronously on slow intervals. Sync per-CPU reads inside hot clauses are flagged expensive.
- Prefer
fentryoverkprobeand tracepoints over kprobes where both exist. - Sample (
profile:hz:99,software:faults:100) rather than record when the question tolerates estimation. - Respect
max_probes(default 1024). Raising it is documented as able to cause high overhead or even freeze the system.
Lifecycle Hygiene¶
Ctrl-C detaches programs. bpftrace does not pin objects, except iter:...:pin probes that pin to /sys/fs/bpf on purpose. After an abnormal exit or a killed wrapper, check for leftovers:
Sources¶
- bpftrace README — LLVM backend, libbpf runtime, install matrix
- docs/language.md — probes, BTF support, map declarations, imports, macros, complex tools
- docs/stdlib.md — per-CPU semantics, invocation modes
- CHANGELOG.md — version-by-version language changes
- docs/dependency_support.md — kernel floor and libbpf policy
- docs/developers.md — build paths, kernel lockdown troubleshooting
- src/aot/aot.cpp — AOT header and version lock