Everything, Everywhere
Verified Specification | Standardized Formulas | Instant Precision
Secure & Private (Zero Data Retention) Free Access • No Sign-Up
Linux eBPF Profiler On-CPU & Off-CPU DWARF & Frame Pointers Zero Overhead Target

Linux Continuous Profiling, Stack Unwinding & FlameGraph Studio

Architect enterprise-scale eBPF continuous profiling pipelines: compare in-kernel Frame Pointer vs compact DWARF vs ORC stack unwinding, calculate profiler CPU tax and map memory budgets, evaluate On-CPU cycles vs Off-CPU blocking stalls, explore interactive SVG FlameGraphs, and synthesize production C/libbpf and Go collectors.

1,520
Samples / Sec (Host)
0.28%
Profiler CPU Tax
4.2 MB
Kernel Map Memory
1.4 MB/s
Fleet Ingest Egress

Interactive SVG FlameGraph Explorer

Click any frame to zoom in; click Reset to restore root view. Hover to inspect sample count and %.

Matched: 0 frames (0.0%)
Palette:
Hover over any stack frame to inspect... -
Total Profile Weight: 2,450 samples FlameGraph Width = Relative Time Spent

How to Read a Production FlameGraph

▪ X-Axis (Width): Represents total sample count or time spent in that function and its descendants. It does NOT represent chronological execution time. A wider frame means more CPU cycles or blocking duration occurred within that call path.

▪ Y-Axis (Depth): Shows the call stack depth from root (bottom or top depending on inversion) to leaf. The top-most frames with wide plateaus are the actual execution bottlenecks (the "top of the stack").

▪ Plateau Hunting: Look for wide frames that have NO children or narrow children. These flat tops are where the CPU is actually executing instructions directly rather than delegating to child functions.

▪ Icicle vs FlameGraph: Classic FlameGraphs place root at the bottom, growing upward. Icicle mode inverts the graph so root is at the top, mimicking modern browser DevTools and IDE call trees.

Stack Unwinding Engine Comparison

Continuous profilers rely on low-level CPU registers and ELF binary sections to reconstruct the sequence of calling functions. Review the exact architectural trade-offs between unwinding mechanisms:

Unwinding Mechanism Target Domain eBPF Kernel Compatibility Application CPU Tax Symbol Quality & Limitations
Frame Pointer (RBP)
-fno-omit-frame-pointer
Native C/C++, Go, Rust, JVM Native bpf_get_stackid() 1.0% - 2.0% throughput penalty (loss of 1 register on x86) Instant (~15ns). Breaks if any shared library (glibc) was compiled with frame pointer omission.
Compact DWARF Tables
Parca / PolarSignals BPF
Stripped native binaries, Go, Rust Kernel BPF Map Array Search 0.0% runtime penalty on target application Evaluates CFA and return addresses via compact 8-byte eBPF tables. Requires ~10-20MB RAM per binary for tables.
ORC (Oops Rewind Capability)
CONFIG_UNWINDER_ORC
Linux Kernel (vmlinux) Built into Linux 4.14+ 0.0% penalty Deterministic, linear lookup. Completely immune to compiler reordering optimizations in kernel space.
JIT / Managed Maps
/tmp/perf-<pid>.map
Node.js (V8), Java (JVM), Erlang Userspace Symbol Bridge 0.5% - 1.5% (symbol map dumping) Bridges dynamically emitted JIT memory blocks to method names. Requires Async-Profiler or V8 --perf-prof flag.
Fedora / Ubuntu Frame Pointer Revolution: Both Fedora (since release 38) and Ubuntu (since release 24.04 LTS) have officially changed their default distribution compiler flags to include -fno-omit-frame-pointer across all packages. This standardized choice unlocks zero-cost continuous profiling across entire cloud fleets without sacrificing debuggability.

The Anatomy of a Stack Frame on x86_64

+------------------------------------+ <-- High Memory Address | Caller Stack Frame Data | +------------------------------------+ | Return Instruction Pointer (RIP) | [RBP + 8] <-- Next instruction in caller +------------------------------------+ | Saved Frame Pointer (Previous RBP) | [RBP] <-- Points to caller's RBP +------------------------------------+ <-- Current RBP register value | Local Function Variables | | ... | +------------------------------------+ <-- Stack Pointer (RSP)

When an eBPF timer fires, the kernel reads the current RBP register, dereferences it to find the previous RBP, and reads (RBP + 8) to record the return address. Repeating this loop 32 times takes less than 300 CPU cycles—orders of magnitude faster than parsing DWARF bytecode.

Continuous Profiling Sizing & Overhead Calculator

Model fleet-wide resource utilization, eBPF map memory allocation, profiler CPU tax, and telemetry ingestion network bandwidth:

Fleet Size (Servers): 100
CPU Cores per Server: 16
Sampling Frequency: 19 Hz (Prime)
Max Stack Depth (Frames): 64 frames

Sizing Architecture Calculation Results

Host Sample Rate
304 samples/s
Fleet-Wide Throughput
30,400 samples/s
Raw Telemetry Rate
15.2 MB/s
Zstd Ingest Bandwidth
450 KB/s
Kernel BPF Map RAM
3.2 MB / host
Estimated Host CPU Tax
0.18%

On-CPU vs Off-CPU Profiling Mechanics

A complete latency investigation requires measuring both where threads are executing and where threads are waiting:

On-CPU Profiling (Burned Cycles)

Active Execution
  • Trigger Mechanism: Timer interrupt via perf_event_open with PERF_COUNT_SW_CPU_CLOCK.
  • Thread State: TASK_RUNNING actively executing on a core.
  • What It Identifies:
    • Excessive JSON serialization / deserialization
    • Regex catastrophic backtracking
    • Hot loop arithmetic & memory hashing
    • JVM Garbage Collection scavenge phases
  • When To Use: System CPU utilization is high (70%-100%) and you need to optimize CPU capacity or reduce server count.

Off-CPU Profiling (Blocked Stalls)

Waiting / Sleeping
  • Trigger Mechanism: Kernel scheduler tracepoint sched:sched_switch.
  • Thread State: TASK_INTERRUPTIBLE or TASK_UNINTERRUPTIBLE.
  • What It Identifies:
    • Lock contention (pthread_mutex_lock, futex)
    • Synchronous storage writes (fsync, WAL commits)
    • Database connection pool exhaustion
    • Network socket reads & external API timeouts
  • When To Use: P99 latency is spiking, users experience sluggishness, but server CPU utilization remains deceptively low (< 20%).

Kernel Scheduler State Machine Flow (Off-CPU Duration Tracking)

[ Thread A: Running ] ---( Calls mutex_lock / read / nanosleep )---> | v [ sched:sched_switch ] Fires in Kernel: 1. eBPF checks: prev_state != 0 (Thread is going to sleep) 2. bpf_get_stackid(&stack_map, ctx, BPF_F_USER_STACK) captures Call Stack 3. bpf_ktime_get_ns() records start_time in BPF Hash Map keyed by TID | |--- ( Thread A is sleeping / blocked on disk / waiting on lock ) --- | v [ sched:sched_switch ] Fires when Thread A is rescheduled back in: 1. eBPF looks up TID in start_time map 2. delta_ns = current_ktime - start_time 3. Accumulates delta_ns into stack trace duration counter map 4. Deletes TID from in-flight start_time map

Production eBPF Profiler Source Code

// Select an artifact above
Sponsored Utility
While You're Here
Sponsored Recommendations
Advertisement