Linux CFS Scheduler, cgroups v2 & CPU Throttling Studio
Architect container CPU performance and eliminate Kubernetes micro-throttling: simulate Completely Fair Scheduler quota exhaustion across multi-threaded bursts, compare cgroups v1 vs cgroups v2 unified resource hierarchies, calculate GOMAXPROCS and JVM ActiveProcessorCount alignment, and synthesize production manifests.
Interactive CFS Period & Quota Throttling Timeline
Visualize how multi-threaded bursts consume container CPU quota early and cause hard latency freezes.
cgroups v1 vs cgroups v2 Unified Hierarchy Architecture
cgroups v2 resolves fundamental flaws in the legacy multi-hierarchy model. Review the structural differences:
| Feature / Controller | cgroups v1 (Legacy) | cgroups v2 (Modern Unified) | Production Impact |
|---|---|---|---|
| Hierarchy Structure | Multiple disjoint hierarchies (/sys/fs/cgroup/cpu, /memory) |
Single Unified Tree (/sys/fs/cgroup) |
Eliminates split-process mapping where CPU and Memory belonged to different tree paths. |
| CPU Quota & Period | cpu.cfs_quota_us & cpu.cfs_period_us |
cpu.max "$QUOTA $PERIOD" |
Single atomic write prevents temporary misconfigurations during dynamic scaling. |
| CPU Relative Weight | cpu.shares (1024 base, non-linear) |
cpu.weight (1 to 10000, 100 default) |
Intuitive proportional sharing: weight 200 gets exactly 2x the CPU of weight 100 under contention. |
| Page Cache Writeback | Attributed to Root cgroup (Broken!) | Attributed to Originating cgroup | Fixes dirty page cache starvation where one container filled RAM without being throttled. |
| OOM Killer Behavior | Kills random individual threads/processes | memory.oom.group = 1 |
Kills the entire container together, preventing half-dead zombie pods that pass liveness probes. |
| Pressure Stall Info (PSI) | None (Only raw CPU usage counters) | cpu.pressure, memory.pressure |
Provides exact percentage of time threads were stalled waiting for CPU runqueues or page faults. |
Kubernetes Pod Sizing & Runtime Processor Alignment
Calculate recommended GOMAXPROCS and JVM thread settings based on host cores and pod allocations:
Runtime Thread Allocation Analysis
Linux Kernel CFS Scheduler Tunables (/proc/sys/kernel/)
Deep-dive into the kernel sysctl knobs governing thread scheduling latency and preemption:
| Kernel Parameter | Default Value | Low-Latency Tuning | Architectural Purpose |
|---|---|---|---|
sched_latency_ns |
6,000,000 ns (6ms) | 4,000,000 ns (4ms) | Target scheduling period during which all runnable tasks are guaranteed to run at least once. |
sched_min_granularity_ns |
750,000 ns (0.75ms) | 500,000 ns (0.5ms) | Minimum time slice a thread runs before the CFS scheduler allows preemption. |
sched_wakeup_granularity_ns |
1,000,000 ns (1ms) | 500,000 ns (0.5ms) | Preemption threshold when a sleeping thread wakes up; lower values improve event response time. |
sched_migration_cost_ns |
500,000 ns (0.5ms) | 250,000 ns (0.25ms) | Protects CPU cache warmth: if a task ran within this window, the scheduler avoids migrating it to another core. |