Everything, Everywhere
Verified Specification | Standardized Formulas | Instant Precision
Secure & Private (Zero Data Retention) Free Access • No Sign-Up
W3C WebGPU 1.0 WGSL Compute Culling Zero G-Buffer VRAM Bandwidth

WebGPU Clustered Forward Shading Studio

Simulate 3D camera frustum voxel clustering, exponential depth distribution, compute shader sphere-AABB light culling, and generate production-grade WGSL shaders for multi-thousand light WebGPU rendering.

1. Frustum & Cluster Grid Topography

2. Dynamic Light Population & Culling Limits

Global Active Lights: 1,500
Total Frustum Clusters
3,456
16 × 9 × 24 grid
Active Non-Empty Clusters
1,420
41.1% frustum fill
Mean Lights / Active Cluster
8.4
Peak: 34 / 64 max
Culling Elimination Rate
99.44%
vs Forward 1,500/frag
G-Buffer Bandwidth Saved
132 MB/f
@ 4K 60fps (7.9 GB/s)

3. Interactive Frustum Cluster Slicer & Light Heatmap

Low Overlap (1-10) Medium (11-30) Dense (31+)
Hover or click any cluster voxel to inspect view-space AABB and light containment list

Selected Cluster: (X: 8, Y: 4, Z: 6)

View-Space Min AABB:
[-2.41m, -1.35m, -14.28m]
View-Space Max AABB:
[+2.41m, +1.35m, -18.72m]
Assigned Light Count:
14 dynamic lights
Light Buffer Offset:
Cluster #1,208 → Index 77,312

4. Production WGSL Compute & Fragment Shaders


      

Frequently Asked Technical Questions

What is Clustered Forward Shading and how does it overcome traditional Forward and Deferred Shading?+
Traditional Forward Shading evaluates every light against every fragment (O(lights * fragments)), bottlenecking heavily with more than a few dozen lights. Deferred Shading solves this by rendering geometry into multiple high-bandwidth screen-space G-Buffers (diffuse, normal, specular, depth), but incurs severe memory bandwidth penalties (unfriendly to mobile GPUs and Apple Silicon tile-based architectures), breaks hardware Multisample Anti-Aliasing (MSAA), and cannot easily handle translucent or refractive surfaces. Clustered Forward Shading sub-divides the 3D camera frustum into a 3D grid of view-space voxels (clusters). A fast GPU compute shader culls thousands of dynamic lights against these clusters once per frame. The forward fragment pass then looks up only the small subset of lights (e.g. 5 to 30) that actually touch its cluster, achieving O(1) lighting scaling with zero G-Buffer bandwidth overhead and full hardware MSAA support.
Why is frustum depth sliced exponentially rather than linearly in cluster generation?+
Perspective projection is inherently non-linear: depth resolution is dense near the near-plane and spreads out drastically toward the far-plane. If frustum depth were partitioned linearly, the vast majority of scene geometry close to the camera would end up bunched into just the first one or two depth slices, defeating the purpose of culling. Exponential depth slicing partitions slice boundaries using z_i = z_near * (z_far / z_near)^(i / N_z). This produces cluster volumes that scale proportionally with perspective foreshortening, keeping view-space screen projected cluster sizes relatively uniform and preventing light count imbalances.
How does the WebGPU compute shader cull lights and record indices atomically?+
In the light culling compute pass (light_cull.wgsl), each workgroup thread evaluates a cluster against global light bounding spheres. If a light intersects the cluster AABB, the thread issues an atomic addition (atomicAdd(&cluster.count, 1u)) to reserve an index slot in a global light-index storage buffer. To prevent buffer overflows from dense light overlap, the addition is clamped to a safe maximum (e.g., 64 or 128 lights per cluster). The resulting buffer layout contains an array of Cluster headers { offset: u32, count: u32 } pointing into the contiguous global light index array.
Why is Clustered Forward Shading ideal for WebGPU on modern mobile and integrated GPUs?+
Modern mobile SoCs (Apple Silicon M-series/A-series, Qualcomm Snapdragon, MediaTek Dimensity) use Tile-Based Deferred Rendering (TBDR) architectures where memory bandwidth across the system bus is the single largest consumer of thermal headroom and battery power. Deferred shading requires constantly streaming 32 to 64 bytes of G-Buffer data per pixel out to system VRAM and reading it back. Clustered Forward Shading keeps all lighting calculations inside the fragment shader with on-chip registers, requiring only a tiny read-only light index buffer (a few hundred kilobytes). This dramatically boosts frame rates on WebGPU-enabled mobile and ultraportable devices.
Sponsored Utility
While You're Here
Sponsored Recommendations
Advertisement