Featured Developer Sponsor • Zero-Token Protection
WebGPU Multi-Draw Indirect & GPU-Driven Rendering Studio
Architect high-throughput GPU-driven rendering pipelines in WebGPU. Model compute shader
frustum culling, Hi-Z occlusion tests, atomic stream compaction, and zero-CPU
drawIndexedIndirect execution.
WebGPU 1.0
GPU-Driven Pipeline
drawIndexedIndirect
Atomic Compaction
Rendering pipeline command dispatch model
Number of 3D object instances in scene space
Horizontal camera view frustum angle
Azimuth rotation of observer frustum
Top-Down Spatial World & Frustum Culling Simulation
Visible (Drawn)
Frustum Culled
Occlusion Culled
Frustum Boundary
Total Instances
2,048
Scene Objects
Visible / Drawn
612
29.9% Drawn
Total Culled
1,436
Frustum: 1,436
CPU Draw Calls
1
drawIndexedIndirect
CPU Command Time
0.04 ms
-99.3% CPU Load
GPU Compute Time
0.02 ms
@workgroup_size(64)
GPUBuffer Memory Layout: DrawIndexedIndirectArgs Struct (20 Bytes)
The WebGPU specification requires five consecutive 32-bit unsigned/signed integers inside the
GPUBufferUsage.INDIRECT buffer. The compute shader dynamically updates
instanceCount via atomic compaction.
| Byte Offset | Field Identifier | WGSL Type | Runtime Value | Hex Memory Dump | Architectural Role |
|---|---|---|---|---|---|
| +0x00 | indexCount | u32 | 1,440 | 0x000005A0 | Indices per instance (e.g. 480 triangles) |
| +0x04 | instanceCount | atomic<u32> | 612 | 0x00000264 | Dynamic: Incremented by GPU compute culling pass |
| +0x08 | firstIndex | u32 | 0 | 0x00000000 | Offset in index buffer to begin reading vertices |
| +0x0C | baseVertex | i32 | 0 | 0x00000000 | Constant added to vertex indices before vertex buffer fetch |
| +0x10 | firstInstance | u32 | 0 | 0x00000000 | Starting instance ID offset passed to vertex shader |
Production WGSL & TypeScript Pipeline Architecture
GPU-Driven Rendering Architecture & Execution Paradigm
1. Frustum Bounding Sphere Math
Each camera frustum is defined by 6 half-space plane equations \(\mathbf{n} \cdot \mathbf{x} + d = 0\). An object bounding sphere with world-space center \(\mathbf{c}\) and radius \(r\) is tested against each plane:
\(D = \mathbf{n}_i \cdot \mathbf{c} + d_i\). If \(D < -r\) for any plane \(i\), the sphere is completely behind the plane and culled in parallel.
2. Atomic Stream Compaction
To prevent fragmented rendering, passing instances are packed into a continuous index array. A compute workgroup invokes
atomicAdd(&indirectArgs.instanceCount, 1u). The returned atomic index grants exclusive write ownership to visibleInstanceIDs[slot] = instanceId.
3. Zero-CPU Synchronization
The indirect buffer remains entirely in GPU VRAM with
GPUBufferUsage.INDIRECT | GPUBufferUsage.STORAGE. The CPU never reads back visibility data, eliminating costly GPU-to-CPU synchronization pipeline bubbles (0 pipeline stalls).
Frequently Asked Technical Questions
What is GPU-driven rendering and how does drawIndexedIndirect eliminate the CPU bottleneck?+
In traditional CPU-driven rendering, the CPU must iterate over thousands of objects every frame, calculate camera frustum visibility, bind uniform buffers, and record individual draw commands (e.g. pass.drawIndexed()). For scenes with tens of thousands of meshes, CPU driver overhead, validation checks, and command buffer recording saturate the CPU core, dropping framerates even if the GPU has idle render capacity. In GPU-driven rendering, visibility determination (frustum culling, occlusion culling, and LOD selection) is offloaded entirely to a WebGPU Compute Shader. The compute shader dynamically writes the draw arguments directly into a GPUBuffer with GPUBufferUsage.INDIRECT. The CPU issues only a single pass.drawIndexedIndirect(indirectBuffer, 0) call, dropping CPU command recording time from O(N) to O(1).
What is the exact binary memory layout of DrawIndexedIndirectArgs in WebGPU?+
The WebGPU specification mandates a strict 20-byte struct layout (five 32-bit words) for DrawIndexedIndirectArgs stored in a GPUBuffer with GPUBufferUsage.INDIRECT: 1) indexCount: u32 (byte offset 0, number of indices to read from index buffer); 2) instanceCount: u32 (byte offset 4, number of visible instances to draw, dynamically populated by compute shader atomic compaction); 3) firstIndex: u32 (byte offset 8, starting offset in the index buffer); 4) baseVertex: i32 (byte offset 12, value added to each index before reading vertex attributes); 5) firstInstance: u32 (byte offset 16, starting instance ID).
How does the Compute Shader perform parallel frustum culling with atomic compaction?+
Each compute invocation evaluates one object instance identified by @builtin(global_invocation_id).x. The shader extracts the object bounding sphere center and radius from a read-only storage buffer, transforms the center into world space, and tests it against the 6 frustum planes (Left, Right, Bottom, Top, Near, Far). For each plane with normal n and distance d, if dot(n, center) + d < -radius, the object lies completely outside the frustum and is culled. If the object passes all 6 planes, the compute shader executes atomicAdd(&indirectArgs.instanceCount, 1u) to allocate a unique slot in a visibleInstancesBuffer, storing the object index. The vertex shader then retrieves instance transform matrices using instance_index into this compacted list.
How does WebGPU handle Multi-Draw Indirect compared to Vulkan vkCmdDrawIndexedIndirectCount?+
Vulkan and DirectX 12 support native multi-draw indirect with dynamic draw count (vkCmdDrawIndexedIndirectCount), allowing a compute shader to emit an arbitrary number of distinct draw calls without the CPU knowing the count. In WebGPU 1.0, renderPass.drawIndexedIndirect() executes a single indirect draw call per invocation (though that single call can render arbitrary instances via instanceCount). For multi-mesh architectures with disparate geometries, WebGPU applications either pack meshes into a giant unified index/vertex buffer (megabuffer) and draw via compacted instance lists, or issue a fixed loop of drawIndexedIndirect() calls reading from consecutive 20-byte offsets within the indirect argument buffer.
Sponsored Utility
While You're Here
Sponsored Recommendations
Advertisement