Everything, Everywhere
Verified Specification | Standardized Formulas | Instant Precision
Secure & Private (Zero Data Retention) Free Access • No Sign-Up

WebGPU Multi-Draw Indirect & GPU-Driven Rendering Studio

Architect high-throughput GPU-driven rendering pipelines in WebGPU. Model compute shader frustum culling, Hi-Z occlusion tests, atomic stream compaction, and zero-CPU drawIndexedIndirect execution.

WebGPU 1.0 GPU-Driven Pipeline drawIndexedIndirect Atomic Compaction
Rendering pipeline command dispatch model
Number of 3D object instances in scene space
Horizontal camera view frustum angle
Azimuth rotation of observer frustum

Top-Down Spatial World & Frustum Culling Simulation

Visible (Drawn) Frustum Culled Occlusion Culled Frustum Boundary
Total Instances
2,048
Scene Objects
Visible / Drawn
612
29.9% Drawn
Total Culled
1,436
Frustum: 1,436
CPU Draw Calls
1
drawIndexedIndirect
CPU Command Time
0.04 ms
-99.3% CPU Load
GPU Compute Time
0.02 ms
@workgroup_size(64)

GPUBuffer Memory Layout: DrawIndexedIndirectArgs Struct (20 Bytes)

The WebGPU specification requires five consecutive 32-bit unsigned/signed integers inside the GPUBufferUsage.INDIRECT buffer. The compute shader dynamically updates instanceCount via atomic compaction.

Byte Offset Field Identifier WGSL Type Runtime Value Hex Memory Dump Architectural Role
+0x00 indexCount u32 1,440 0x000005A0 Indices per instance (e.g. 480 triangles)
+0x04 instanceCount atomic<u32> 612 0x00000264 Dynamic: Incremented by GPU compute culling pass
+0x08 firstIndex u32 0 0x00000000 Offset in index buffer to begin reading vertices
+0x0C baseVertex i32 0 0x00000000 Constant added to vertex indices before vertex buffer fetch
+0x10 firstInstance u32 0 0x00000000 Starting instance ID offset passed to vertex shader

Production WGSL & TypeScript Pipeline Architecture


      

GPU-Driven Rendering Architecture & Execution Paradigm

1. Frustum Bounding Sphere Math Each camera frustum is defined by 6 half-space plane equations \(\mathbf{n} \cdot \mathbf{x} + d = 0\). An object bounding sphere with world-space center \(\mathbf{c}\) and radius \(r\) is tested against each plane: \(D = \mathbf{n}_i \cdot \mathbf{c} + d_i\). If \(D < -r\) for any plane \(i\), the sphere is completely behind the plane and culled in parallel.
2. Atomic Stream Compaction To prevent fragmented rendering, passing instances are packed into a continuous index array. A compute workgroup invokes atomicAdd(&indirectArgs.instanceCount, 1u). The returned atomic index grants exclusive write ownership to visibleInstanceIDs[slot] = instanceId.
3. Zero-CPU Synchronization The indirect buffer remains entirely in GPU VRAM with GPUBufferUsage.INDIRECT | GPUBufferUsage.STORAGE. The CPU never reads back visibility data, eliminating costly GPU-to-CPU synchronization pipeline bubbles (0 pipeline stalls).

Frequently Asked Technical Questions

What is GPU-driven rendering and how does drawIndexedIndirect eliminate the CPU bottleneck?+
In traditional CPU-driven rendering, the CPU must iterate over thousands of objects every frame, calculate camera frustum visibility, bind uniform buffers, and record individual draw commands (e.g. pass.drawIndexed()). For scenes with tens of thousands of meshes, CPU driver overhead, validation checks, and command buffer recording saturate the CPU core, dropping framerates even if the GPU has idle render capacity. In GPU-driven rendering, visibility determination (frustum culling, occlusion culling, and LOD selection) is offloaded entirely to a WebGPU Compute Shader. The compute shader dynamically writes the draw arguments directly into a GPUBuffer with GPUBufferUsage.INDIRECT. The CPU issues only a single pass.drawIndexedIndirect(indirectBuffer, 0) call, dropping CPU command recording time from O(N) to O(1).
What is the exact binary memory layout of DrawIndexedIndirectArgs in WebGPU?+
The WebGPU specification mandates a strict 20-byte struct layout (five 32-bit words) for DrawIndexedIndirectArgs stored in a GPUBuffer with GPUBufferUsage.INDIRECT: 1) indexCount: u32 (byte offset 0, number of indices to read from index buffer); 2) instanceCount: u32 (byte offset 4, number of visible instances to draw, dynamically populated by compute shader atomic compaction); 3) firstIndex: u32 (byte offset 8, starting offset in the index buffer); 4) baseVertex: i32 (byte offset 12, value added to each index before reading vertex attributes); 5) firstInstance: u32 (byte offset 16, starting instance ID).
How does the Compute Shader perform parallel frustum culling with atomic compaction?+
Each compute invocation evaluates one object instance identified by @builtin(global_invocation_id).x. The shader extracts the object bounding sphere center and radius from a read-only storage buffer, transforms the center into world space, and tests it against the 6 frustum planes (Left, Right, Bottom, Top, Near, Far). For each plane with normal n and distance d, if dot(n, center) + d < -radius, the object lies completely outside the frustum and is culled. If the object passes all 6 planes, the compute shader executes atomicAdd(&indirectArgs.instanceCount, 1u) to allocate a unique slot in a visibleInstancesBuffer, storing the object index. The vertex shader then retrieves instance transform matrices using instance_index into this compacted list.
How does WebGPU handle Multi-Draw Indirect compared to Vulkan vkCmdDrawIndexedIndirectCount?+
Vulkan and DirectX 12 support native multi-draw indirect with dynamic draw count (vkCmdDrawIndexedIndirectCount), allowing a compute shader to emit an arbitrary number of distinct draw calls without the CPU knowing the count. In WebGPU 1.0, renderPass.drawIndexedIndirect() executes a single indirect draw call per invocation (though that single call can render arbitrary instances via instanceCount). For multi-mesh architectures with disparate geometries, WebGPU applications either pack meshes into a giant unified index/vertex buffer (megabuffer) and draw via compacted instance lists, or issue a fixed loop of drawIndexedIndirect() calls reading from consecutive 20-byte offsets within the indirect argument buffer.
Sponsored Utility
While You're Here
Sponsored Recommendations
Advertisement