Apache Iceberg & Parquet Columnar Lakehouse Storage Studio
An architectural deep-dive, interactive metadata snapshot tree simulator, and query pushdown modeler for modern data lakehouses. Trace transactions across Table Metadata, Manifest Lists, and Manifest Files, simulate 99%+ I/O reduction through Parquet columnar projections and dictionary pruning, compare Copy-on-Write vs Merge-on-Read, and synthesize production PyIceberg, DuckDB, and Rust pipelines.
Apache Iceberg Metadata Tree & Atomic Commit Simulator
Step through table lifecycle operations (Appends, Overwrites, Deletes, Compaction). Observe how the 4-tier metadata tree guarantees ACID isolation without filesystem renames or directory locks.
Parquet Columnar Projection & Predicate Pushdown Modeler
Calculate physical bytes scanned across storage formats. Contrast row-oriented CSV/JSON scans against Parquet column projection, row group stats min/max pruning, and dictionary filtering.
Apache Parquet Binary File Anatomy
| Component | Typical Size | Contents & Metadata | Pruning / Optimization Role |
|---|---|---|---|
| Header | 4 Bytes | Magic bytes PAR1 |
Format verification |
| Row Groups (1..N) | 128 MB – 512 MB | Horizontal partition of table rows | Enables parallel reading across worker cores |
| Column Chunks | 1 MB – 64 MB | All data for 1 column in row group | Column Projection: unrequested chunks skipped entirely |
| Data Pages | 1 MB | Values compressed with Snappy / ZSTD | RLE bit-packing and dictionary encoded |
| Dictionary Pages | 10 KB – 500 KB | Set of distinct string/symbol values | Dictionary filtering skips entire page if value absent |
| File Metadata Footer | 10 KB – 100 KB | Schema, Row Group offsets, Min/Max stats | Statistics Pruning: Read once at end of file; prunes row groups |
Copy-on-Write (CoW) vs Merge-on-Read (MoR) Architecture
How Iceberg handles record updates, GDPR row deletions, and change-data-capture (CDC) streams:
| Evaluation Dimension | Copy-on-Write (CoW) | Merge-on-Read (MoR) |
|---|---|---|
| Update / Delete Mechanism | Rewrites entire Parquet data file containing affected rows. | Appends a lightweight Positional Delete or Equality Delete file. |
| Write Amplification | High (1 row update rewrites 500MB) | Ultra-Low (Writes few KB delete file) |
| Read Query Latency | Fastest (Direct Parquet scan) | Moderate (Requires in-memory anti-join) |
| Compaction Requirement | Minimal (data files remain clean) | Mandatory (periodic compaction merges deletes into base files) |
| Ideal Production Workload | Batch data pipelines, BI dashboards, read-intensive reporting. | Real-time streaming CDC ingest (Debezium), GDPR right-to-be-forgotten. |
Open Lakehouse Formats: Apache Iceberg vs Delta Lake vs Apache Hudi
A rigorous engineering comparison of the big three open-source table formats:
| Feature / Dimension | Apache Iceberg | Delta Lake | Apache Hudi |
|---|---|---|---|
| Metadata Architecture | Hierarchical Avro tree (Metadata -> Manifest List -> Manifests) | JSON transaction commit log with periodic Parquet checkpoints | Timeline metadata (Avro commits, instants, and delta logs) |
| Partitioning Model | Hidden Partitioning + Partition Evolution | Physical folder paths (/key=val/) + Liquid Clustering | Physical directory paths |
| Catalog Decoupling | Universal (REST, Polaris, Unity, Nessie, Glue) | Historically tied to Databricks / Unity Catalog | Tied to Hive Metastore / AWS Glue |
| Engine Independence | Trino, DuckDB, Snowflake, Spark, Flink, StarRocks | Native Spark; Delta Kernel for multi-engine | Native Spark and Flink |
| Schema Evolution | Full (Rename, reorder, add, drop with column IDs) | Full (Column mapping mode enabled) | Supported with restrictions |
Production Implementation Blueprints
Syntax-validated, memory-efficient implementations for Iceberg lakehouse pipelines and Parquet queries.
// Select a blueprint above