Distributed Consensus: Multi-Paxos & ZooKeeper Zab Architecture Studio
Architect, simulate, and verify state machine replication (SMR) protocols. Model Multi-Paxos Phase 1/Phase 2 message passing with slot hole repair, simulate ZooKeeper Atomic Broadcast (Zab) epochs, 64-bit zxid sequences, and DIFF/TRUNC recovery pipelines, test network partition split-brain quorums, and synthesize production ZooKeeper and Go consensus blueprints.
Interactive 5-Node State Machine Replication Simulator
Inject partitions, crash leaders, and propose new client commands to inspect how Multi-Paxos and Zab maintain safety and quorum intersection.
REAL-TIME CONSENSUS MESSAGE BUS & REPLICATION TRACE
Multi-Paxos Phase 1 (Prepare) & Phase 2 (Accept) Deep Dive
Examine the exact message formats, ballot number monotonicity rules, and slot hole repair algorithms.
Phase 1: Leader Establishment (1 RTT Amortized)
Proposer generates a unique, strictly increasing ballot number
(round, node_id) and broadcasts to all acceptors.1b. Promise(ballot_num, max_accepted_slots)
If
ballot_num > min_ballot, acceptor promises never to accept future proposals with ballots smaller than ballot_num, returning all previously accepted (slot_id, vr, val) entries.Once promised by a majority quorum (3 of 5), the proposer becomes the stable leader and skips Phase 1 for all subsequent slots!
Phase 2: Log Replication (1 RTT Normal Case)
Leader assigns command to
slot_id and broadcasts Accept. If an existing value was promised in Phase 1 for this slot, the leader MUST propose that value instead to preserve safety!2b. Accepted(slot_id, ballot_num)
Acceptor accepts proposal if
ballot_num >= min_ballot and acknowledges to leader.2c. Commit / Learn(slot_id, command)
Once accepted by a quorum, the leader marks
slot_id committed and notifies state machines and learners.
Multi-Paxos Log Slot Hole Repair Mechanism
When a new leader assumes command, it observes gaps where earlier leaders failed mid-replication. State machines cannot execute out-of-order slots; the leader injects no-op commands to heal holes.
Slot 2: [SET y=20] → Committed (Ballot 1)
Slot 3: [HOLE - UNCOMMITTED] → New Leader detects no quorum accepted a value → Proposes [NO-OP] with Ballot 2 → Committed
Slot 4: [SET z=30] → Quorum accepted → Reproposed by Leader with Ballot 2 → Committed
Result: State machine applies Slot 1, Slot 2, Slot 3 (no-op skipped), and Slot 4 sequentially with zero state corruption.
ZooKeeper Atomic Broadcast (Zab) Protocol Deep Dive
Explore the 64-bit zxid anatomy, Fast Leader Election (FLE), and synchronization recovery states (DIFF, TRUNC, SNAP).
Anatomy of a 64-Bit Zab Transaction Identifier (zxid)
0x000000050000014A. When a new leader takes over, it increments the Epoch by 1 and resets the Counter to 0 (0x0000000600000000).
Phase 1: Discovery
Followers connect to prospective leader. Leader establishes newest epoch (CEPOCH & NEWEPOCH) by exchanging prospective epochs with a majority quorum.
Phase 2: Synchronization
Leader aligns follower logs before serving clients: sends DIFF to catch up followers, TRUNC to discard uncommitted orphan logs, or SNAP for full tree transfer.
Phase 3: Broadcast
Normal linear atomic broadcast: Leader sends PROPOSAL(zxid, v) → Followers reply with ACK(zxid) → Leader sends COMMIT(zxid) upon majority quorum.
Consensus Protocols Architecture Matrix: Multi-Paxos vs Zab vs Raft
Comparative analysis of core consensus protocols across leader election, log structures, hole handling, and client read guarantees.
| Protocol Dimension | Multi-Paxos | ZooKeeper Zab | Raft (Ongaro & Ousterhout) |
|---|---|---|---|
| Primary Design Goal | General state machine consensus | Hierarchical tree primary-backup | Understandability & formal verification |
| Leader Election Rule | Any node can propose; ballot collision | Fast Leader Election (highest zxid/epoch) | Candidate must have up-to-date log |
| Log Continuity | Allows out-of-order slots & holes | Strict contiguous FIFO stream | Strict contiguous log (Log Matching Invariant) |
| Hole Filling on Recovery | Explicit no-op proposals for gaps | Truncation (TRUNC) of uncommitted proposals | Overwrites follower entries from leader log |
| Steady State Latency | 1 RTT (Phase 2 Accept/Accepted) | 1 RTT (Proposal/Ack) | 1 RTT (AppendEntries/Ack) |
| Linearizable Reads | Quorum Leases or ReadIndex | sync() API (Phase 2 broadcast) | ReadIndex or LeaseRead |
| Notable Real-World Systems | Google Spanner, Chubby, CockroachDB | Apache ZooKeeper, ClickHouse Keeper | etcd (Kubernetes), HashiCorp Consul/Nomad |
Production Configuration & Cluster Sizing Blueprints
Production-hardened ZooKeeper zoo.cfg configs, Kubernetes 5-node StatefulSets, and Go Multi-Paxos state machine scaffolding.
Distributed Consensus Production Traps & Split-Brain Pitfalls
System.currentTimeMillis()) for leader leases means an NTP leap second or VM hypervisor pause can cause the old leader to believe its lease is still valid while the remaining nodes already timed out and elected a new leader. Two leaders simultaneously accept writes or serve conflicting reads! Solution: Use monotonic clock timers (CLOCK_MONOTONIC) with bounded drift guard intervals.
dataLogDir) apart from snapshot storage (dataDir), and enable group commit batching.