Elasticsearch & OpenSearch Architecture Studio
Size production Elasticsearch and OpenSearch clusters: calculate primary/replica shard counts, prevent JVM heap oversharding, synthesize ILM/ISM policies, and audit Lucene page cache headroom.
Ingestion & Shard Sizing Parameters
Data Node Hardware Specification
Cluster runs at 16.1 shards per GB heap (well below the 20 shards/GB safety threshold). Lucene segment memory will not overwhelm JVM heap.
• Disk Watermark Threshold: 85% Low Watermark (6,800 GB max usable per 8TB node)
• Total Cluster Memory: 384 GB RAM (186 GB JVM / 198 GB PageCache)
• Split-Brain Resilience: 3 Master nodes guarantee continuous quorum if 1 master dies.
Index Lifecycle Policy Synthesizer
Automate the progression of time-series indices through Hot, Warm, Cold, and Delete phases to optimize NVMe performance and reduce storage costs by up to 70%.
Merging Lucene segments to 1 in the Warm phase purges soft-deleted documents and compresses inverted index term dictionaries, freeing 10% to 25% disk space.
The 50/50 Rule & 32GB CompressedOops Limit
Exceeding 31.5GB of JVM heap triggers 64-bit object pointer expansion, destroying CPU cache locality and wasting 6GB to 8GB of memory.
Heap is $le 31.5 ext{GB}$. JVM aligns pointers to 8-byte boundaries, packing 32GB of addressing space into compact 32-bit pointers. CPU cache efficiency is maximized.
Lucene avoids garbage-collected JVM heap for stored fields and postings lists. By letting the Linux kernel manage Lucene segments in RAM page cache, query execution avoids GC pauses and leverages zero-copy kernel transfers.
Multi-AZ Shard Allocation & Node Failure Simulator
Visualize shard distribution across 3 Availability Zones and simulate node failure self-healing mechanics.
5 Architectural Showdowns & Decision Matrices
Elasticsearch/OpenSearch: Full-text inverted index + BKD trees; best for high-cardinality search, phrase matching, and complex aggregations. ClickHouse: Pure columnar database; 5x faster analytical aggregations and 70% lower disk footprint, but poor for unstructured text. Grafana Loki: Only indexes metadata labels (like grep); 90% cheaper storage, but queries on unindexed log bodies require brute-force streaming scans.
Legacy Daily Indices (e.g. logs-2026.09.20): Creates fixed shards regardless of volume, producing 200MB shards on weekends and 120GB shards on Black Friday. Data Streams (.ds-*): Automatically rolls over indices based strictly on shard size (e.g. exactly 50GB per primary shard), eliminating oversharding and uneven disk hotspotting.
In small test clusters, nodes share both data and master roles. In production (>5 nodes), dedicated master nodes (3x small VMs) are mandatory. Heavy search aggregations on data nodes trigger high CPU/GC pauses; if that node is also a master, cluster heartbeats timeout, falsely declaring the master dead and triggering catastrophic re-election storms.
Doc Values are column-oriented data structures stored on disk and loaded into Linux OS page cache for sorting and aggregations on keywords/numbers. Fielddata is an in-memory un-inverted index loaded directly into the JVM heap for analyzed text fields. Enabling fielddata: true on high-cardinality text fields is the single most common cause of fatal cluster OOM crashes.
5 Fatal Production Elasticsearch Pitfalls
Creating hundreds of small shards (under 1GB each) wastes 20MB to 50MB of JVM heap per shard solely on Lucene segment readers and FST term pointers. 10,000 shards consume 30GB of heap with zero active queries, triggering incessant GC pauses. Remedy: Target 30GB–50GB shards for logs, shrink old indices, and enforce the 20 shards/GB heap rule.
Setting
-Xmx32g or -Xmx33g causes the JVM to switch from 32-bit compressed references to 64-bit uncompressed pointers. You lose ~8GB of memory to pointer overhead, making 33GB perform worse than 31GB. Remedy: Always set max heap strictly to -Xmx31g (or check JVM logs for "Compressed Oops enabled").
Ingesting arbitrary JSON objects with dynamic keys (e.g. user IDs or timestamps as JSON keys) expands index mappings into thousands of distinct field definitions. The cluster state explodes into hundreds of megabytes and synchronization fails across nodes. Remedy: Set
dynamic: false or dynamic: strict on nested telemetry objects and enforce index.mapping.total_fields.limit: 1000.
Configuring 2 master-eligible nodes with quorum $N/2 + 1 = 2$ means if a single node fails, the cluster drops below quorum and refuses all write requests. Remedy: Always deploy exactly 3 dedicated master nodes across 3 distinct failure zones.
*search)Executing a query like
{"wildcard": {"user_agent": "*Chrome*"}} bypasses the Lucene B-Tree / FST prefix index, forcing Lucene to iterate through every single term in the dictionary sequentially across all shards, causing 100% CPU spikes. Remedy: Use wildcard field types with n-gram indexing or reject unanchored wildcard queries via query DSL linters.