Featured Developer Sponsor • Zero-Token Protection
RFC 4180 Compliant
Auto-Delimiter Detection
Zero Server Uploads
CSV to JSON Converter & Delimited Table Parser
Convert CSV, TSV, and delimited spreadsheet data into structured JSON arrays, matrices, or keyed objects. Features RFC 4180 state machine tokenization, automatic delimiter detection, intelligent type casting, interactive preview table, and 100% client-side execution.
Need to convert in the opposite direction? Transform JSON arrays back into RFC 4180 CSV spreadsheets with our JSON to CSV Converter.
Tip: Drag and drop any .csv, .tsv, or .txt file directly into the editor
Rows: 0
Cols: 0
Payload: 0 B
Delimiter: Auto
Ready
📊 Live Parsed Data Table
Visual table preview of parsed records📐 RFC 4180 State Machine Architecture & Type Casting
Parsing delimited tabular data reliably requires a deterministic finite state machine rather than simple String.split(',') calls. Here is how our browser engine processes your dataset:
1. UTF-8 Byte Order Mark Stripping: Files exported from Excel on Windows begin with the invisible UTF-8 BOM sequence (
0xEF, 0xBB, 0xBF or Unicode \uFEFF). If not stripped, the first column header becomes "\uFEFFid" instead of "id", causing silent lookup failures in code.
2. Quote State Tracking (RFC 4180 §2.5): When encountering an opening quotation mark, the tokenizer enters Quoted State. Inside this state, commas, tabs, semicolons, and CRLF line breaks are treated as literal cell content. Escaped quotes (
"") are unescaped into single quotes.
3. Auto-Delimiter Histogram Scoring: The engine samples up to 4,000 characters and computes frequency distributions for commas, semicolons, tabs, and pipes outside of quote boundaries to determine the true delimiter without user intervention.
4. Strict Primitive Coercion: Values matching
/^-?\d+(\.\d+)?([eE][+-]?\d+)?$/ without arbitrary leading zeros (except decimal fractions) are cast to IEEE 754 numbers; boolean strings (true/false) become native booleans; empty or null tokens resolve to JavaScript null.
⚠️ 5 Fatal Traps in CSV to JSON Conversion & ETL Pipelines
1. The Unescaped Double Quote & Commas Inside Quotes Trap
Naively splitting lines on commas with
line.split(',') catastrophically shatters any record with an address ("123 Main St, Apt 4") or product title into split columns. Subsequent fields shift to the wrong headers, corrupting relational schemas and database insertions. Always use an RFC 4180 compliant state-machine tokenizer.
2. The UTF-8 BOM Phantom Header Bug
When Excel exports UTF-8 CSVs, it injects a 3-byte Byte Order Mark (
\uFEFF) at index 0. If your script accesses record.id or record["id"], it returns undefined because the actual key in memory is record["\uFEFFid"]. This is one of the most frustrating hidden bugs in backend Node.js and Python ETL scripts.
3. European Semicolon (;) & Decimal Comma Confusion
In European locales, commas denote decimals (e.g.
€1.499,50) and semicolons separate columns. Parsing a German CSV with a comma-delimiter splits every financial number into two separate broken cells. Our auto-detection ensures European semicolon formats are parsed smoothly.
4. Leading Zero Stripping on ZIP Codes & Account Numbers
Casting all numeric tokens to numbers destroys data integrity for US postal ZIP codes (e.g. New Jersey
07030 becomes integer 7030), phone numbers (+011...), and international IBAN/routing numbers. If your CSV contains leading-zero identifiers, uncheck "Parse Numbers" to preserve strings intact.
5. Duplicate Column Header Overwrite in JSON Objects
If a CSV contains two columns with the same name (for example, two tables joined in SQL yielding two
created_at columns), converting to an Array of Objects causes the second value to silently overwrite the first. Our converter automatically appends numerical suffixes (created_at_2) or allows conversion to a 2D Array matrix where duplicates are preserved.
Sponsored Utility
While You're Here
Sponsored Recommendations
Advertisement