PDF Page Counter & Metadata Inspector
Quickly check page count, PDF version, author, and security properties instantly without installing software.
SELECT PDF FILE TO INSPECT
Drop file or click to choose from your device
PDF Binary Format & Object Tree Derivations
Every standard PDF document (ISO 32000-1) is structured into four sequential binary sections: the Header (%PDF-1.x), the Body of Indirect Objects (Fonts, Pages, Images), the Cross-Reference Table (xref byte offsets), and the Trailer Dictionary referencing the /Root Document Catalog.
10-digit byte offset: e.g. 0000049210
5-digit generation number: 00000
f = Free object • n = In-use object
Node: /Kids [ref1, ref2, ...] • /Count N
Leaf: /Type /Page • /Parent ref • /MediaBox [0 0 w h]
Total Pages = Sum of all Leaf nodes
| PDF Section | Marker / Syntax | Metadata Contained | Inspection Priority |
|---|---|---|---|
| Header | %PDF-1.4 to %PDF-2.0 | Format version & binary safety flag | Engine capability validation |
| Body Objects | id gen obj ... endobj | Page trees, raster images, embedded fonts | Resource profiling |
| Info Dictionary | /Title /Author /Producer | Author, software generator, timestamp | Document provenance |
| Trailer | trailer << /Size /Root >> %%EOF | Byte offset to xref table & encryption dictionary | Primary entry point |
5 Critical PDF Page Counting & Metadata Traps
1. Corrupted Page Tree Counts (/Count vs Actual /Kids Array)
In malformed or patched PDFs, the /Pages << /Count N >> value may report 50 pages, but the actual /Kids tree contains only 35 valid leaf nodes. Simple regex scrapers that extract the /Count integer give inaccurate results. Our tool recursively parses the actual object tree to count guaranteed renderable pages.
2. Incremental Updates Hiding True Document History
When digitally signing or editing a PDF, many applications append new objects to the end of the file along with a new xref and trailer rather than rewriting the file. A document may contain multiple conflicting /Info dictionaries; inspecting only the top of the file misses recent revisions.
3. Discrepancies Between /Info and XMP Metadata Packets
PDF documents maintain two independent metadata systems: legacy document info dictionaries (/Title, /Author) and modern XML-based Adobe XMP metadata streams. Often, sanitization tools wipe the /Info dictionary but leave sensitive author names and GPS locations intact within the embedded XMP stream.
4. Truncated Trailing Bytes and Broken %%EOF Markers
Incomplete downloads often clip the final 1,024 bytes containing the %%EOF tag and startxref pointer. Without these, standard operating system PDF viewers declare the file corrupt and fail to open it. PDF.js includes error recovery routines that rebuild broken cross-reference tables from raw body scans.
5. Hidden Annotations and Redacted Text Leaks
Placing a black rectangle over sensitive text in a PDF does not redact it. Unless the underlying character stream was permanently removed, the text still exists in the page object's content stream and will be reported by metadata and text inspectors.