How Many PDF Pages Is 1 Million Tokens?
One million tokens can process about 650 visually-rendered PDF pages (preserving layout, tables, and figures at ~1,500 tokens per page). That means the IPCC AR6 Synthesis Report (81-page Longer Report, often cited as 85 pages) fits roughly 8 times within a single context window.
Why this matters
PDFs combine text, tables, and figures, making them one of the most token-intensive formats. A single visually-rendered PDF page costs roughly 1,500 tokens β so the 81-page IPCC AR6 Synthesis Report uses about 121,500 tokens, while a 200-page SEC filing uses around 300,000. Understanding the per-page token cost helps enterprises estimate the expense of processing contracts, research papers, and financial reports through AI.

1 block = 1 PDF page
The answer
β 650 PDF pages
visually-processed PDF pages
The math
Basis assumption
~1,500 tokens per page with layout
Calculation
When a model reads a PDF page as both text and image (preserving layout, tables and figures), each page costs ~1,500 tokens. Text-only extraction is cheaper: ~2,000 pages per million tokens. For example, the IPCC AR6 Synthesis Report (81-page Longer Report, often cited as 85 pages) is about 121,500 tokens, so a million-token model could read it roughly 8 times over.
Result
1,000,000 tokens = 650 visually-processed PDF pages
Real PDFs, measured in tokens
Token counts for well-known documents, read with full layout (~1,500 tokens/page) or as extracted text (~500 tokens/page). Cost is to read the document once as input.
| Document | Pages | Tokens (layout) | Tokens (text only) | Fits in 1M | GPT-6 Luna | Claude Fable 5.1 |
|---|---|---|---|---|---|---|
| Bitcoin whitepaper | 9 | 13,500 | 4,500 | 74.1Γ | <$0.01 | $0.14 |
| "Attention Is All You Need" | 15 | 22,500 | 7,500 | 44.4Γ | <$0.01 | $0.22 |
| IPCC AR6 Synthesis Report (Longer Report) | 81 | 121,500 | 40,500 | 8.2Γ | $0.01 | $1.21 |
| GDPR (Official Journal) | 88 | 132,000 | 44,000 | 7.6Γ | $0.01 | $1.32 |
| GPT-4 Technical Report | 100 | 150,000 | 50,000 | 6.7Γ | $0.01 | $1.50 |
| Typical SEC 10-K filing | 200 | 300,000 | 100,000 | 3.3Γ | $0.03 | $3.00 |
| Mueller Report | 448 | 672,000 | 224,000 | 1.5Γ | $0.07 | $6.72 |
Page counts are for the published PDF editions; token counts are estimates and vary by tokenizer and page density.
Real-world references
To put this in perspective, here are some well-known examples for comparison.
Frequently asked questions
- How many tokens is one PDF page?
- About 1,500 tokens when the model reads the page with its layout, tables and figures, and about 500 tokens when only the extracted text is sent. A dense page of small print can run higher.
- How many tokens is the IPCC AR6 Synthesis Report?
- The IPCC AR6 Synthesis Report Longer Report is 81 pages (often cited as 85): roughly 121,500 tokens read with full layout, or about 40,500 tokens as extracted text. A 1 million token context window holds it about 8.2 times.
- Is it cheaper to extract the text from a PDF first?
- Yes, usually about 3x cheaper. Text-only extraction drops the page images, so a page costs ~500 tokens instead of ~1,500. The trade-off is that tables, charts and scanned pages can lose information.
- Can an AI model read a 400-page PDF in one request?
- Yes, if its context window is large enough. A 448-page document like the Mueller Report is about 672,000 tokens with layout, which fits in a 1 million token context window but not in a 200,000 token one.