Context windows · October 2026
A context window is the maximum number of tokens an AI model can work with in a single request — your prompt, any documents or conversation history, and the model's reply combined. As of October 2026, the largest belong to GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna at 1,050,000 tokens, roughly 787,500 words.
It is measured in tokens, not words. Below: every current model's window, what fits inside it, and what changes when you fill it.
Converted with 1 token ≈ 0.75 English words, 500 words per page, and 90,000 words per novel.
| Model | Context window | Words | Pages | Fits |
|---|---|---|---|---|
| GPT-6 Astra | 1,050,000 | 787,500 | 1,574 | all of the King James Bible |
| GPT-6 Sol | 1,050,000 | 787,500 | 1,574 | all of the King James Bible |
| GPT-6 Luna | 1,050,000 | 787,500 | 1,574 | all of the King James Bible |
| Claude Fable 5.1 | 1,000,000 | 750,000 | 1,499 | 1.3× War and Peace |
| Claude Opus 5 | 1,000,000 | 750,000 | 1,499 | 1.3× War and Peace |
| Claude Opus 5.5 | 1,000,000 | 750,000 | 1,499 | 1.3× War and Peace |
| Gemini 3.1 Pro | 1,000,000 | 750,000 | 1,499 | 1.3× War and Peace |
| Claude Sonnet 5 | 1,000,000 | 750,000 | 1,499 | 1.3× War and Peace |
| Gemini 3.8 Flash | 1,000,000 | 750,000 | 1,499 | 1.3× War and Peace |
| Gemini 3.5 Flash-Lite | 1,000,000 | 750,000 | 1,499 | 1.3× War and Peace |
| Claude Haiku 4.5 | 200,000 | 150,000 | 300 | 1.7× a typical 90,000-word novel |
1,050,000 tokens
GPT-6 Astra, GPT-6 Sol, GPT-6 Luna
1,000,000 tokens
Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5, Gemini 3.1 Pro, Claude Sonnet 5, Gemini 3.8 Flash, Gemini 3.5 Flash-Lite
200,000 tokens
Claude Haiku 4.5
Every request to a model is assembled into one sequence of tokens: the system prompt, tool definitions, earlier turns of the conversation, any documents or search results you attach, your new message — and then the reply the model writes. All of it has to fit inside the context window.
When a request would exceed the window, it can't be processed as-is. Applications handle this by trimming or summarizing older conversation turns, retrieving only the relevant parts of long documents, or splitting the work across several requests. That is why long chats sometimes "forget" early details.
The window is shared by input and output, but the output also has its own, much smaller limit. Claude Opus 5.5 and Claude Fable 5.1, for example, accept up to 1M tokens in total but write at most 128K tokens in one response. A big window lets a model read a lot; it doesn't mean it will write that much.
You pay per token, so filling a large window is expensive: a 1M-token prompt on a model charging $2 per million input tokens costs $2 every time you send it. Some providers also add a long-context surcharge —Gemini 3.1 Pro charges more for prompts above 200K tokens — while Anthropic bills the full 1M window of Claude 4.6 and later models at the standard rate. Prompt caching cuts the cost of re-sending the same long context; see the Claude, OpenAI andGemini pricing pages for cached-input prices.
To see how much of a window your own text would use, paste it into the token counter.
A context window is the maximum number of tokens an AI model can work with in a single request — your prompt, any documents or conversation history, and the model's reply combined. As of October 2026, the largest belong to GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna at 1,050,000 tokens, roughly 787,500 words.
As of October 2026, GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna have the largest context windows at 1,050,000 tokens. That is about 787,500 words, or all of the King James Bible.
Yes. The context window covers input and output together, so a long reply leaves less room for input. Models also have a separate cap on output length that is far smaller than the window — Claude Opus 5.5 and Claude Fable 5.1, for example, have a 1M-token window but can write at most 128K tokens in one response.
You pay for every token you send, so a longer prompt always costs more in total. Some models also raise the per-token price for very long prompts: Gemini 3.1 Pro charges more above 200K tokens, while Anthropic bills the full 1M-token window of Claude 4.6 and later models at the standard rate.
No. A model only sees what is inside the context window of the current request. Chat apps that seem to remember earlier conversations do it by re-sending past messages, summaries or retrieved notes as part of the context each time.