Context windows · October 2026

What is a context window?

A context window is the maximum number of tokens an AI model can work with in a single request — your prompt, any documents or conversation history, and the model's reply combined. As of October 2026, the largest belong to GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna at 1,050,000 tokens, roughly 787,500 words.

It is measured in tokens, not words. Below: every current model's window, what fits inside it, and what changes when you fill it.

Context window sizes for every current model

Converted with 1 token ≈ 0.75 English words, 500 words per page, and 90,000 words per novel.

ModelContext windowWordsPagesFits
GPT-6 Astra1,050,000787,5001,574all of the King James Bible
GPT-6 Sol1,050,000787,5001,574all of the King James Bible
GPT-6 Luna1,050,000787,5001,574all of the King James Bible
Claude Fable 5.11,000,000750,0001,4991.3× War and Peace
Claude Opus 51,000,000750,0001,4991.3× War and Peace
Claude Opus 5.51,000,000750,0001,4991.3× War and Peace
Gemini 3.1 Pro1,000,000750,0001,4991.3× War and Peace
Claude Sonnet 51,000,000750,0001,4991.3× War and Peace
Gemini 3.8 Flash1,000,000750,0001,4991.3× War and Peace
Gemini 3.5 Flash-Lite1,000,000750,0001,4991.3× War and Peace
Claude Haiku 4.5200,000150,0003001.7× a typical 90,000-word novel

What each context window holds

1,050,000 tokens

GPT-6 Astra, GPT-6 Sol, GPT-6 Luna

  • 787,500 words
  • 1,574 single-spaced pages
  • 8.8 average novels
  • all of the King James Bible

1,000,000 tokens

Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5, Gemini 3.1 Pro, Claude Sonnet 5, Gemini 3.8 Flash, Gemini 3.5 Flash-Lite

  • 750,000 words
  • 1,499 single-spaced pages
  • 8.3 average novels
  • 1.3× War and Peace

200,000 tokens

Claude Haiku 4.5

  • 150,000 words
  • 300 single-spaced pages
  • 1.7 average novels
  • 1.7× a typical 90,000-word novel

How a context window works

Every request to a model is assembled into one sequence of tokens: the system prompt, tool definitions, earlier turns of the conversation, any documents or search results you attach, your new message — and then the reply the model writes. All of it has to fit inside the context window.

When a request would exceed the window, it can't be processed as-is. Applications handle this by trimming or summarizing older conversation turns, retrieving only the relevant parts of long documents, or splitting the work across several requests. That is why long chats sometimes "forget" early details.

Context window vs. maximum output

The window is shared by input and output, but the output also has its own, much smaller limit. Claude Opus 5.5 and Claude Fable 5.1, for example, accept up to 1M tokens in total but write at most 128K tokens in one response. A big window lets a model read a lot; it doesn't mean it will write that much.

Does a bigger context window cost more?

You pay per token, so filling a large window is expensive: a 1M-token prompt on a model charging $2 per million input tokens costs $2 every time you send it. Some providers also add a long-context surcharge —Gemini 3.1 Pro charges more for prompts above 200K tokens — while Anthropic bills the full 1M window of Claude 4.6 and later models at the standard rate. Prompt caching cuts the cost of re-sending the same long context; see the Claude, OpenAI andGemini pricing pages for cached-input prices.

To see how much of a window your own text would use, paste it into the token counter.

Frequently asked questions

What is a context window?

A context window is the maximum number of tokens an AI model can work with in a single request — your prompt, any documents or conversation history, and the model's reply combined. As of October 2026, the largest belong to GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna at 1,050,000 tokens, roughly 787,500 words.

Which AI model has the largest context window?

As of October 2026, GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna have the largest context windows at 1,050,000 tokens. That is about 787,500 words, or all of the King James Bible.

Does the output count against the context window?

Yes. The context window covers input and output together, so a long reply leaves less room for input. Models also have a separate cap on output length that is far smaller than the window — Claude Opus 5.5 and Claude Fable 5.1, for example, have a 1M-token window but can write at most 128K tokens in one response.

Does using more of the context window cost more?

You pay for every token you send, so a longer prompt always costs more in total. Some models also raise the per-token price for very long prompts: Gemini 3.1 Pro charges more above 200K tokens, while Anthropic bills the full 1M-token window of Claude 4.6 and later models at the standard rate.

Is a context window the same as memory?

No. A model only sees what is inside the context window of the current request. Chat apps that seem to remember earlier conversations do it by re-sending past messages, summaries or retrieved notes as part of the context each time.