Context window comparison, 2026
A context window caps how much input a model can attend to in a single request. Max output caps how long a single response can be. The two limits are separate, and providers publish them separately. This page collects both numbers for the current models, verified against provider documentation.
Window size also sets the ceiling on cost per request. Input billing scales with the tokens you send, so a model that accepts a million tokens can also bill you for a million tokens in one call. A bigger window is a bigger budget ceiling, not a discount. See the pricing pages for per-token rates.
Every current model, side by side
Verified 2026-08-10.
| Model | Context window (tokens) | Max output (tokens) |
|---|---|---|
| GPT-5.6 Sol | 1,050,000 | 128,000 |
| Gemini 2.5 Pro | 1,048,576 | 65,536 |
| Gemini 2.5 Flash | 1,048,576 | — |
| GPT-5.5 | 1,000,000 | — |
| Claude Fable 5 | 1,000,000 | 128,000 |
| Claude Opus 5 | 1,000,000 | 128,000 |
| Claude Opus 4.8 | 1,000,000 | 128,000 |
| Claude Sonnet 5 | 1,000,000 | 128,000 |
| Claude Sonnet 4.6 | 1,000,000 | 128,000 |
| Kimi K3 | 1,048,576 | — |
| DeepSeek V4 Flash | 1,000,000 | 384,000 |
| DeepSeek V4 Pro | 1,000,000 | 384,000 |
| Grok 4.3 | 1,000,000 | — |
| Grok 4.5 | 500,000 | — |
| GPT-5 | 400,000 | 128,000 |
| Kimi K2.6 | 262,144 | — |
| Grok 4 | 256,000 | — |
| Claude Haiku 4.5 | 200,000 | 64,000 |
| GPT-4o | 128,000 | 16,384 |
A dash means the provider does not publish a verified max output figure that we track for that model.
Will your input fit?
Paste text or enter a token count to check it against any window in the table.
Are there models with a 10 million token context window in 2026?
Not among the models verified here. As of 2026-08-10, no model in the table above ships a 10,000,000-token window. The largest verified windows are around one million tokens: GPT-5.6 Sol at 1,050,000, Gemini 2.5 Pro, Gemini 2.5 Flash, and Kimi K3 at 1,048,576, and GPT-5.5, the current Claude models, DeepSeek V4, and Grok 4.3 at 1,000,000. If you see a 10M claim, treat it as unverified marketing until the provider documents the number in its API reference. We will update this table when a provider does.
Two cautions before you fill a big window
First, cost. A full one-million-token prompt is billed as one million input tokens. On Gemini 2.5 Pro the per-token price itself rises once a prompt passes 200,000 input tokens, which is easy to miss when budgeting long-context workloads. Details are on the Gemini 2.5 Pro page.
Second, quality. Long prompts can degrade retrieval of facts placed in the middle of the window, an effect well documented in long-context evaluations. A window that accepts your document does not guarantee the model uses all of it equally well. When accuracy matters, keep the key material near the start or end of the prompt and test at your real prompt length.
Model pages
For per-token rates on these models, see API pricing. To count tokens in your own text, start at the token calculator.