Grok 4 context window: 256,000 tokens
Grok 4, from xAI, has a context window of 256,000 tokens. The context window caps the total input the model can attend to in a single request. Anything past the limit has to be cut, summarized, or split across requests before the model sees it.
Verified 2026-08-10.
| Model | Context window (tokens) |
|---|---|
| Grok 4 | 256,000 |
Where 256K sits among current models
Grok 4 lands in the middle band of the current lineup. It is double the 128,000-token window of GPT-4o, close to Kimi K2.6 at 262,144, and above Claude Haiku 4.5 at 200,000. Within xAI's own lineup it is now the smallest window: Grok 4.5 takes 500,000 tokens and Grok 4.3 reaches 1,000,000, per xAI's model documentation.
Above it, GPT-5 offers 400,000 tokens, and the top of the table is the million-token group: GPT-5.5, the current Claude models, DeepSeek V4, and Grok 4.3 at 1,000,000, Gemini 2.5 Pro, Gemini 2.5 Flash, and Kimi K3 at 1,048,576, and GPT-5.6 Sol at 1,050,000. Grok 4 holds about a quarter of what those models accept in one request.
| Model | Context window (tokens) |
|---|---|
| GPT-4o | 128,000 |
| Claude Haiku 4.5 | 200,000 |
| Grok 4 | 256,000 |
| Kimi K2.6 | 262,144 |
| GPT-5 | 400,000 |
| Grok 4.5 | 500,000 |
| GPT-5.5 | 1,000,000 |
| Claude Sonnet 5 | 1,000,000 |
| DeepSeek V4 Flash | 1,000,000 |
| Grok 4.3 | 1,000,000 |
| Kimi K3 | 1,048,576 |
| Gemini 2.5 Pro | 1,048,576 |
| GPT-5.6 Sol | 1,050,000 |
Is 256K enough?
For most single-document work, yes. A 256K window handles long reports, sizable transcripts, and multi-file code questions in one request. The workloads that need more are the whole-corpus jobs: full repositories, large document collections, or long multi-session agent histories. Those are the cases where the million-token models earn their place.
Two general points apply at any size. Input billing scales with the tokens you send, so a fuller window is a costlier request. And long prompts can degrade retrieval of facts placed in the middle of the window, an effect well documented in long-context evaluations, so test at your real prompt length rather than assuming full-window accuracy.
For every current model side by side, see the context window comparison.