13k.euES
Menu

Explainer · AI Explained

What is a token in AI, and what is a context window?

By 13k.eu editorsUpdated and checked 3 min read

Short answer

A token is the unit a model reads and writes: a character, part of a word or a whole word. In English, OpenAI estimates 100 tokens at about 75 words and Google at 60 to 80, but counts change with the model and the language. Context windows, answer limits and API prices are all measured in tokens, and reasoning tokens are billed as output even though you do not see them.

A 3D lattice of thin steel lines and nodes with a few pathways lit in orange, like a neural network

Prices, limits and features change often. We date every figure and link to its source: check the vendor's page before you buy or build. How we make money.

A token is the unit a language model reads and writes. Models do not see letters or words directly: they split text into tokens, which can be a single character, part of a word, a whole word or a punctuation mark. Everything about using a model is measured in tokens: how much text it can take in at once (its context window), how long its answer can be and how much the API charges.

How text becomes tokens

OpenAI describes the process in three steps: the text is divided into tokens, the model processes them, and the model generates output tokens. Google gives the examples of a single character such as "z" or a whole word such as "cat", explains that "long words are broken up into several tokens", and calls the set of all tokens a model knows its vocabulary.

Details matter more than you might expect. OpenAI notes that "red", "Red" and " red" (with a leading space) are different text and can become different tokens. Token IDs also depend on the model's encoding.

How many words is a token?

Each provider publishes a rule of thumb for English:

Source Rule of thumb 1,000 English words is about
OpenAI 1 token ≈ 4 characters ≈ ¾ of a word; 100 tokens ≈ 75 words 1,330 tokens
Google (Gemini) 1 token ≈ 4 characters; 100 tokens ≈ 60 to 80 words 1,250 to 1,670 tokens

These are estimates, and both companies say so. OpenAI adds that "the same text can produce different token counts depending on the model, its encoding, and the language", and that other languages can have different ratios between characters, words and tokens. Even within one company, counts change between model generations: Anthropic says the tokenizer in Claude 4.7 and later models "produces approximately 30% more tokens for the same text" than its previous one.

Images, audio and video are counted in tokens too. Google's documentation gives about 6,000 tokens for a 60-second video and about 1,920 tokens for 60 seconds of audio on Gemini models.

Input, output, cached and reasoning tokens

APIs bill tokens in several categories:

  • Input tokens: everything you send, including instructions, documents, conversation history and tool definitions.
  • Output tokens: what the model writes. They usually cost several times more than input; see LLM API pricing compared.
  • Cached input tokens: repeated input the provider can reuse, billed at a lower rate.
  • Reasoning (thinking) tokens: reasoning models "think" before answering. OpenAI says these tokens are not shown as answer text but "count toward output usage and are billed as output tokens", and Anthropic says the same of Claude's thinking tokens. A short answer can therefore use far more tokens than it seems.

What a context window is

The context window is the maximum number of tokens a model can work with in one request: your input plus the output it generates. On October 1, 2026:

Model or plan Context window
Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 (API) 1M tokens, with up to 128K output tokens per request
Other Claude models, such as Sonnet 4.5 200K tokens
GPT-6.1 Sol (API) 1,050,000 tokens, with up to 128,000 output tokens
ChatGPT Plus (app), fast model 54K tokens
ChatGPT Plus (app), reasoning models 256K tokens

The window in a chat app is not always the model's full window. OpenAI notes that in ChatGPT the space for your input is smaller than the total, because system instructions, memory and the model's own reasoning share it. When a conversation outgrows the window, older parts have to be dropped or summarized, which is why long chats can "forget" their beginning. For a plan-by-plan comparison, see Claude vs ChatGPT.

Running models on your own computer, the context window also costs memory: on Llama 3.1 8B, the cache that holds the context takes 1 GiB at 8,192 tokens and 16 GiB at 131,072. Our VRAM calculator works it out for other models.

How to count tokens exactly

  • OpenAI: its Tokenizer page shows how text splits; for code, the tiktoken library with the encoding for your model; for a full request including images, files and tools, its input-token counting API.
  • Anthropic: a token counting API that is "free to use", subject to rate limits.
  • Google: a token counting method in the Gemini API, and usage figures returned with each response.

For anything that affects a budget, count a sample of your real prompts with the provider's own tool instead of relying on rules of thumb.

What we checked

  • OpenAI: 1 token is about 4 characters or three-quarters of a word, 100 tokens about 75 words; counts depend on the model, encoding and language; reasoning tokens are billed as output; Tokenizer, tiktoken and an input-token counting API. (OpenAI Help Center, )
  • Google: for Gemini, 1 token is about 4 characters and 100 tokens about 60 to 80 English words; long words become several tokens; a 60-second video is about 6,000 tokens and 60 seconds of audio about 1,920; count_tokens method. (Google AI for Developers, )
  • Anthropic's tokenizer for Claude 4.7 and later produces about 30% more tokens for the same text. (Anthropic, )
  • Anthropic token counting is free to use, subject to rate limits. (Anthropic, )
  • Claude Fable 5.1, Opus 5.5 and Sonnet 5.5 have a 1M-token context window with up to 128K output tokens; other models such as Sonnet 4.5 have 200K; thinking tokens are billed as output. (Anthropic, )
  • GPT-6.1 Sol has a 1,050,000-token context window and 128,000 max output tokens. (OpenAI, )
  • ChatGPT Plus: 54K tokens for the fast model and 256K for reasoning models; input space is smaller than the total window. (OpenAI, )

What may change

  • Context windows and tokenizers change with new models.
  • Chat apps may change how much of the window they make available on each plan.

Frequently asked questions

How many words is 1,000 tokens?

In English, about 750 words by OpenAI's rule of thumb (100 tokens ≈ 75 words) and 600 to 800 by Google's (100 tokens ≈ 60 to 80 words). Other languages and models give different counts.

What is a context window?

The maximum number of tokens a model can handle in one request, counting your input and its output. On October 1, 2026, current Claude models and GPT-6.1 Sol offered about 1 million tokens through their APIs.

Why did my short answer use so many tokens?

Reasoning models think before they answer. Those reasoning tokens are not shown but are billed as output tokens, according to both OpenAI and Anthropic.

How do I count tokens exactly?

Use the provider's own tools: OpenAI's Tokenizer, tiktoken or input-token counting API; Anthropic's free token counting API; or count_tokens in the Gemini API.

Sources

  1. Understanding and counting tokens, OpenAI Help Center. Accessed October 1, 2026.
  2. Understand and count tokens (Gemini API), Google AI for Developers. Accessed October 1, 2026.
  3. Claude API pricing: models, prompt caching, batch, long context and tokenizer note, Anthropic. Accessed October 1, 2026.
  4. Token counting (Claude API), Anthropic. Accessed October 1, 2026.
  5. Context windows (Claude API): sizes by model and thinking tokens, Anthropic. Accessed October 1, 2026.
  6. GPT-6.1 Sol model page: context window, cached input and long-context pricing, OpenAI. Accessed October 1, 2026.
  7. ChatGPT pricing: plans, models, context windows and features, OpenAI. Accessed October 1, 2026.

Spotted an error or an outdated price? Tell us and we will fix it.

Change history

  • : First published, with token rules of thumb from OpenAI and Google and context windows as of October 2026.

Next review: .

Part of our AI Explained guide.