13k.euES
Menu

Calculator · AI APIs & Costs

LLM API pricing compared, with a cost calculator

By 13k.eu editorsUpdated and checked 4 min read

Short answer

Per million tokens, prices run from $0.10 input and $0.50 output (GPT-6 Luna) to $10 and $50 (GPT-6 Astra, Claude Fable 5.1). List prices alone mislead: GPT-6.1 Sol and Claude Sonnet 5.5 both list $2 and $10, yet a sample workload costs $194 a month on one and $256 on the other once cache rates and Anthropic's 30% tokenizer change are counted. Batch halves the bill where offered.

A data center aisle with streams of small orange data packets flowing between server racks

Prices, limits and features change often. We date every figure and link to its source: check the vendor's page before you buy or build. How we make money.

Every LLM provider lists a price per million input and output tokens, but those two numbers rarely predict the bill. Caching, batch discounts, long prompts, reasoning tokens and even how each model counts tokens can move the real cost by a factor of two or more. This page compares 12 current models from OpenAI, Anthropic, Google and DeepSeek, with prices taken from each provider's own pricing page on October 1, 2026, and a calculator that applies the rules they publish.

Your workload (example values)

Prompt, instructions, documents and conversation history. Up to 200,000.

Include reasoning tokens: providers bill them as output.

%

For repeated instructions or documents. Cache writes and storage are not included.

Options

Estimated monthly cost, cheapest first

  • GPT-6 Luna (OpenAI)$9.84$0.0005 per request
  • DeepSeek V4.1 Flash (DeepSeek)$25.34$0.0013 per request · peak-hour rate (off-peak is half)
  • Gemini 3.1 Flash-Lite (Google)$27.60$0.0014 per request
  • Gemini 3.8 Flash (Google)$73.80$0.0037 per request · price until 2026-12-31, then doubles
  • DeepSeek V4 Pro (DeepSeek)$96.10$0.0048 per request · peak-hour rate (off-peak is half)
  • Claude Haiku 4.5 (Anthropic)$98.40$0.0049 per request
  • GPT-6.1 Sol (OpenAI)$194.40$0.0097 per request
  • Gemini 3.1 Pro (preview) (Google)$220.80$0.011 per request · preview model
  • Claude Sonnet 5.5 (Anthropic)$255.84$0.0128 per request · +30% tokens
  • Claude Opus 5.5 (Anthropic)$505.44$0.0253 per request · +30% tokens
  • GPT-6 Astra (OpenAI)$984.00$0.0492 per request
  • Claude Fable 5.1 (Anthropic)$1,255.80$0.0628 per request · +30% tokens
  • Prices in US dollars per million tokens from each provider's pricing page, checked October 1, 2026, before tax.
  • Up to 200,000 input tokens per request: above that, Google (Gemini Pro) and OpenAI charge long-context rates that this calculator does not model.
  • Different providers count the same text as different numbers of tokens. The only adjustment applied is the one Anthropic documents for its own models.
  • Nothing you enter leaves your browser.

Price per million tokens

API prices per million tokens (US dollars, standard rates, prompts up to 200K tokens)
ModelInputCached inputOutputBatch
GPT-6 Astra (OpenAI)$10.00$1.00$50.0050% off
GPT-6.1 Sol (OpenAI)$2.00$0.10$10.0050% off
GPT-6 Luna (OpenAI)$0.10$0.01$0.5050% off
Claude Fable 5.1 (Anthropic)$10.00$0.25$50.0050% off
Claude Opus 5.5 (Anthropic)$4.00$0.20$20.0050% off
Claude Sonnet 5.5 (Anthropic)$2.00$0.20$10.0050% off
Claude Haiku 4.5 (Anthropic)$1.00$0.10$5.0050% off
Gemini 3.1 Pro (preview) (Google)$2.00$0.20$12.00Not listed
Gemini 3.8 Flash (Google)$0.75$0.075$3.7550% off
Gemini 3.1 Flash-Lite (Google)$0.25$0.025$1.50Not listed
DeepSeek V4 Pro (DeepSeek)$1.32$0.044$3.96Not listed
DeepSeek V4.1 Flash (DeepSeek)$0.30$0.006$1.20Not listed

DeepSeek's prices are its peak-hour rates; off-peak rates are half. Gemini 3.1 Pro is listed by Google as a preview model.

Same list price, different bill

GPT-6.1 Sol and Claude Sonnet 5.5 both list $2 per million input tokens and $10 per million output tokens. Take the calculator's example workload: 20,000 requests a month, each with 3,000 input tokens (40% of them repeated instructions read from cache) and 600 output tokens.

GPT-6.1 Sol Claude Sonnet 5.5
Monthly cost at list prices $194.40 $196.80
Adjusted for Anthropic's newer tokenizer (+30% tokens) n/a $255.84
With batch processing (50% off) $97.20 $127.92

Two things explain the gap:

  1. Cache reads cost different amounts. OpenAI charges 5% of the input price for cached tokens on GPT-6.1 Sol ($0.10); Anthropic charges 10% on Sonnet 5.5 ($0.20).
  2. The same text becomes more tokens. Anthropic's own pricing page says its tokenizer for Claude 4.7 and later models "produces approximately 30% more tokens for the same text" than its previous one. A price per token is only comparable if the token counts are.

The 30% figure compares Anthropic's two tokenizers, not Anthropic with OpenAI or Google, and the exact change "depends on the content". For a firm number, count tokens for a sample of your real prompts with each provider's tools before you commit. For the basics of what a token is, see what is a token in AI.

What else moves the bill

Output costs several times more than input. Across these models, output tokens cost between 3 and 6 times as much as input tokens, and 5 times for most of them. Reasoning models also bill their "thinking" as output: Google's table says so explicitly ("including thinking tokens"). Long answers and heavy reasoning dominate the bill more than long prompts.

Caching. Repeated content (system prompts, documents, tool definitions) can be cached:

  • OpenAI: cached input at 5% of the input price on GPT-6.1 Sol; writing to the cache costs $2.50 per million tokens on that model.
  • Anthropic: reading from cache costs 0.1x the input price (0.05x on Opus 5.5 and 0.025x on Fable 5.1); writing costs 1.25x for a 5-minute cache or 2x for a 1-hour cache. Anthropic notes that caching "pays off after one cache read" for the 5-minute duration.
  • Google: reduced input price for cached content plus a storage charge per million tokens per hour.
  • DeepSeek: a separate, much lower price for cache hits.

The calculator includes cache reads but not cache writes or storage, so it slightly flatters heavy caching.

Batch processing. OpenAI, Anthropic and Google (for Gemini 3.8 Flash) offer 50% off for requests you can wait for. Anthropic adds that batch and caching discounts can be combined.

Long prompts. Thresholds differ:

Provider Rule above the threshold
OpenAI (GPT-6.1 Sol) Over 272K input tokens: 2x input and cache rates and 1.5x output, for the whole request
Google (Gemini 3.1 Pro) Over 200K tokens: $4 input and $18 output instead of $2 and $12
Anthropic (Claude 4.6 and later) The full 1M-token window at standard prices

If you routinely send very long prompts, Anthropic's flat pricing changes the comparison; the calculator stops at 200,000 input tokens for that reason.

Where the data is processed. OpenAI adds 10% for regional processing (data residency) on models released on or after March 5, 2026. Anthropic charges 1.1x when you require US-only inference on Claude 4.6 and later models.

Time of day. DeepSeek charges half price outside its peak hours, which are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.

Announced changes. Google lists Gemini 3.8 Flash at $0.75 input and $3.75 output through December 31, 2026, and $1.50 and $7.50 from January 1, 2027.

Free tiers and your data

Google's Gemini API has a free tier for some models, but its pricing table says data from the free tier is "used to improve our products" while paid-tier data is not. If you send customer or company data, use a paid tier and read the provider's data terms.

How to choose

  1. Estimate your real token counts with a sample of production prompts, including the system prompt and history.
  2. Run them through the calculator with your cache share; check whether batch works for your use case.
  3. Shortlist two or three models across price tiers and test quality on your own tasks, starting with the cheaper ones. Anthropic's own guidance for developers suggests routing easy or common requests to smaller, cheaper models and hard ones to more capable models.

How we checked

Prices were read from the providers' own pages on October 1, 2026: OpenAI's API pricing page (Standard and Batch tables) and model page for GPT-6.1 Sol, Anthropic's Claude pricing documentation, Google's Gemini Developer API pricing page and DeepSeek's models and pricing page. The calculator's formula and the example figures are covered by automated tests. We have not compared output quality, speed or rate limits.

What we checked

  • OpenAI standard prices per 1M tokens: GPT-6 Astra $10/$1 cached/$50; GPT-6.1 Sol $2/$0.10/$10 (cache writes $2.50); GPT-6 Luna $0.10/$0.01/$0.50; Batch at half price; regional processing +10% for models released on or after March 5, 2026. (OpenAI, )
  • GPT-6.1 Sol: 1,050,000-token context, cached input at 5% of the input rate, prompts over 272K input tokens at 2x input and cache and 1.5x output for the full request. (OpenAI, )
  • Anthropic prices per 1M tokens: Fable 5.1 $10/$50, Opus 5.5 $4/$20, Sonnet 5.5 $2/$10, Haiku 4.5 $1/$5; cache reads 0.1x (0.05x Opus 5.5, 0.025x Fable 5.1), writes 1.25x or 2x; Batch 50%; full 1M context at standard prices on Claude 4.6+; US-only inference 1.1x; Claude 4.7+ tokenizer produces about 30% more tokens. (Anthropic, )
  • Google prices per 1M tokens (paid tier): Gemini 3.1 Pro preview $2/$12 up to 200K tokens and $4/$18 above; Gemini 3.8 Flash $0.75/$3.75 through Dec 31, 2026 and $1.50/$7.50 from Jan 1, 2027, Batch at half; Gemini 3.1 Flash-Lite $0.25/$1.50; output includes thinking tokens; free-tier data used to improve products, paid-tier data not. (Google AI for Developers, )
  • DeepSeek peak prices per 1M tokens: V4.1 Flash $0.30 input, $0.006 cache hit, $1.20 output; V4 Pro $1.32, $0.044, $3.96; off-peak is half; peak 01:00-04:00 and 06:00-10:00 UTC on weekdays. (DeepSeek, )

What may change

  • API prices and models change often; Google has already announced a price increase for Gemini 3.8 Flash on January 1, 2027.
  • Promotional prices (such as OpenAI's for GPT-5.6 Sol) are not included.
  • Tokenizer differences between providers are not published; only Anthropic documents its own change.

Frequently asked questions

What is the cheapest LLM API?

Of the 12 models compared on October 1, 2026, GPT-6 Luna has the lowest list price ($0.10 input, $0.50 output per million tokens), followed by Gemini 3.1 Flash-Lite and DeepSeek V4.1 Flash. Cheapest is not always best value: test quality on your own tasks.

Why does the same price per token give different bills?

Because cache prices differ and models count the same text as different numbers of tokens. Anthropic says its tokenizer for Claude 4.7 and later produces about 30% more tokens for the same text than its previous one.

How much does the Batch API save?

50% at OpenAI, Anthropic and Google (for Gemini 3.8 Flash), for requests that can wait for an asynchronous result. Anthropic says batch and caching discounts can be combined.

Do long prompts cost more?

At some providers. OpenAI charges 2x input and 1.5x output above 272K input tokens on GPT-6.1 Sol; Google charges more above 200K on Gemini 3.1 Pro; Anthropic charges standard rates across the full 1M window on Claude 4.6 and later.

Sources

  1. OpenAI API pricing (Standard and Batch, per 1M tokens), OpenAI. Accessed October 1, 2026.
  2. GPT-6.1 Sol model page: context window, cached input and long-context pricing, OpenAI. Accessed October 1, 2026.
  3. Claude API pricing: models, prompt caching, batch, long context and tokenizer note, Anthropic. Accessed October 1, 2026.
  4. Gemini Developer API pricing, Google AI for Developers. Accessed October 1, 2026.
  5. DeepSeek API: models and pricing (peak and off-peak), DeepSeek. Accessed October 1, 2026.

Spotted an error or an outdated price? Tell us and we will fix it.

Change history

  • : First published with 12 models from OpenAI, Anthropic, Google and DeepSeek and a cost calculator.

Next review: .

Part of our AI APIs & Costs guide.