Skip to content

LLM API Cost Calculator

Compare what a workload actually costs across Anthropic, OpenAI, Google and xAI, including the prompt caching, batch rates and long-context tiers most calculators leave out.

Share
Start from
Prompt cache hit rate60%
Share of input tokens served from cache. Cache reads are typically a tenth of the input rate, so this is the single biggest lever on a repeated-prefix workload.
Providers
ModelInputOutputMonthly totalvs cheapest
GPT-5 nano
OpenAIgpt-5-nano
$13.80$6.40$20.201.0x
Gemini 2.5 Flash-Lite
Googlegemini-2.5-flash-lite
$27.60$6.40$34.001.7x
GPT-5.6 Luna
OpenAIgpt-5.6-luna
$55.20$19.20$74.403.7x
GPT-5.4 nano
OpenAIgpt-5.4-nano
$55.20$20.00$75.203.7x
Gemini 3.1 Flash-Lite
Googlegemini-3.1-flash-lite
$69.00$24.00$93.004.6x
GPT-5 mini
OpenAIgpt-5-mini
$69.00$32.00$101.005.0x
Gemini 3.5 Flash-Lite
Googlegemini-3.5-flash-lite
$82.80$40.00$122.806.1x
Gemini 2.5 Flash
Googlegemini-2.5-flash
$82.80$40.00$122.806.1x
Gemini 3.7 Flash
Googlegemini-3.7-flash
$207.00$60.00$267.0013.2x
Gemini 3.6 Flash
Googlegemini-3.6-flash
$207.00$60.00$267.0013.2x
GPT-5.4 mini
OpenAIgpt-5.4-mini
$207.00$72.00$279.0013.8x
Grok Build 0.1
xAIgrok-build-0.1
$312.00$32.00$344.0017.0x
Claude Haiku 4.5
Anthropicclaude-haiku-4-5
$276.00$80.00$356.0017.6x
Grok 4.3
xAIgrok-4.3
$372.00$40.00$412.0020.4x
GPT-5.1
OpenAIgpt-5.1
$345.00$160.00$505.0025.0x
Gemini 3.5 Flash
Googlegemini-3.5-flash
$414.00$144.00$558.0027.6x
Grok 4.5
xAIgrok-4.5
$588.00$96.00$684.0033.9x
Claude Sonnet 5
Anthropicclaude-sonnet-5 intro rate
$552.00$160.00$712.0035.2x
GPT-5.6 Terra
OpenAIgpt-5.6-terra
$552.00$192.00$744.0036.8x
Gemini 3.1 Pro (Preview)
Googlegemini-3.1-pro-preview
$552.00$192.00$744.0036.8x
Grok 4.6
xAIgrok-4.6
$660.00$96.00$756.0037.4x
GPT-4.1
OpenAIgpt-4.1
$660.00$128.00$788.0039.0x
GPT-5.4
OpenAIgpt-5.4
$690.00$240.00$930.0046.0x
Claude Sonnet 4.6
Anthropicclaude-sonnet-4-6
$828.00$240.00$1,06852.9x
GPT-4o
OpenAIgpt-4o
$1,050$160.00$1,21059.9x
Claude Opus 5
Anthropicclaude-opus-5
$1,380$400.00$1,78088.1x
Claude Opus 4.8
Anthropicclaude-opus-4-8
$1,380$400.00$1,78088.1x
GPT-5.6 Sol
OpenAIgpt-5.6-sol
$1,380$480.00$1,86092.1x
GPT-5.5
OpenAIgpt-5.5
$1,380$480.00$1,86092.1x
Claude Fable 5
Anthropicclaude-fable-5
$2,760$800.00$3,560176.2x
GPT-5.5 Pro
OpenAIgpt-5.5-pro
$18,000$2,880$20,8801033.7x

Across the models shown, this workload ranges from $20.20 to $20,880 a month, a 1034x spread. Model choice is almost always a larger lever than prompt optimisation.

Prices are USD per million tokens, taken from each provider's own pricing page and last verified on 2026-08-15. Standard tier only: enterprise discounts, committed-use pricing, and free tiers are not modelled. Verify against Anthropic, OpenAI, Google, xAI before committing budget.

Comparing two providers specifically? There are head-to-head pages for Claude vs GPT, Gemini vs GPT and Claude vs Gemini.

How to use LLM API Cost Calculator

1

Describe one request

Enter the input and output tokens for a single typical request, or pick a preset workload to start from.

2

Set volume and caching

Add your monthly request count, then set the share of input tokens served from the prompt cache and whether the work can run as a batch.

3

Compare the monthly bill

Models are ranked cheapest first, with the multiple over the cheapest option so you can see whether a swap is worth it.

Model choice usually beats prompt optimisation

The spread between the cheapest and most expensive model for the same workload is routinely 50x or more. Before spending a week shaving tokens off a prompt, check whether a smaller model handles the task at all, because that single decision moves the bill further than any amount of prompt engineering.

The reverse also holds. Routing everything to the cheapest model to save money often costs more overall, because a weaker model needs more retries, longer prompts, and more human correction. The useful comparison is cost per successfully completed task, not cost per million tokens.

A common structure is to route by difficulty: a small model handles the bulk of routine requests, and a frontier model handles the remainder that the small one cannot. Model the two tiers separately here and add them together.

The three costs people forget

Output is far more expensive than input, typically by 5x, and reasoning tokens are billed as output on models that produce them. A request with 100 visible output tokens can bill for thousands if the model reasoned at length first. If your workload uses extended or adaptive thinking, raise the output figure well above what you see in the response.

System prompts are charged on every request. A 4,000-token system prompt across a million requests is four billion input tokens, which is a meaningful line item on its own. This is exactly the content prompt caching exists to make cheap, which is why the cache slider is on this page rather than hidden behind an advanced toggle.

Retries and failures still bill for whatever was generated before they failed. A pipeline with a 5 percent retry rate is 5 percent more expensive than a naive calculation suggests, and rate-limit backoff loops can be considerably worse if they are retrying long prompts.

Everything runs in your browser

The pricing table ships with the page and all arithmetic happens locally. Nothing about your workload, volumes, or token counts is uploaded, which matters because request volumes and prompt sizes are commercially sensitive information that has no business being typed into someone else's server.

Frequently Asked Questions

How current are these prices?

Every rate was read off the provider's own pricing page and last verified on 2026-08-15, which is printed under the results. If that date looks stale, treat the numbers as a starting point and check the linked provider pages. LLM pricing changes often enough that an undated calculator is worthless.

Why is my real bill higher than this estimate?

The three usual causes are reasoning tokens, retries, and system prompts. Thinking or reasoning tokens are billed as output on most providers and are easy to forget. Failed requests that produced output are still billed. And a long system prompt is charged on every single request, which is why the cache slider matters so much.

What does the prompt cache hit rate actually change?

Providers charge roughly a tenth of the input rate for tokens they can serve from a cached prefix. If your requests share a large stable prefix, such as a long system prompt or a fixed document, a high hit rate can cut the input half of your bill by 80 to 90 percent. Models with no published cache rate ignore the slider.

What is a long-context rate?

Some providers charge more once a single prompt crosses a size threshold. Every xAI Grok model doubles its rate at 200,000 tokens, for example. Most calculators ignore this and under-report long-context workloads by half, so this one applies the higher tier automatically and labels the row when it does.

Does this include enterprise or committed-use discounts?

No. These are standard published list rates. Negotiated pricing, committed-use agreements, and free tiers are not modelled, so treat the figures as an upper bound if you have a contract.