Free tool
What does your AI usage actually cost?
Start from a real-world workload, or enter your own monthly token volume, and see the estimated cost across Claude, GPT, and Gemini models side by side — computed live in your browser, nothing submitted or stored.
Free tool
What would your AI usage cost on each model?
Pick a model and drop in your real monthly token volume. Everything below computes in your browser — nothing you type here is sent anywhere.
Not sure of your volume? Start from a real-world workload:
Grouped by provider.
Everything you send to the model — prompts, context, tool results.
What the model generates back — usually priced several times higher than input.
Share of input tokens served from a prompt cache (cache reads cost ≈10% of a normal input token).
- Annualized
- $3,600
You could save ~$273/mo by switching to Gemini 3.1 Flash-Lite for this workload.
List prices, not a quote — verify current pricing with each provider before budgeting. Prices as of 2026-07-08: Anthropic, OpenAI, Google. Computes entirely in your browser — nothing typed here is sent anywhere.
Compare every model at your usage
| Model | Tier | Input $/1M | Output $/1M | Context | Your est. monthly cost |
|---|---|---|---|---|---|
| Claude Opus 4.8 | Frontier | $5.00 | $25.00 | 1M | $500 |
| Claude Sonnet 5† | Balanced | $3.00 | $15.00 | 1M | $300 |
| Claude Haiku 4.5 | Fast | $1.00 | $5.00 | 200K | $100 |
| GPT-5.5 | Frontier | $5.00 | $30.00 | 1.1M | $550 |
| GPT-5.4 | Balanced | $2.50 | $15.00 | 1.1M | $275 |
| GPT-5.4 mini | Fast | $0.75 | $4.50 | 400K | $83 |
| Gemini 3.1 Pro Preview† | Frontier | $2.00 | $12.00 | 1M | $220 |
| Gemini 3.5 Flash | Balanced | $1.50 | $9.00 | 1M | $165 |
| Gemini 3.1 Flash-Lite | Fast | $0.25 | $1.50 | 1M | $28 |
- † Claude Sonnet 5: intro pricing $2 / $10 thru 2026-08-31
- † Gemini 3.1 Pro Preview: preview pricing, ≤200K token prompts
Just need the raw numbers? See the LLM API pricing reference for every model's list price in one table, no inputs required.
FAQ
Common questions about LLM pricing
I don't know my token volume — where do I start?
Use one of the workload presets above the calculator (support chatbot, RAG assistant, coding assistant, batch summarization, classification). Each one fills in a realistic monthly input/output token volume and cache-hit rate from a stated, transparent scenario — pick the closest match, then adjust the numbers once you have your own usage data.
How is LLM API cost calculated?
Model providers bill per token, separately for input (what you send) and output (what the model generates), each at its own per-million-token rate. Monthly cost is roughly (input tokens ÷ 1,000,000 × input rate) + (output tokens ÷ 1,000,000 × output rate) — this calculator does that math live for whichever model you pick.
What's the difference between input and output tokens?
Input tokens are everything sent to the model — the prompt, system instructions, retrieved context, prior conversation turns, tool results. Output tokens are what the model generates back. Output is typically priced several times higher than input, so response length and format matter as much as prompt size.
Does prompt caching really cut cost?
Yes, when your requests repeat the same large system prompt, tool definitions, or reference document. Cache reads typically cost a small fraction of a normal input token, so a high cache-hit rate on a stable prefix can meaningfully cut your input bill. Try the caching slider above to see the effect on your own volume.
Which model is cheapest for my workload?
It depends on your input/output mix and how much of that input is repeated (cacheable). The comparison table above recalculates every model's cost at your entered usage and highlights the cheapest option — it's rarely the same answer for every workload, which is why it's worth checking rather than assuming.
How do I track what I'm actually spending?
This calculator estimates cost from list prices and a workload you enter by hand — useful for planning, but it isn't your real bill. CostMon connects directly to your provider's cost API (starting with the Anthropic Admin cost report) and normalizes actual AI spend into the same dashboard as your cloud and SaaS costs, so you see what you're really spending, daily, not a monthly total that arrives too late to act on.
See what your stack really costs.
Connect your first providers in minutes and give finance and engineering one number they both trust. Flat pricing — never a percentage of your bill or your savings.