AI & LLM cost
Cost per 1M tokens
The standard unit LLM providers use to publish pricing: the dollar cost of one million input or output tokens, since a single token is too small a unit to price meaningfully.
Last updated
Definition
Because a single token is worth a small fraction of a cent, providers publish rate cards in dollars per million tokens, typically as separate input and output rates. Comparing models on this number alone is a useful first filter but not the full picture. A cheaper per-token rate on a less capable model can lose to a pricier model that needs fewer follow-up calls or produces a usable answer on the first attempt.
The number that predicts your bill is this rate multiplied by your real token volume per model and per workload, which is what a normalized cost view needs to expose to be useful.
Where it shows up
Cost per 1M tokens is the unit every provider publishes on its own pricing page, listed as a separate input and output rate per model, and it's the rate a usage-based invoice applies to the total token counts reported across the whole billing period, model by model.
What makes it expensive
Teams get burned by picking a model purely on its per-token rate without checking how many follow-up calls or retries it takes to get a usable answer. The rate card prices a token. The bill prices a finished task. Those two rankings are not always in the same order.
Related
Token spend moves every day.
Model prices and rate limits change between invoices. CostMon tracks your AI spend daily, so a change shows up before the bill does.