How we verify prices
Every price on this site comes from the provider’s own pricing page and shows the date it was last checked. Automated feeds only tell us when to look again; they never change a price on their own.
Where prices come from
For each model we record the price on the developer’s official pricing page or documentation, a link to that page, and the date we checked it. Open-weight models (gpt-oss, Gemma, Qwen open models, Nemotron and others) are sold by many hosts at different prices; for those we show the cheapest host listed on OpenRouter and say so on the page.
When a provider hasn’t published something (a cache price, a release date, a context window) we leave it blank rather than estimate. A few prices come only from third-party listings; those pages say so.
What “verified” means
The verified date is the last time a person compared our figure with the provider’s page. The whole table was verified on 27 Sep 2026. A handful of older models still show an earlier date because their pricing pages no longer list them in full.
Catching price changes
Providers change prices with little notice, so we also compare our data with two public price feeds: LiteLLM’s model price list and OpenRouter’s model API. When a feed disagrees with us, or lists a new model from a provider we track, a person checks the provider’s page before anything changes. The feeds are often wrong in their own ways (third-party hosts, stale entries), which is why they can’t overwrite a price.
Pricing details we handle
- Cached input: the per-million price for repeated prompt prefixes, where a provider publishes one.
- Batch: the discount for jobs that return within hours.
- Long prompts: some models charge more above a length threshold (for example Gemini 3.1 Pro above 200K tokens, GPT-5.4 and later above 272K).
- Promotions and peak pricing: temporary prices (such as Gemini 3.8 Flash until 31 December 2026) show their end date; DeepSeek’s peak and off-peak prices are both listed, with peak shown in tables.
- Regions: we show US dollar list prices. Alibaba’s are Singapore-region prices. Taxes are not included. Other currencies are converted at European Central Bank reference rates (currently 25 Sep 2026).
Tiers and scores
Frontier, workhorse and small-and-fast are our grouping by intended use, not a ranking of quality. Where we show a quality score it is a named, dated figure: the Arena text leaderboard or the Artificial Analysis Intelligence Index. We don’t publish our own quality estimates.
The token price index
The index is the median blended price (3 input : 1 output) of current first-party models in each tier. Its method is on that page.
Local speed estimates
The local LLM calculator estimates generation speed from memory bandwidth, calibrated against published llama.cpp results. Estimates are labelled as estimates; measured results will be labelled with the hardware, backend and settings used.
Provider logos
Logos help you spot each provider’s models at a glance. They are trademarks of their owners and are used only to identify those models; their use doesn’t imply any endorsement. The logo artwork comes from the open-source LobeHub Icons set (MIT licence).
Corrections
We correct errors in the data file, which regenerates every page, and record price changes in the price history.
Using the data
All pricing data is free to reuse under CC BY 4.0: JSON, CSV, changelog. Please credit AI Token Price with a link.