AI Token Price

Local LLM speed and cost calculator

How many tokens per second an open model runs at on your GPU or Mac, whether it fits in memory, and whether owning the hardware beats paying for the same model through an API.

Estimates from memory bandwidth, calibrated against published llama.cpp benchmarks.

Generation speed
82 tok/s
likely 61–98
Memory needed
18.2 GB
of 31 GB usable · KV cache 0.5 GB
Fit
Fits in memory
Runs fully on the device
500-token answer
6.1 s
generation only

Calibrated against published results: gpt-oss-20b on an RTX 5090 measures 282 tok/s (estimate ≈290); gpt-oss-120b on a DGX Spark 55–60 (≈55); Qwen3.5-35B-A3B on an RTX 5090 194 (≈191).

The same model on every device

Qwen3.8 27B at Q4_K_M, 8K context. Select a row to switch hardware.

RTX 5090
RTX 5080
RTX 5070 Ti
RTX 5060 Ti 16 GB
RTX 4090 (used)
RTX 3090 (used)
RTX 3060 12 GB
RTX PRO 6000 Blackwell
Radeon RX 9070 XT
Radeon RX 7900 XTX
Radeon AI PRO R9700
Intel Arc Pro B70
Mac Studio M5 Max 128 GB
Mac Studio M5 Ultra 512 GB
Ryzen AI Max+ 395 128 GB
NVIDIA DGX Spark

Does owning the hardware pay off?

Against hosted prices for the same open model, a local GPU rarely pays for itself on electricity savings alone. It makes sense for privacy, offline work, heavy steady use, or when the alternative is a closed model at a higher price.

US street price, September 2026. Edit for your market.
API cost per month
$140
Qwen3.8 27B · $0.42 / $3 per 1M
Electricity per month
$13.11
455 W while busy
Busy time per day
3.8 h
generation and prompt processing
Hardware pays back in
3.6 years
saves $127 a month after power

Prompt processing counted at about 30× generation speed. Excludes idle power, your time, the PC around a graphics card, and resale value.

How the estimate works

To generate each token, the hardware reads every active weight from memory once. Generation speed is therefore limited by memory bandwidth. Mixture-of-experts models read only their active experts per token, so a 120B MoE model can run faster than a 70B dense one while still needing memory for all 120B parameters.

weights (GB)   = total params × bits per weight ÷ 8
KV cache (GB)  = KV bytes per token × context
memory needed  = weights + KV cache + ~1 GB
read per token = active params × bits ÷ 8 + KV cache
tokens/s       ≈ 1 ÷ (read ÷ (efficiency × bandwidth) + overhead)

Efficiency and per-token overhead are set per device and per architecture so that the estimates reproduce published measurements. Small-active MoE models reach only 35–60% of peak bandwidth on fast GPUs, and a flat efficiency factor would overestimate them by about 2×. Speeds also fall faster at long context than the cache reads alone predict.

Calibration points

Hardware and modelMeasured
RTX 5090gpt-oss-20b (MXFP4)282 tok/s Source
NVIDIA DGX Sparkgpt-oss-120b (MXFP4)55–60 tok/s Source
RTX 5090Qwen3.5-35B-A3B (Q4_K_M)194 tok/s Source
RTX 3060 12 GBLlama 2 7B (Q4_0)75.6 tok/s Source
Ryzen AI Max+ 395 128 GBgpt-oss-120b (MXFP4)52–54 tok/s Source

Hardware prices are approximate US street prices in September 2026, when memory shortages pushed many GPUs 45–240% above list price.

More calculators