Local LLM speed and cost calculator
How many tokens per second an open model runs at on your GPU or Mac, whether it fits in memory, and whether owning the hardware beats paying for the same model through an API.
Estimates from memory bandwidth, calibrated against published llama.cpp benchmarks.
Calibrated against published results: gpt-oss-20b on an RTX 5090 measures 282 tok/s (estimate ≈290); gpt-oss-120b on a DGX Spark 55–60 (≈55); Qwen3.5-35B-A3B on an RTX 5090 194 (≈191).
The same model on every device
Qwen3.8 27B at Q4_K_M, 8K context. Select a row to switch hardware.
Does owning the hardware pay off?
Against hosted prices for the same open model, a local GPU rarely pays for itself on electricity savings alone. It makes sense for privacy, offline work, heavy steady use, or when the alternative is a closed model at a higher price.
Prompt processing counted at about 30× generation speed. Excludes idle power, your time, the PC around a graphics card, and resale value.
How the estimate works
To generate each token, the hardware reads every active weight from memory once. Generation speed is therefore limited by memory bandwidth. Mixture-of-experts models read only their active experts per token, so a 120B MoE model can run faster than a 70B dense one while still needing memory for all 120B parameters.
weights (GB) = total params × bits per weight ÷ 8
KV cache (GB) = KV bytes per token × context
memory needed = weights + KV cache + ~1 GB
read per token = active params × bits ÷ 8 + KV cache
tokens/s ≈ 1 ÷ (read ÷ (efficiency × bandwidth) + overhead)
Efficiency and per-token overhead are set per device and per architecture so that the estimates reproduce published measurements. Small-active MoE models reach only 35–60% of peak bandwidth on fast GPUs, and a flat efficiency factor would overestimate them by about 2×. Speeds also fall faster at long context than the cache reads alone predict.
Calibration points
| Hardware and model | Measured |
|---|---|
| RTX 5090gpt-oss-20b (MXFP4) | 282 tok/s Source |
| NVIDIA DGX Sparkgpt-oss-120b (MXFP4) | 55–60 tok/s Source |
| RTX 5090Qwen3.5-35B-A3B (Q4_K_M) | 194 tok/s Source |
| RTX 3060 12 GBLlama 2 7B (Q4_0) | 75.6 tok/s Source |
| Ryzen AI Max+ 395 128 GBgpt-oss-120b (MXFP4) | 52–54 tok/s Source |
Hardware prices are approximate US street prices in September 2026, when memory shortages pushed many GPUs 45–240% above list price.