Alibaba · 8.2B parameters · 41K context · commercial OK
Local (your GPU) $0 · ~9 GB VRAM | Rent a GPU from $0.06/hr | Your API key $0.29/1M |
Best next step
Already have 9 GB of VRAM? Run Q4_K_M locally for $0 compute. Open the local setup guide →
No matching GPU? For a 4,000-token prompt + 2,000-token answer, an API is about $0.0014 — no GPU setup. Open the API option →
Decision basis: 4,000 input + 2,000 output tokens, rate table dated 2026-08-24. Size-based API estimates never beat a provider-sourced rate here; referral status never changes the winner.
Novita AI offers model APIs and GPU instances. We do not include its prices in Spanvero's ranking yet, so compare Novita's current price and fit yourself.
Check Novita AI — referral link. You use your own account and pay Novita's normal price; Novita may pay Spanvero a commission.
Watch this price — free
Get an email when Qwen3-8B's tracked price drops — checked against the append-only daily record. No account needed.
Free accounts watch one price. Exact multi-price targets remain available to existing Pro accounts; the old subscription is no longer sold — see current products.
Dense 8B Qwen3 with hybrid thinking modes — a popular, laptop-friendly Apache-2.0 model.
$0.286/1M blended input + output
Published price logged daily since June 23, 2026. See the full price record.
Above ~11,000,000 tokens/day, a used RTX 3090 beats the cheapest API for Qwen3-8B ($0.29/1M).
Local is never $0 — every figure here includes electricity and hardware amortization. Speeds come from published 7-8B Q4 benchmarks (your model, quant, and context will differ); break-even volumes are rounded to 2 significant figures.
| Parameters | 8.2B |
|---|---|
| Context window | 41K tokens |
| Recommended quant | Q4_K_M |
| VRAM requirement | ~9 GB (at Q4_K_M, 16.4K context) |
| Download size | ~5 GB |
| License | Commercial use OK |
Open the free LLM cost calculator → to model your workload, monthly cost, and break-even assumptions.
Three paths: $0 marginal compute on already-owned hardware if the calculated 9 GB VRAM baseline at Q4_K_M fits your exact build and runner (electricity and allocated hardware excluded); a rented GPU from $0.06/hr (NVIDIA RTX 3060 12GB at the direct Vast.ai price); or about $0.29 per 1M blended tokens via your own API key (provider-sourced rate dated 2026-08-24). Spanvero adds $0 markup on every path.
Spanvero calculates a baseline of about 9 GB of VRAM/unified memory for Qwen3-8B at Q4_K_M with 16.4K context. Actual fit depends on the exact build, runner, context, and overhead. After download it can run offline; marginal compute can be $0 on hardware you already own, excluding electricity and allocated hardware cost.
Spanvero calculates about 9 GB of VRAM/unified memory for Qwen3-8B at Q4_K_M; actual fit depends on the exact build, runner, context, and overhead. The smallest tracked consumer preset above that baseline is a used NVIDIA RTX 3080 10GB, but verify your configuration before downloading.
Browse: How to run Qwen3-8B locally · More Qwen (Alibaba) models · All models · Compare
Spanvero · What's new · Prices as of 2026-08-24. We're an honest advisor — $0 markup, your own accounts, we never resell compute. © 2026 Cynosure LLC.
A short email of real AI price moves, straight from the daily log — no hype. We're collecting the list now; the first issue goes out when it opens. Unsubscribe with one click.