A100 80GB for local LLMs — what fits in 80 GB & the real cost
Rent a A100 80GB from $1.39/hr ($0 markup, as of 2026-07-09). 487 of the 565 models in our catalog fit its 80 GB — plus the honest own-vs-rent-vs-API math, electricity and amortization included.
Three honest ways to run AI on a A100 80GB
Own it: buy the card and pay electricity + amortization — local is real money, not "$0".
Rent it: from $1.39/GPU-hr on your own vendor account (as of 2026-07-09, $0 markup — we never resell compute).
Skip it: run a comparable open model through your own API key, paying per million tokens (prices as of 2026-08-10).
Rent a A100 80GB — real vendor rates, $0 markup
RunPod: $1.39/GPU-hr — vendor's own price as of 2026-07-09; you pay them directly.
Crusoe: $2.00/GPU-hr — vendor's own price as of 2026-07-09; you pay them directly.
Links marked "referral link" are disclosed referrals — the vendor pays us a small cut and the price you pay is unchanged. Referrals never affect a rate or a ranking. See /affiliate-disclosure/.
What fits in 80 GB
487 of the 565 models in the Spanvero catalog fit on one A100 80GB at a sensible quant (context capped at 16k for the estimate). Most capable first:
gpt-oss-120b (OpenAI, 117B) — ~80 GB at Q4_K_M; own-key API alternative ≈ $0.10/1M tokens.
GLM 4.5 Air (zai-org, 110.5B) — ~76 GB at Q4_K_M; own-key API alternative ≈ $0.49/1M tokens.
GLM 4.5 Air (unsloth, 110.5B) — ~76 GB at Q4_K_M; own-key API alternative ≈ $0.98/1M tokens (size estimate).
Llama 4 Scout (17B-16E) (Meta, 109B) — ~77 GB at Q4_K_M; own-key API alternative ≈ $0.20/1M tokens (last-known).
sarvam 105b (sarvamai, 106B) — ~80 GB at Q4_K_M; own-key API alternative ≈ $0.95/1M tokens (size estimate).
Command R+ (08-2024) (Cohere, 104B) — ~75 GB at Q4_K_M; own-key API alternative ≈ $0.93/1M tokens (size estimate).
LLaDA2.1 flash (inclusionAI, 102.9B) — ~71 GB at Q4_K_M; own-key API alternative ≈ $0.92/1M tokens (size estimate).
Solar Open 100B (upstage, 102.7B) — ~71 GB at Q4_K_M; own-key API alternative ≈ $0.92/1M tokens (size estimate).
MiniMax M2.7 REAP 172B A10B NVFP4 GB10 (scottgl, 97.6B) — ~68 GB at Q4_K_M; own-key API alternative ≈ $0.88/1M tokens (size estimate).
gpt oss puzzle 88B (nvidia, 90.8B) — ~62 GB at Q4_K_M; own-key API alternative ≈ $0.83/1M tokens (size estimate).
Qwen3 Next 80B A3B Instruct (Qwen, 81.3B) — ~56 GB at Q4_K_M; own-key API alternative ≈ $0.60/1M tokens.
Qwen3 Next 80B A3B Thinking (Qwen, 81.3B) — ~56 GB at Q4_K_M; own-key API alternative ≈ $0.67/1M tokens.
Qwen3 Coder Next (Qwen, 79.7B) — ~55 GB at Q4_K_M; own-key API alternative ≈ $0.46/1M tokens.
Qwen2.5 72B Instruct (Alibaba, 72B) — ~54 GB at Q4_K_M; own-key API alternative ≈ $0.68/1M tokens (size estimate).
Meta Llama 3.1 70B Instruct quantized.w4a16 (RedHatAI, 70.6B) — ~54 GB at Q4_K_M; own-key API alternative ≈ $0.66/1M tokens (size estimate).
As an Amazon Associate we earn from qualifying purchases. Each link opens an Amazon search for the category — not an endorsement of any specific product — and Cynosure LLC may earn a commission at no extra cost to you.
A short email of real AI price moves, straight from the daily log — no hype. We're collecting the list now; the first issue goes out when it opens. Unsubscribe with one click.