RTX A5000 for local LLMs — what fits in 24 GB & the real cost
Rent a RTX A5000 from $0.27/hr ($0 markup, as of 2026-07-09). 402 of the 565 models in our catalog fit its 24 GB — plus the honest own-vs-rent-vs-API math, electricity and amortization included.
Three honest ways to run AI on a RTX A5000
Own it: buy the card and pay electricity + amortization — local is real money, not "$0".
Rent it: from $0.27/GPU-hr on your own vendor account (as of 2026-07-09, $0 markup — we never resell compute).
Skip it: run a comparable open model through your own API key, paying per million tokens (prices as of 2026-08-10).
Rent a RTX A5000 — real vendor rates, $0 markup
RunPod: $0.27/GPU-hr — vendor's own price as of 2026-07-09; you pay them directly.
Links marked "referral link" are disclosed referrals — the vendor pays us a small cut and the price you pay is unchanged. Referrals never affect a rate or a ranking. See /affiliate-disclosure/.
What fits in 24 GB
402 of the 565 models in the Spanvero catalog fit on one RTX A5000 at a sensible quant (context capped at 16k for the estimate). Most capable first:
Laguna XS.2 (poolside, 33.4B) — ~24 GB at Q4_K_M; own-key API alternative ≈ $0.15/1M tokens (last-known).
Laguna XS 2.1 NVFP4 (poolside, 33.4B) — ~24 GB at Q4_K_M; own-key API alternative ≈ $0.37/1M tokens (size estimate).
Laguna XS 2.1 (poolside, 33.4B) — ~24 GB at Q4_K_M; own-key API alternative ≈ $0.09/1M tokens.
sarvam 30b (sarvamai, 32.2B) — ~22 GB at Q4_K_M; own-key API alternative ≈ $0.36/1M tokens (size estimate).
llm jp 4 32b a3b thinking (llm-jp, 32.1B) — ~23 GB at Q4_K_M; own-key API alternative ≈ $0.36/1M tokens (size estimate).
NVIDIA Nemotron 3 Nano 30B A3B BF16 (nvidia, 31.6B) — ~22 GB at Q4_K_M; own-key API alternative ≈ $0.35/1M tokens (size estimate).
Nemotron Cascade 2 30B A3B (nvidia, 31.6B) — ~22 GB at Q4_K_M; own-key API alternative ≈ $0.35/1M tokens (size estimate).
Qwen3 30B A3B (Qwen, 30.5B) — ~22 GB at Q4_K_M; own-key API alternative ≈ $0.31/1M tokens.
Qwen3 Coder 30B A3B Instruct (Qwen, 30.5B) — ~22 GB at Q4_K_M; own-key API alternative ≈ $0.17/1M tokens.
Qwen3 30B A3B Instruct 2507 (Qwen, 30.5B) — ~22 GB at Q4_K_M; own-key API alternative ≈ $0.12/1M tokens.
Qwen3 30B A3B Thinking 2507 (Qwen, 30.5B) — ~22 GB at Q4_K_M; own-key API alternative ≈ $1.30/1M tokens.
Tongyi DeepResearch 30B A3B (Alibaba-NLP, 30.5B) — ~22 GB at Q4_K_M; own-key API alternative ≈ $0.34/1M tokens (size estimate).
Qwen3 30B A3B Base (Qwen, 30.5B) — ~22 GB at Q4_K_M; own-key API alternative ≈ $0.34/1M tokens (size estimate).
lynx instruct 30b (bineric, 30.5B) — ~22 GB at Q4_K_M; own-key API alternative ≈ $0.34/1M tokens (size estimate).
North Mini Code 1.0 (CohereLabs, 30.5B) — ~22 GB at Q4_K_M; own-key API alternative ≈ $0.34/1M tokens (size estimate).
As an Amazon Associate we earn from qualifying purchases. Each link opens an Amazon search for the category — not an endorsement of any specific product — and Cynosure LLC may earn a commission at no extra cost to you.
A short email of real AI price moves, straight from the daily log — no hype. We're collecting the list now; the first issue goes out when it opens. Unsubscribe with one click.