Can Used AMD Radeon RX 7800 XT 16GB run Llama, Qwen & DeepSeek? 326 models that fit
326 of the 493 models in the Spanvero catalog fit Used AMD Radeon RX 7800 XT 16GB's 16 GB VRAM (at a sensible quant, 16k context). For each: run it locally ($0 compute + electricity), rent an equivalent GPU ($0 markup, as of 2026-07-09), or pay per-token via your own API key (as of 2026-08-17).
Three honest ways to run each model on Used AMD Radeon RX 7800 XT 16GB
Run it locally: $0 in compute — you pay only electricity (~263 W under load on this card). Local is real money, never a fake "$0".
Rent an equivalent GPU: from a $0-markup vendor rate (as of 2026-07-09) — you rent on your own account and pay the vendor directly; we never resell compute.
Skip the box: run the same model through your own API key, paying per million tokens (prices as of 2026-08-17).
What fits Used AMD Radeon RX 7800 XT 16GB (16 GB VRAM)
326 of the 493 notable models in the Spanvero catalog fit Used AMD Radeon RX 7800 XT 16GB at a sensible quant (context capped at 16k for the estimate). Most capable first:
ERNIE 4.5 21B A3B Thinking (baidu, 21.8B) — needs ~16 GB at Q4_K_M: run it locally for $0 compute + ~$0.6135/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.27/1M via your own API key (size estimate).
gpt oss safeguard 20b (openai, 21.5B) — needs ~16 GB at Q4_K_M: run it locally for $0 compute + ~$0.6068/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.19/1M via your own API key.
gpt-oss-20b (OpenAI, 21B) — needs ~15 GB at Q4_K_M: run it locally for $0 compute + ~$0.5954/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.08/1M via your own API key.
gpt oss 20b BF16 (unsloth, 20.9B) — needs ~15 GB at Q4_K_M: run it locally for $0 compute + ~$0.5932/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.27/1M via your own API key (size estimate).
NVIDIA Nemotron 3 Nano 30B A3B NVFP4 (nvidia, 18.2B) — needs ~13 GB at Q4_K_M: run it locally for $0 compute + ~$0.531/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.25/1M via your own API key (size estimate).
NVIDIA Nemotron 3.5 Lightning 30B A3B NVFP4 (nvidia, 17.8B) — needs ~13 GB at Q4_K_M: run it locally for $0 compute + ~$0.5217/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.24/1M via your own API key (size estimate).
Qwen3 30B A3B NVFP4 (RedHatAI, 17.5B) — needs ~13 GB at Q4_K_M: run it locally for $0 compute + ~$0.5146/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.24/1M via your own API key (size estimate).
Qwen3 32B NVFP4 (nvidia, 17.2B) — needs ~15 GB at Q4_K_M: run it locally for $0 compute + ~$0.5076/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.24/1M via your own API key (size estimate).
deepseek moe 16b base (deepseek-ai, 16.4B) — needs ~12 GB at Q4_K_M: run it locally for $0 compute + ~$0.4886/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.23/1M via your own API key (size estimate).
deepseek moe 16b chat (deepseek-ai, 16.4B) — needs ~12 GB at Q4_K_M: run it locally for $0 compute + ~$0.4886/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.23/1M via your own API key (size estimate).
LLaDA2.0 mini (inclusionAI, 16.3B) — needs ~12 GB at Q4_K_M: run it locally for $0 compute + ~$0.4862/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.23/1M via your own API key (size estimate).
LLaDA2.1 mini (inclusionAI, 16.3B) — needs ~12 GB at Q4_K_M: run it locally for $0 compute + ~$0.4862/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.23/1M via your own API key (size estimate).
Ling mini 2.0 (inclusionAI, 16.3B) — needs ~12 GB at Q4_K_M: run it locally for $0 compute + ~$0.4862/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.23/1M via your own API key (size estimate).
Moonlight 16B A3B Instruct (moonshotai, 16B) — needs ~13 GB at Q4_K_M: run it locally for $0 compute + ~$0.479/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.23/1M via your own API key (size estimate).
Moonlight 16B A3B (moonshotai, 16B) — needs ~13 GB at Q4_K_M: run it locally for $0 compute + ~$0.479/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.23/1M via your own API key (size estimate).
DeepSeek-Coder-V2-Lite Instruct (DeepSeek, 15.7B) — needs ~11 GB at Q4_K_M: run it locally for $0 compute + ~$0.4718/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.23/1M via your own API key (size estimate).
DeepSeek V2 Lite Chat (deepseek-ai, 15.7B) — needs ~15 GB at Q4_K_M: run it locally for $0 compute + ~$0.4718/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.23/1M via your own API key (size estimate).
DeepSeek V2 Lite (deepseek-ai, 15.7B) — needs ~15 GB at Q4_K_M: run it locally for $0 compute + ~$0.4718/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.23/1M via your own API key (size estimate).
Qwen3 30B A3B NVFP4 (nvidia, 15.6B) — needs ~12 GB at Q4_K_M: run it locally for $0 compute + ~$0.4694/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.22/1M via your own API key (size estimate).
Qwen2.5 Coder 14B Instruct (Qwen, 14.8B) — needs ~14 GB at Q4_K_M: run it locally for $0 compute + ~$0.4501/1M in electricity, rent 2× NVIDIA RTX 3060 12GB from $0.12/hr ($0 markup), or ~$0.22/1M via your own API key (size estimate).
The honest cost of owning Used AMD Radeon RX 7800 XT 16GB
Street price $460.00 (as of 2026-07-06; videocardprices.com eBay used tracker, retrieved 2026-07-06 (~$460)) — amortized over 3 years that's ~$0.4201/day whether or not you're generating.
Electricity: ~263 W under sustained inference at $0.1883/kWh (EIA Electric Power Monthly Table 5.6.A — U.S. residential average, Apr 2026, as of 2026-07-02) — the per-1M-token figures above already include this at each model's speed.
Straight talk: for the small models a 16 GB VRAM box runs, hosted APIs are often cheaper per token. Own local for privacy, offline use, and unlimited runs — not to save money on tokens.
Too big for Used AMD Radeon RX 7800 XT 16GB — rent or use an API instead
These need more than the 16 GB VRAM on this card. Closest first — you can still run them on a rented GPU ($0 markup) or via your own API key:
solar pro preview instruct (upstage, 22.1B) — needs ~17 GB; rent 2× NVIDIA RTX 3060 12GB from $0.12/hr, or ~$0.28/1M via your own API key (size estimate).
gpt neox 20b (EleutherAI, 20.7B) — needs ~17 GB; rent 2× NVIDIA RTX 3060 12GB from $0.12/hr, or ~$0.27/1M via your own API key (size estimate).
Gemma 4 26B A4B NVFP4 (nvidia, 14.4B) — needs ~17 GB; rent 2× NVIDIA RTX 3060 12GB from $0.12/hr, or ~$0.22/1M via your own API key (size estimate).
diffusiongemma 26B A4B it NVFP4 (nvidia, 14.4B) — needs ~17 GB; rent 2× NVIDIA RTX 3060 12GB from $0.12/hr, or ~$0.22/1M via your own API key (size estimate).
LFM2 24B A2B (LiquidAI, 23.8B) — needs ~18 GB; rent 2× NVIDIA RTX 3060 12GB from $0.12/hr, or ~$0.29/1M via your own API key (size estimate).
Gemma 4 26B A4B it NVFP4 (bg-digitalservices, 15.1B) — needs ~18 GB; rent 2× NVIDIA RTX 3060 12GB from $0.12/hr, or ~$0.22/1M via your own API key (size estimate).
A short email of real AI price moves, straight from the daily log — no hype. We're collecting the list now; the first issue goes out when it opens. Unsubscribe with one click.