H200 for local LLMs — what fits in 141 GB & the real cost
Rent a H200 from $3.44/hr ($0 markup, as of 2026-07-09). 507 of the 565 models in our catalog fit its 141 GB — plus the honest own-vs-rent-vs-API math, electricity and amortization included.
Three honest ways to run AI on a H200
Own it: buy the card and pay electricity + amortization — local is real money, not "$0".
Rent it: from $3.44/GPU-hr on your own vendor account (as of 2026-07-09, $0 markup — we never resell compute).
Skip it: run a comparable open model through your own API key, paying per million tokens (prices as of 2026-08-10).
Rent a H200 — real vendor rates, $0 markup
DigitalOcean: $3.44/GPU-hr — vendor's own price as of 2026-07-09; you pay them directly.
Hyperstack: $3.99/GPU-hr — vendor's own price as of 2026-07-09; you pay them directly.
Verda: $4.00/GPU-hr — vendor's own price as of 2026-07-09; you pay them directly.
Crusoe: $4.29/GPU-hr — vendor's own price as of 2026-07-09; you pay them directly.
RunPod: $4.39/GPU-hr (confidential-compute capable) — vendor's own price as of 2026-07-09; you pay them directly.
Nebius: $4.50/GPU-hr — vendor's own price as of 2026-07-09; you pay them directly.
Links marked "referral link" are disclosed referrals — the vendor pays us a small cut and the price you pay is unchanged. Referrals never affect a rate or a ranking. See /affiliate-disclosure/.
What fits in 141 GB
507 of the 565 models in the Spanvero catalog fit on one H200 at a sensible quant (context capped at 16k for the estimate). Most capable first:
DeepSeek V4 Flash NVFP4 (nvidia, 166.7B) — ~113 GB at Q4_K_M; own-key API alternative ≈ $1.43/1M tokens (size estimate).
DeepSeek V4 Flash DSpark (deepseek-ai, 165.3B) — ~112 GB at Q4_K_M; own-key API alternative ≈ $1.42/1M tokens (size estimate).
DeepSeek V4 Flash (deepseek-ai, 158.1B) — ~107 GB at Q4_K_M; own-key API alternative ≈ $0.21/1M tokens.
Solar Open2 250B Nota NVFP4 (nota-ai, 144.6B) — ~99 GB at Q4_K_M; own-key API alternative ≈ $1.26/1M tokens (size estimate).
Qwen3 235B A22B NVFP4 (nvidia, 132.8B) — ~91 GB at Q4_K_M; own-key API alternative ≈ $1.16/1M tokens (size estimate).
NVIDIA Nemotron 3 Super 120B A12B BF16 (nvidia, 123.6B) — ~85 GB at Q4_K_M; own-key API alternative ≈ $1.09/1M tokens (size estimate).
Mistral Large 2 (2407) (Mistral AI, 123B) — ~92 GB at Q4_K_M; own-key API alternative ≈ $1.08/1M tokens (size estimate).
gpt oss 120b (unsloth, 120.4B) — ~82 GB at Q4_K_M; own-key API alternative ≈ $1.06/1M tokens (size estimate).
Laguna S 2.1 (poolside, 117.6B) — ~81 GB at Q4_K_M; own-key API alternative ≈ $0.14/1M tokens.
gpt-oss-120b (OpenAI, 117B) — ~80 GB at Q4_K_M; own-key API alternative ≈ $0.10/1M tokens.
MiniMax M2.7 NVFP4 (nvidia, 116.3B) — ~81 GB at Q4_K_M; own-key API alternative ≈ $1.03/1M tokens (size estimate).
MiniMax M2.5 NVFP4 (nvidia, 116.3B) — ~81 GB at Q4_K_M; own-key API alternative ≈ $1.03/1M tokens (size estimate).
GLM 4.5 Air (zai-org, 110.5B) — ~76 GB at Q4_K_M; own-key API alternative ≈ $0.49/1M tokens.
GLM 4.5 Air (unsloth, 110.5B) — ~76 GB at Q4_K_M; own-key API alternative ≈ $0.98/1M tokens (size estimate).
Llama 4 Scout (17B-16E) (Meta, 109B) — ~77 GB at Q4_K_M; own-key API alternative ≈ $0.20/1M tokens (last-known).
As an Amazon Associate we earn from qualifying purchases. Each link opens an Amazon search for the category — not an endorsement of any specific product — and Cynosure LLC may earn a commission at no extra cost to you.
A short email of real AI price moves, straight from the daily log — no hype. We're collecting the list now; the first issue goes out when it opens. Unsubscribe with one click.