L40S for local LLMs — what fits in 48 GB & the real cost
Rent a L40S from $0.99/hr ($0 markup, as of 2026-07-09). 456 of the 565 models in our catalog fit its 48 GB — plus the honest own-vs-rent-vs-API math, electricity and amortization included.
Three honest ways to run AI on a L40S
Own it: buy the card and pay electricity + amortization — local is real money, not "$0".
Rent it: from $0.99/GPU-hr on your own vendor account (as of 2026-07-09, $0 markup — we never resell compute).
Skip it: run a comparable open model through your own API key, paying per million tokens (prices as of 2026-08-10).
Rent a L40S — real vendor rates, $0 markup
RunPod: $0.99/GPU-hr — vendor's own price as of 2026-07-09; you pay them directly.
Crusoe: $1.50/GPU-hr — vendor's own price as of 2026-07-09; you pay them directly.
DigitalOcean: $1.57/GPU-hr — vendor's own price as of 2026-07-09; you pay them directly.
Links marked "referral link" are disclosed referrals — the vendor pays us a small cut and the price you pay is unchanged. Referrals never affect a rate or a ranking. See /affiliate-disclosure/.
What fits in 48 GB
456 of the 565 models in the Spanvero catalog fit on one L40S at a sensible quant (context capped at 16k for the estimate). Most capable first:
StableBeluga2 (petals-team, 69B) — ~48 GB at Q4_K_M; own-key API alternative ≈ $0.65/1M tokens (size estimate).
Laguna S 2.1 NVFP4 (poolside, 67.9B) — ~48 GB at Q4_K_M; own-key API alternative ≈ $0.64/1M tokens (size estimate).
NVIDIA Nemotron 3 Super 120B A12B NVFP4 (nvidia, 67.2B) — ~47 GB at Q4_K_M; own-key API alternative ≈ $0.64/1M tokens (size estimate).
Kimi Linear 48B A3B Instruct (moonshotai, 49.1B) — ~36 GB at Q4_K_M; own-key API alternative ≈ $0.49/1M tokens (size estimate).
Mixtral 8x7B Instruct (Mistral AI, 46.7B) — ~34 GB at Q4_K_M; own-key API alternative ≈ $0.24/1M tokens (last-known).
NVIDIA Nemotron Labs 3 Puzzle 75B A9B NVFP4 (nvidia, 44.5B) — ~33 GB at Q4_K_M; own-key API alternative ≈ $0.46/1M tokens (size estimate).
Phi 3.5 MoE instruct (microsoft, 41.9B) — ~31 GB at Q4_K_M; own-key API alternative ≈ $0.44/1M tokens (size estimate).
Karnak 40B v1.0 (Applied-Innovation-Center, 40.7B) — ~29 GB at Q4_K_M; own-key API alternative ≈ $0.43/1M tokens (size estimate).
Seed OSS 36B Instruct (ByteDance-Seed, 36.2B) — ~27 GB at Q4_K_M; own-key API alternative ≈ $0.39/1M tokens (size estimate).
Hermes 4.3 36B (NousResearch, 36.2B) — ~27 GB at Q4_K_M; own-key API alternative ≈ $0.39/1M tokens (size estimate).
automotive (flywheel-ai, 34.7B) — ~40 GB at Q4_K_M; own-key API alternative ≈ $0.38/1M tokens (size estimate).
Qwen AgentWorld 35B A3B (Qwen, 34.7B) — ~40 GB at Q4_K_M; own-key API alternative ≈ $0.38/1M tokens (size estimate).
Yi-1.5-34B-Chat (01.AI, 34.4B) — ~25 GB at Q4_K_M; own-key API alternative ≈ $0.38/1M tokens (size estimate).
Laguna XS.2 (poolside, 33.4B) — ~24 GB at Q4_K_M; own-key API alternative ≈ $0.15/1M tokens (last-known).
Laguna XS 2.1 NVFP4 (poolside, 33.4B) — ~24 GB at Q4_K_M; own-key API alternative ≈ $0.37/1M tokens (size estimate).
As an Amazon Associate we earn from qualifying purchases. Each link opens an Amazon search for the category — not an endorsement of any specific product — and Cynosure LLC may earn a commission at no extra cost to you.
A short email of real AI price moves, straight from the daily log — no hype. We're collecting the list now; the first issue goes out when it opens. Unsubscribe with one click.