SpanveroCheck my hardwareCompare GPUsModels by VRAMCompare costsHow it works

What LLMs fit in 24 GB of VRAM?

426 open models are calculated to use no more than 24 GB of VRAM at the stated default quant and context — a budget associated with a high-end card like an RTX 3090 / 4090. Largest first; verify the selected build and runner on your exact machine.

How to choose an LLM for a 24 GB GPU

The 426-model list below is a calculated fit inventory, not a quality ranking. Start with fit, verify your exact card and runner, then compare the cost of owning, renting, or using an API.

  1. Test 24 GB against the full catalog.
  2. Check RTX 3090 or RTX 4090 fit and cost details.
  3. Compare local, rented GPU, and API routes.

Memory use changes with quantization, context length, runtime overhead, and the exact model build. A calculated fit is a starting point—not a hardware guarantee.

Our shortlist: Best LLMs for 24 GB VRAM →

Other GPU budgets

8 GB of VRAM · 16 GB of VRAM · 48 GB of VRAM · All models

Open the free LLM cost calculator → · Rate table dated 2026-08-24. $0 markup, your own accounts, we never resell compute. © 2026 Cynosure LLC.

The weekly price index

A short email of real AI price moves, straight from the daily log — no hype. We're collecting the list now; the first issue goes out when it opens. Unsubscribe with one click.