SpanveroHow it worksFind a modelBest by useCompare modelsPricing

State of Inference Costs — July 2026

Published July 9, 2026 · data June 20, 2026 → August 23, 2026 (65 daily snapshots) · live edition — extends with each daily snapshot until the next report ships

What running AI actually cost this month — computed straight from Spanvero's append-only daily price record, the same data you can download. No estimates, no smoothing: every figure carries its date.

Daily snapshots
65
Jun 20 → Aug 23
Prices logged
33,729
rows appended to the log
Models tracked
417
live market, latest day
Real price moves
60
37 cuts · 23 rises

What the data shows

Between June 20, 2026 and August 23, 2026, Spanvero logged 33,729 prices across 65 daily snapshots: the full live market listing (417 priced models on the latest day), 85 GPU rental rates, and the dated per-model prices Spanvero publishes. The log is append-only — a day, once written, is never revised.

The floor held. The cheapest paid model on the market cost $0.02 per 1M blended tokens on June 20, 2026 (inclusionAI: Ling-2.6-flash) and $0.02/1M on August 23, 2026 (still inclusionAI: Ling-2.6-flash) — no meaningful move at the bottom of the market.

Above the floor, prices did move: 60 real day-over-day changes cleared our 0.5% noise threshold — 37 cuts against 23 rises. The largest single move was Thinking Machines: Inkling (batch), rising from $1.2625 to $2.525/1M on August 19, 2026 (+100%); the largest cut was DeepSeek: DeepSeek V4 Pro 0423, from $2.4 to $0.646/1M on August 22, 2026 (−73%).

On the rental side, the cheapest tracked GPU went from $0.26/hr on June 20, 2026 (NVIDIA RTX 3090 24GB) to $0.06/hr on August 23, 2026 (NVIDIA RTX 3060 12GB). Read that with care: our GPU coverage grew from 10 to 85 SKUs over the window, and the new floor comes from a card that entered the record mid-window — so much of that drop is wider tracking, not vendor repricing. Like-for-like — only the SKUs tracked for the whole window — the floor held at $0.26/hr (NVIDIA RTX 3090 24GB).

Finally, the market ended the window 81 models larger than it began (entries minus exits), and truly-free endpoints ($0 in, $0 out) went from 27 to 22.

The cheapest way to run AI, start vs end

Biggest market moves

Single-day changes on the live market listing, largest first — 60 in this window. The full feed lives on the price record.

NameBeforeAfterChangeDate
Thinking Machines: Inkling (batch)$1.2625/1M$2.525/1M+100%Aug 19
Z.ai: GLM 5.2 (batch)$1.45/1M$2.9/1M+100%Aug 19
MoonshotAI: Kimi K2.7 Code (batch)$1.2375/1M$2.475/1M+100%Aug 19
NVIDIA: Nemotron 3 Ultra (batch)$1.05/1M$2.1/1M+100%Aug 19
MiniMax: MiniMax M3 (batch)$0.375/1M$0.75/1M+100%Aug 19
OpenAI: GPT-5.6 Luna Pro$0.35/1M$0.7/1M+100%Aug 18
OpenAI: GPT-5.6 Luna$0.35/1M$0.7/1M+100%Aug 18
OpenAI: GPT-5.6 Terra Pro$3.5/1M$7/1M+100%Aug 18
OpenAI: GPT-5.6 Terra$3.5/1M$7/1M+100%Aug 18
DeepSeek: DeepSeek V3.1$0.6/1M$1.1/1M+83%Aug 22

Published-model movers

Window-start vs window-end change in the dated price Spanvero publishes per model — each links to that model's live cost page. 31 of 70 tracked models moved.

NameBeforeAfterChangeLast changed
Qwen3 30B A3B Thinking 2507 (entered Jun 23)$0.24/1M$1.3/1M+442%Aug 4
DeepSeek V4 Pro$0.655/1M$2.64/1M+303%Aug 18
Hy3 preview (entered Jun 23)$0.1365/1M$0.39/1M+186%Aug 18
Llama 3.1 8B Instruct$0.025/1M$0.065/1M+160%Jul 21
Gemma 3 27B (entered Jun 23)$0.12/1M$0.265/1M+121%Jul 28
Qwen2.5 7B Instruct$0.07/1M$0.15/1M+114%Aug 4
DeepSeek V4 Flash 0731 (entered Aug 4)$0.135/1M$0.21/1M+56%Aug 18
Qwen3 Next 80B A3B Thinking (entered Jul 7)$0.4388/1M$0.675/1M+54%Aug 4
Hy3 (entered Jul 21)$0.5/1M$0.33/1M−34%Jul 28
DeepSeek-V3$0.5/1M$0.6432/1M+29%Aug 4

GPU rental movers

Hourly rates that changed over the window, per tracked SKU — each links to that card's rate page.

NameBeforeAfterChangeLast changed
NVIDIA A100 40GB$1.10/hr$1.99/hr+81%Jul 3
NVIDIA H200 141GB$3.49/hr$4.39/hr+26%Jul 3
NVIDIA RTX A5000 24GB$0.36/hr$0.27/hr−25%Jul 3
NVIDIA H100 80GB SXM$2.69/hr$3.29/hr+22%Jul 3
NVIDIA L40S 48GB$0.86/hr$0.99/hr+15%Jul 3
NVIDIA L4 24GB$0.43/hr$0.39/hr−9.3%Jul 3

Sources, labeled: the market floor and single-day moves come from the live OpenRouter listing (the full market); published-model rows are the dated prices Spanvero publishes; GPU rates are the vendors' own posted prices. We sell no compute and mark nothing up. How we stay honest →

Cite this report

Free to quote with a link. The underlying record: history.json · history.csv · prices.json · gpus.json · the full feed →

Spanvero, “State of Inference Costs — July 2026” (data 2026-06-20 → 2026-08-23) — spanvero.com/reports/inference-costs-2026-07/

Want the Outcome Lab run against your own evidence? Flat-fee Outcome Cost audits are $49 / $249 / $499. How the audits work →

All reports → · The daily price record · Outcome Lab →

Data June 20, 2026 → August 23, 2026; published July 9, 2026. $0 markup, your own accounts, we never resell compute. © 2026 Cynosure LLC.

The weekly price index

A short email of real AI price moves, straight from the daily log — no hype. We're collecting the list now; the first issue goes out when it opens. Unsubscribe with one click.