PRICING ── COMPARED
LLM API Pricing Table
11 vendors · 31 models · each row hand-verified on 2026-08-07
Three value tiers: maximum savings — DeepSeek v4-flash, Doubao seed-2.0-mini, GPT-5.4-nano; balanced workhorses — MiniMax-M3, the Gemini Flash line, Qwen3.7-plus, Claude Haiku 4.5; flagships — Claude Fable 5, GPT-5.6-sol, Kimi k3, GLM-5.2. The same workload spreads ~90× between the priciest and cheapest.
INTL · 18
International models (USD / 1M tokens)
A price range means input-length-tiered pricing. Cache prices are read/hit rates; separate write/storage fees are flagged in the notes.
| Vendor | Model | Input $/M | Output $/M | Cache read | Context | Notes |
|---|---|---|---|---|---|---|
| OpenAI | gpt-5.6-sol | $5 | $30 | $0.5 | — | Flagship tier; 50% off across the line via Batch |
| OpenAI | gpt-5.6-terra | $2 | $12 | $0.2 | — | Cut ~20% on Jul 30 (was $2.50/$15) |
| OpenAI | gpt-5.6-luna | $0.2 | $1.2 | $0.02 | — | Standard tier shown. Official Flex/Batch is $0.10/$0.60 and Fast is $0.40/$2.40; >272K long context steps up again (input doubles, output $1.80). Azure still bills the old $1/$6. Cut from $1/$6 on Jul 30 |
| OpenAI | gpt-5.5-pro | $30 | $180 | — | — | High-reasoning tier; no cache pricing |
| OpenAI | gpt-5.4-mini | $0.75 | $4.5 | $0.075 | — | |
| OpenAI | gpt-5.4-nano | $0.2 | $1.25 | $0.02 | — | Budget tier |
| Anthropic | Claude Fable 5 | $10 | $50 | $1 | 1M | 50% off via Batch; no long-context surcharge at 1M |
| Anthropic | Claude Opus 5 | $5 | $25 | $0.5 | 1M | Launched Jul 24; Batch at $2.50/$12.50 |
| Anthropic | Claude Opus 4.8 | $5 | $25 | $0.5 | 1M | Batch at $2.50/$12.50 |
| Anthropic | Claude Sonnet 5 | $2 | $10 | $0.2 | 1M | Intro price through Aug 31; $3/$15 afterwards |
| Anthropic | Claude Haiku 4.5 | $1 | $5 | $0.1 | 200K | Batch at $0.50/$2.50 |
| Gemini 3.1 Pro (Preview) | $2–4 | $12–18 | $0.2–0.4 | — | Tiered at 200K; cache storage billed separately; 50% off via Batch; free tier in AI Studio | |
| Gemini 3.6 Flash | $1.5 | $7.5 | $0.15 | — | Cache storage billed separately; free tier available | |
| Gemini 3.5 Flash-Lite | $0.3 | $2.5 | $0.03 | — | Cheapest GA tier; cache storage billed separately | |
| xAI | Grok 4.5 | $2–4 | $6–12 | $0.3–0.6 | 500K | Segmented by input length (<200K / ≥200K) |
| xAI | Grok 4.3 | $1.25–2.5 | $2.5–5 | $0.2–0.4 | 1M | Segmented by input length |
| Mistral | Mistral Medium 3.5 | $1.5 | $7.5 | — | — | No published cache pricing |
| Mistral | Mistral Large 3 | $0.5 | $1.5 | — | — | Marketing page still shows the old $2/$6; the API pricing table is authoritative |
CN · 13
Chinese models (CNY / 1M tokens)
CN-side pricing runs messier: input-length tiers, limited-time (or “permanent”) promos, snapshots excluded from discounts — see the change log for what moved.
| Vendor | Model | Input ¥/M | Output ¥/M | Cache read | Context | Notes |
|---|---|---|---|---|---|---|
| DeepSeek | deepseek-v4-flash | ¥1 | ¥2 | ¥0.02 | 1M | ⚠️ Official pricing page posted a price-rise notice on Aug 6 ("a substantial increase"); the numbers have not moved yet. R series folded into V4 thinking mode; legacy names deprecated Jul 24 |
| DeepSeek | deepseek-v4-pro | ¥3 | ¥6 | ¥0.025 | 1M | ⚠️ Also covered by the Aug 6 price-rise notice; plan and effective date pending official notice |
| Alibaba Qwen | qwen3.7-max | ¥6 | ¥18 | — | 1M | Limited-time 50% off (list ¥12/¥36); cache hits at 10% of input; Batch half price; snapshots excluded |
| Alibaba Qwen | qwen3.7-plus | ¥1.6–4.8 | ¥6.4–19.2 | — | 1M | Limited-time 20% off added Aug 6 (list ¥2/¥6 in, ¥8/¥24 out); tiered by input length (≤256K / >256K); cache hits at 10% of input |
| Alibaba Qwen | qwen3.8-max | ¥12 | ¥36 | — | 1M | New flagship at list price; has replaced qwen3.7-max as the promoted model (the latter's 50% off looks like a run-out) |
| Alibaba Qwen | qwen3.6-flash | ¥1.2–4.8 | ¥7.2–28.8 | — | 1M | Tiered by input length; cache hits at 10% of input |
| Moonshot Kimi | kimi-k3 | ¥20 | ¥100 | ¥2 | 1M | Always-reasoning; adjustable effort |
| Moonshot Kimi | kimi-k2.7-code | ¥6.5 | ¥27 | ¥1.3 | 256K | Coding model; highspeed variant costs double |
| Zhipu GLM | GLM-5.2 | ¥8 | ¥28 | ¥2 | 1M | Cache storage free for a limited time |
| Zhipu GLM | GLM-5 | ¥4–6 | ¥18–22 | ¥1–1.5 | — | Tiered by input length (<32K / ≥32K) |
| ByteDance Doubao | doubao-seed-2.1-pro | ¥6 | ¥30 | ¥1.2 | 256K | Cache storage billed at ¥0.017/1M/hour |
| ByteDance Doubao | doubao-seed-2.0-mini | ¥0.2–0.8 | ¥2–8 | ¥0.04–0.16 | 256K | Budget model; three input-length tiers. Output price corrected Aug 6 (our Jul 22 row mistakenly copied the audio-input column, ¥3/6/12) |
| MiniMax | MiniMax-M3 | ¥2.1 | ¥8.4 | ¥0.42 | 1M | "Permanent 50% off" (list ¥4.2 input); ≤512K tier; natively multimodal; priority tier 1.5× |
EXAMPLE · MONTHLY
One workload, eleven bills
300M input + 30M output tokens per month (~10M in / 1M out daily), no cache or Batch discounts, CNY at 1 USD ≈ ¥7.2 — computed directly from the tables above:
| Model | Monthly cost (USD) |
|---|---|
| deepseek-v4-flash | $50 |
| gpt-5.6-luna | $96 |
| gpt-5.4-nano | $98 |
| MiniMax-M3 | $123 |
| Gemini 3.5 Flash-Lite | $165 |
| qwen3.7-max | $325 |
| GLM-5.2 | $450 |
| Claude Haiku 4.5 | $450 |
| kimi-k3 | $1,250 |
| gpt-5.6-sol | $2,400 |
| Claude Fable 5 | $4,500 |
Roughly 90× between the extremes. Caching (hits at 2%–10% of input price) and Batch (~50% off at several vendors) stack to cut another order of magnitude.
CHANGELOG · 15
Price change log
A table like this is right today and wrong next month — so every change observed during re-verification is logged, append-only.
- 2026-08-07OpenAICorrecting and completing our own Aug 6 wording: "gpt-5.6-luna cut 80%" described only one tier. The official pricing page lists three service tiers — Flex/Batch $0.10/$0.60, Standard $0.20/$1.20, Fast $0.40/$2.40 (a 4× spread) — plus a >272K long-context step-up (Standard becomes $0.40/$1.80, input doubles). Our table shows Standard and now footnotes the rest.
- 2026-08-07OpenAIThe same luna carries eight different quotes across channels and tiers today, a 22× spread (checked the morning of Aug 7). OpenRouter currently stacks an extra 50% promo on all three OpenAI tiers (its endpoints API reports "discount": 0.5, exactly half the list price): flex bills $0.05/$0.30, standard $0.10/$0.60, priority $0.20/$1.20. Amazon Bedrock is $0.22/$1.32 with no discount. Azure, meanwhile, still charges $1.00/$6.00 and Azure EU $1.10/$6.60 with discount 0 — a week after the cut, still not synced. So: $0.05 via OpenRouter flex, $1.10 via Azure EU. Note the OpenRouter 50% is a promotion and can be withdrawn at any time.
- 2026-08-06DeepSeekThe official pricing page added a notice: DeepSeek plans a substantial across-the-board increase to API pricing, with the plan and effective date to follow. Note: the numbers on the page had NOT moved as of today (v4-flash still ¥1/¥2, v4-pro still ¥3/¥6); neither the size of the increase nor its start date has been published.
- 2026-08-06DeepSeekThe full story on "2x peak-hour pricing" (reconstructed from archived snapshots): the notice did appear on DeepSeek's official pricing page, but it was a forward-looking announcement throughout — and it was taken down today. Timeline: a Jul 31 10:20 (Beijing) snapshot still had a single footnote; later that same day, alongside the V4-Flash-0731 GA, a footnote was added saying DeepSeek "will soon adopt" peak/off-peak pricing at 2x during peak hours for all billable items, effective date pending official notice, with peak defined as 09:00–12:00 and 14:00–18:00 Beijing time daily. It stayed word-for-word identical through Aug 6, when it was replaced wholesale by the price-rise notice. Two details the internet gets wrong: the official text said "daily," not "weekdays" — the weekday/weekend carve-out was added by commentators; and only peak hours were ever defined, never off-peak. Corroboration: all six resellers checked today (Alibaba Model Studio, Volcano Ark, Tencent Cloud, Baidu Qianfan, SiliconFlow, OpenRouter) publish a single flat price with no time-of-day tier. Bottom line: don't budget at 2x — it never took effect.
- 2026-08-06OpenAIgpt-5.6-luna cut 80%: input $1→$0.20, output $6→$1.20, cache read $0.10→$0.02. In the same batch gpt-5.6-terra fell ~20%: $2.50/$15→$2.00/$12 (cache read $0.25→$0.20). Flagship gpt-5.6-sol unchanged. Announced Jul 30; confirmed on our Aug 6 re-verification.
- 2026-08-06AnthropicClaude Opus 5 launched Jul 24 at $5/$25 (cache read $0.50, Batch $2.50/$12.50) and is now in the table. Also: the official pricing page states that Claude 4.7 and later models (including Fable 5, Opus 4.8, Sonnet 5) use a new tokenizer that yields ~30% more tokens for the same text — this moves your bill more than the sticker price does when comparing across generations.
- 2026-08-06Alibaba Qwenqwen3.7-plus added a limited-time 20% off: input ¥2/¥6→¥1.6/¥4.8, output ¥8/¥24→¥6.4/¥19.2 (tiered by input length). The 50% off on qwen3.7-max is still running (no end date published on either the pricing or campaign page). Meanwhile the new flagship qwen3.8-max has launched at list price ¥12/¥36 and taken over the campaign page — the 3.7-max discount looks like a run-out; watching it closely.
- 2026-08-06ByteDance DoubaoCorrection (ours, not a vendor change): doubao-seed-2.0-mini output is ¥2/¥4/¥8 across the three tiers; our Jul 22 row mistakenly copied the vendor table's audio-input column (¥3/¥6/¥12). Fixed against the Aug 6 official page. Input ¥0.2/¥0.4/¥0.8 and cache-hit ¥0.04–0.16 were correct.
- 2026-08-06All vendorsFull re-verification (11 vendors, 31 models). Countdown: Claude Sonnet 5's intro price $2/$10 expires Aug 31, reverting to $3/$15 on Sep 1 (Batch $1/$5→$1.50/$7.50) — wording unchanged on the official page, 25 days out.
- 2026-07-24DeepSeekLegacy model names deepseek-chat and deepseek-reasoner officially deprecated; the R series is no longer sold separately — reasoning folded into V4's thinking mode.
- 2026-07-22Alibaba Qwenqwen3.7-max running a limited-time 50% off: input ¥12→¥6, output ¥36→¥18; snapshot models excluded.
- 2026-07-22MiniMaxMiniMax-M3 marked "permanently 50% off": input ¥4.2→¥2.1, output ¥8.4 (≤512K tier).
- 2026-07-22AnthropicClaude Sonnet 5 intro price $2/$10 through Aug 31; $3/$15 afterwards.
- 2026-07-22MistralTwo conflicting prices on the official site: the marketing page still lists Large at $2/$6 while the API pricing table says $0.5/$1.5. The API table is authoritative.
- 2026-07-22All vendorsBaseline established: 29 models across 11 vendors, each row hand-checked against the official pricing page.
FAQ
Frequently asked
Which LLM API is cheapest in 2026?
Checked against official pricing pages: DeepSeek v4-flash at ¥1 in / ¥2 out per 1M tokens (cache hits just ¥0.02) is the cheapest mainstream tier, followed by ByteDance Doubao seed-2.0-mini (from ¥0.2) and OpenAI gpt-5.4-nano ($0.20/$1.25).
How much can the same workload differ across models?
For 300M input + 30M output tokens per month (no cache/Batch discounts): the priciest flagship runs about $4,500/month vs about $50/month at the low end — roughly 90×. Decide whether the task truly needs a flagship first.
How much do caching and Batch save?
Cache-hit rates typically run 2%–10% of the input price (DeepSeek as low as ¥0.02/1M); OpenAI, Anthropic, Google and Qwen offer ~50% off via Batch. The two stack, cutting repeat-heavy workloads by another order of magnitude.
Sources are the vendors’ official pricing pages (verified 2026-08-07): OpenAI · Anthropic · Google · xAI · Mistral · DeepSeek · Alibaba Model Studio · Moonshot Kimi · Zhipu · ByteDance Doubao · MiniMax. Prices move often — treat the official page as live truth; corrections welcome.