PRICING ── COMPARED
LLM API Pricing Table
11 vendors · 29 models · each row hand-verified on 2026-07-22
Three value tiers: maximum savings — DeepSeek v4-flash, Doubao seed-2.0-mini, GPT-5.4-nano; balanced workhorses — MiniMax-M3, the Gemini Flash line, Qwen3.7-plus, Claude Haiku 4.5; flagships — Claude Fable 5, GPT-5.6-sol, Kimi k3, GLM-5.2. The same workload spreads ~90× between the priciest and cheapest.
INTL · 17
International models (USD / 1M tokens)
A price range means input-length-tiered pricing. Cache prices are read/hit rates; separate write/storage fees are flagged in the notes.
| Vendor | Model | Input $/M | Output $/M | Cache read | Context | Notes |
|---|---|---|---|---|---|---|
| OpenAI | gpt-5.6-sol | $5 | $30 | $0.5 | — | Flagship tier; 50% off across the line via Batch |
| OpenAI | gpt-5.6-terra | $2.5 | $15 | $0.25 | — | |
| OpenAI | gpt-5.6-luna | $1 | $6 | $0.1 | — | |
| OpenAI | gpt-5.5-pro | $30 | $180 | — | — | High-reasoning tier; no cache pricing |
| OpenAI | gpt-5.4-mini | $0.75 | $4.5 | $0.075 | — | |
| OpenAI | gpt-5.4-nano | $0.2 | $1.25 | $0.02 | — | Budget tier |
| Anthropic | Claude Fable 5 | $10 | $50 | $1 | 1M | 50% off via Batch; no long-context surcharge at 1M |
| Anthropic | Claude Opus 4.8 | $5 | $25 | $0.5 | 1M | Batch at $2.50/$12.50 |
| Anthropic | Claude Sonnet 5 | $2 | $10 | $0.2 | 1M | Intro price through Aug 31; $3/$15 afterwards |
| Anthropic | Claude Haiku 4.5 | $1 | $5 | $0.1 | 200K | Batch at $0.50/$2.50 |
| Gemini 3.1 Pro (Preview) | $2–4 | $12–18 | $0.2–0.4 | — | Tiered at 200K; cache storage billed separately; 50% off via Batch; free tier in AI Studio | |
| Gemini 3.6 Flash | $1.5 | $7.5 | $0.15 | — | Cache storage billed separately; free tier available | |
| Gemini 3.5 Flash-Lite | $0.3 | $2.5 | $0.03 | — | Cheapest GA tier; cache storage billed separately | |
| xAI | Grok 4.5 | $2–4 | $6–12 | $0.3–0.6 | 500K | Segmented by input length (<200K / ≥200K) |
| xAI | Grok 4.3 | $1.25–2.5 | $2.5–5 | $0.2–0.4 | 1M | Segmented by input length |
| Mistral | Mistral Medium 3.5 | $1.5 | $7.5 | — | — | No published cache pricing |
| Mistral | Mistral Large 3 | $0.5 | $1.5 | — | — | Marketing page still shows the old $2/$6; the API pricing table is authoritative |
CN · 12
Chinese models (CNY / 1M tokens)
CN-side pricing runs messier: input-length tiers, limited-time (or “permanent”) promos, snapshots excluded from discounts — see the change log for what moved.
| Vendor | Model | Input ¥/M | Output ¥/M | Cache read | Context | Notes |
|---|---|---|---|---|---|---|
| DeepSeek | deepseek-v4-flash | ¥1 | ¥2 | ¥0.02 | 1M | R series folded into V4 thinking mode; legacy names deepseek-chat/reasoner deprecated Jul 24 |
| DeepSeek | deepseek-v4-pro | ¥3 | ¥6 | ¥0.025 | 1M | |
| Alibaba Qwen | qwen3.7-max | ¥6 | ¥18 | — | 1M | Limited-time 50% off (list ¥12/¥36); cache hits at 10% of input; Batch half price; snapshots excluded |
| Alibaba Qwen | qwen3.7-plus | ¥2–6 | ¥8–24 | — | 1M | Tiered by input length (≤256K / >256K); cache hits at 10% of input |
| Alibaba Qwen | qwen3.6-flash | ¥1.2–4.8 | ¥7.2–28.8 | — | 1M | Tiered by input length; cache hits at 10% of input |
| Moonshot Kimi | kimi-k3 | ¥20 | ¥100 | ¥2 | 1M | Always-reasoning; adjustable effort |
| Moonshot Kimi | kimi-k2.7-code | ¥6.5 | ¥27 | ¥1.3 | 256K | Coding model; highspeed variant costs double |
| Zhipu GLM | GLM-5.2 | ¥8 | ¥28 | ¥2 | 1M | Cache storage free for a limited time |
| Zhipu GLM | GLM-5 | ¥4–6 | ¥18–22 | ¥1–1.5 | — | Tiered by input length (<32K / ≥32K) |
| ByteDance Doubao | doubao-seed-2.1-pro | ¥6 | ¥30 | ¥1.2 | 256K | Cache storage billed at ¥0.017/1M/hour |
| ByteDance Doubao | doubao-seed-2.0-mini | ¥0.2–0.8 | ¥3–12 | ¥0.04–0.16 | 256K | Budget model; three input-length tiers |
| MiniMax | MiniMax-M3 | ¥2.1 | ¥8.4 | ¥0.42 | 1M | "Permanent 50% off" (list ¥4.2 input); ≤512K tier; natively multimodal; priority tier 1.5× |
EXAMPLE · MONTHLY
One workload, eleven bills
300M input + 30M output tokens per month (~10M in / 1M out daily), no cache or Batch discounts, CNY at 1 USD ≈ ¥7.2 — computed directly from the tables above:
| Model | Monthly cost (USD) |
|---|---|
| deepseek-v4-flash | $50 |
| gpt-5.4-nano | $98 |
| MiniMax-M3 | $123 |
| Gemini 3.5 Flash-Lite | $165 |
| qwen3.7-max | $325 |
| GLM-5.2 | $450 |
| Claude Haiku 4.5 | $450 |
| kimi-k3 | $1,250 |
| gpt-5.6-sol | $2,400 |
| Claude Fable 5 | $4,500 |
Roughly 90× between the extremes. Caching (hits at 2%–10% of input price) and Batch (~50% off at several vendors) stack to cut another order of magnitude.
CHANGELOG · 6
Price change log
A table like this is right today and wrong next month — so every change observed during re-verification is logged, append-only.
- 2026-07-24DeepSeekLegacy model names deepseek-chat and deepseek-reasoner officially deprecated; the R series is no longer sold separately — reasoning folded into V4's thinking mode.
- 2026-07-22Alibaba Qwenqwen3.7-max running a limited-time 50% off: input ¥12→¥6, output ¥36→¥18; snapshot models excluded.
- 2026-07-22MiniMaxMiniMax-M3 marked "permanently 50% off": input ¥4.2→¥2.1, output ¥8.4 (≤512K tier).
- 2026-07-22AnthropicClaude Sonnet 5 intro price $2/$10 through Aug 31; $3/$15 afterwards.
- 2026-07-22MistralTwo conflicting prices on the official site: the marketing page still lists Large at $2/$6 while the API pricing table says $0.5/$1.5. The API table is authoritative.
- 2026-07-22All vendorsBaseline established: 29 models across 11 vendors, each row hand-checked against the official pricing page.
FAQ
Frequently asked
Which LLM API is cheapest in 2026?
Checked against official pricing pages: DeepSeek v4-flash at ¥1 in / ¥2 out per 1M tokens (cache hits just ¥0.02) is the cheapest mainstream tier, followed by ByteDance Doubao seed-2.0-mini (from ¥0.2) and OpenAI gpt-5.4-nano ($0.20/$1.25).
How much can the same workload differ across models?
For 300M input + 30M output tokens per month (no cache/Batch discounts): the priciest flagship runs about $4,500/month vs about $50/month at the low end — roughly 90×. Decide whether the task truly needs a flagship first.
How much do caching and Batch save?
Cache-hit rates typically run 2%–10% of the input price (DeepSeek as low as ¥0.02/1M); OpenAI, Anthropic, Google and Qwen offer ~50% off via Batch. The two stack, cutting repeat-heavy workloads by another order of magnitude.
Sources are the vendors’ official pricing pages (verified 2026-07-22): OpenAI · Anthropic · Google · xAI · Mistral · DeepSeek · Alibaba Model Studio · Moonshot Kimi · Zhipu · ByteDance Doubao · MiniMax. Prices move often — treat the official page as live truth; corrections welcome.