PRICING ── COMPARED

LLM API Pricing Table

11 vendors · 29 models · each row hand-verified on 2026-07-22

Three value tiers: maximum savings — DeepSeek v4-flash, Doubao seed-2.0-mini, GPT-5.4-nano; balanced workhorses — MiniMax-M3, the Gemini Flash line, Qwen3.7-plus, Claude Haiku 4.5; flagships — Claude Fable 5, GPT-5.6-sol, Kimi k3, GLM-5.2. The same workload spreads ~90× between the priciest and cheapest.

Subscribe to price-change alerts← Back to the intel desk

INTL · 17

International models (USD / 1M tokens)

A price range means input-length-tiered pricing. Cache prices are read/hit rates; separate write/storage fees are flagged in the notes.

VendorModelInput $/MOutput $/MCache readContextNotes
OpenAIgpt-5.6-sol$5$30$0.5Flagship tier; 50% off across the line via Batch
OpenAIgpt-5.6-terra$2.5$15$0.25
OpenAIgpt-5.6-luna$1$6$0.1
OpenAIgpt-5.5-pro$30$180High-reasoning tier; no cache pricing
OpenAIgpt-5.4-mini$0.75$4.5$0.075
OpenAIgpt-5.4-nano$0.2$1.25$0.02Budget tier
AnthropicClaude Fable 5$10$50$11M50% off via Batch; no long-context surcharge at 1M
AnthropicClaude Opus 4.8$5$25$0.51MBatch at $2.50/$12.50
AnthropicClaude Sonnet 5$2$10$0.21MIntro price through Aug 31; $3/$15 afterwards
AnthropicClaude Haiku 4.5$1$5$0.1200KBatch at $0.50/$2.50
GoogleGemini 3.1 Pro (Preview)$2–4$12–18$0.2–0.4Tiered at 200K; cache storage billed separately; 50% off via Batch; free tier in AI Studio
GoogleGemini 3.6 Flash$1.5$7.5$0.15Cache storage billed separately; free tier available
GoogleGemini 3.5 Flash-Lite$0.3$2.5$0.03Cheapest GA tier; cache storage billed separately
xAIGrok 4.5$2–4$6–12$0.3–0.6500KSegmented by input length (<200K / ≥200K)
xAIGrok 4.3$1.25–2.5$2.5–5$0.2–0.41MSegmented by input length
MistralMistral Medium 3.5$1.5$7.5No published cache pricing
MistralMistral Large 3$0.5$1.5Marketing page still shows the old $2/$6; the API pricing table is authoritative

CN · 12

Chinese models (CNY / 1M tokens)

CN-side pricing runs messier: input-length tiers, limited-time (or “permanent”) promos, snapshots excluded from discounts — see the change log for what moved.

VendorModelInput ¥/MOutput ¥/MCache readContextNotes
DeepSeekdeepseek-v4-flash¥1¥2¥0.021MR series folded into V4 thinking mode; legacy names deepseek-chat/reasoner deprecated Jul 24
DeepSeekdeepseek-v4-pro¥3¥6¥0.0251M
Alibaba Qwenqwen3.7-max¥6¥181MLimited-time 50% off (list ¥12/¥36); cache hits at 10% of input; Batch half price; snapshots excluded
Alibaba Qwenqwen3.7-plus¥2–6¥8–241MTiered by input length (≤256K / >256K); cache hits at 10% of input
Alibaba Qwenqwen3.6-flash¥1.2–4.8¥7.2–28.81MTiered by input length; cache hits at 10% of input
Moonshot Kimikimi-k3¥20¥100¥21MAlways-reasoning; adjustable effort
Moonshot Kimikimi-k2.7-code¥6.5¥27¥1.3256KCoding model; highspeed variant costs double
Zhipu GLMGLM-5.2¥8¥28¥21MCache storage free for a limited time
Zhipu GLMGLM-5¥4–6¥18–22¥1–1.5Tiered by input length (<32K / ≥32K)
ByteDance Doubaodoubao-seed-2.1-pro¥6¥30¥1.2256KCache storage billed at ¥0.017/1M/hour
ByteDance Doubaodoubao-seed-2.0-mini¥0.2–0.8¥3–12¥0.04–0.16256KBudget model; three input-length tiers
MiniMaxMiniMax-M3¥2.1¥8.4¥0.421M"Permanent 50% off" (list ¥4.2 input); ≤512K tier; natively multimodal; priority tier 1.5×

EXAMPLE · MONTHLY

One workload, eleven bills

300M input + 30M output tokens per month (~10M in / 1M out daily), no cache or Batch discounts, CNY at 1 USD ≈ ¥7.2 — computed directly from the tables above:

ModelMonthly cost (USD)
deepseek-v4-flash$50
gpt-5.4-nano$98
MiniMax-M3$123
Gemini 3.5 Flash-Lite$165
qwen3.7-max$325
GLM-5.2$450
Claude Haiku 4.5$450
kimi-k3$1,250
gpt-5.6-sol$2,400
Claude Fable 5$4,500

Roughly 90× between the extremes. Caching (hits at 2%–10% of input price) and Batch (~50% off at several vendors) stack to cut another order of magnitude.

CHANGELOG · 6

Price change log

A table like this is right today and wrong next month — so every change observed during re-verification is logged, append-only.

FAQ

Frequently asked

Which LLM API is cheapest in 2026?

Checked against official pricing pages: DeepSeek v4-flash at ¥1 in / ¥2 out per 1M tokens (cache hits just ¥0.02) is the cheapest mainstream tier, followed by ByteDance Doubao seed-2.0-mini (from ¥0.2) and OpenAI gpt-5.4-nano ($0.20/$1.25).

How much can the same workload differ across models?

For 300M input + 30M output tokens per month (no cache/Batch discounts): the priciest flagship runs about $4,500/month vs about $50/month at the low end — roughly 90×. Decide whether the task truly needs a flagship first.

How much do caching and Batch save?

Cache-hit rates typically run 2%–10% of the input price (DeepSeek as low as ¥0.02/1M); OpenAI, Anthropic, Google and Qwen offer ~50% off via Batch. The two stack, cutting repeat-heavy workloads by another order of magnitude.

Sources are the vendors’ official pricing pages (verified 2026-07-22): OpenAI · Anthropic · Google · xAI · Mistral · DeepSeek · Alibaba Model Studio · Moonshot Kimi · Zhipu · ByteDance Doubao · MiniMax. Prices move often — treat the official page as live truth; corrections welcome.