12 models · 5 providers · prices verified 2026-10-08
LLM API Cost & Token Calculator
Compare per-1M-token prices for Anthropic, OpenAI, Google, DeepSeek, xAI models and simulate your real monthly bill with prompt caching. Set your workload once; every number updates instantly, in USD or KRW.
- Cheapest input
- Claude Haiku 5.5
- $0.1 / 1M
- Priciest input
- Claude Fable 5.1
- $10 / 1M
- Price spread
- 100×
- input price, most vs least
- Cache discount
- up to 98%
- on cached input tokens
Monthly workload
Prompt + context sent to the model each month, in millions.
Tokens generated by the model each month, in millions.
Share of input tokens served from the prompt cache. Agents: 60-80%. Plain chat: 0-30%.
Monthly cost ranking
Cheapest first · click a row to pick it for the VS panel
Spec & pricing table
USD per 1M tokens · tick two rows to compare · hover a name for notes
| Select | Model | Provider | Context | Input | Cached input | Output | Key strength |
|---|---|---|---|---|---|---|---|
Claude Fable 5.1 claude-fable-5-1 | Anthropic | 1M | $10.00 | $0.25(-98%) | $50.00 | Most Capable | |
Claude Opus 5.5 claude-opus-5-5 | Anthropic | 1M | $4.00 | $0.20(-95%) | $20.00 | Flagship | |
Claude Sonnet 5.5 claude-sonnet-5-5 | Anthropic | 1M | $2.00 | $0.10(-95%) | $10.00 | Best for Coding | |
Claude Haiku 5.5 claude-haiku-5-5 | Anthropic | 1M | $0.10 | $0.010(-90%) | $0.50 | Budget | |
GPT-6 Astra gpt-6-astra | OpenAI | 1.05M | $10.00 | $1.00(-90%) | $50.00 | Most Capable | |
GPT-6.1 Sol gpt-6.1-sol | OpenAI | 1.05M | $2.00 | $0.10(-95%) | $10.00 | Flagship | |
GPT-6 Luna gpt-6-luna | OpenAI | 1.05M | $0.10 | $0.010(-90%) | $0.50 | Budget | |
Gemini 3.1 Pro (Preview) gemini-3.1-pro-preview | 1M | $2.00 | $0.20(-90%) | $12.00 | Preview | ||
Gemini 3.8 Flash gemini-3.8-flash | 1M | $0.75 | $0.075(-90%) | $3.75 | Best Value | ||
DeepSeek V4 Pro deepseek-v4-pro | DeepSeek | 1M | $1.32 | $0.044(-97%) | $3.96 | Budget Frontier | |
DeepSeek V4.1 Flash deepseek-flash | DeepSeek | 1M | $0.30 | $0.006(-98%) | $1.20 | Budget | |
Grok 4.7 grok-4.7 | xAI | 500K | $2.00 | $0.50(-75%) | $6.00 | Flagship |
Pricing notes per model
- Claude Fable 5.1 — Anthropic's most capable model. Cache reads are 0.025x input (cheapest cache ratio of any model listed). Thinking always on. source (checked 2026-10-08)
- Claude Opus 5.5 — Current Opus for long-running agentic coding and knowledge work. Cache reads 0.05x input. 128K max output. source (checked 2026-10-08)
- Claude Sonnet 5.5 — Everyday coding and agent workhorse. Cache reads 0.05x input. 128K max output. source (checked 2026-10-08)
- Claude Haiku 5.5 — Price shown is for prompts up to 100K tokens; prompts over 100K are billed at $0.50 / $2.50. High-volume classification, routing, extraction. source (checked 2026-10-08)
- GPT-6 Astra — OpenAI's most capable model for complex reasoning, coding, computer use and research. Price shown is for prompts under 272K tokens. source (checked 2026-10-08)
- GPT-6.1 Sol — Near-Astra performance at a fifth of the price. Cache reads 0.05x input. Price shown is for prompts under 272K tokens. source (checked 2026-10-08)
- GPT-6 Luna — Efficient model for focused, high-volume tasks. Same list price as Claude Haiku 5.5. source (checked 2026-10-08)
- Gemini 3.1 Pro (Preview) — Price shown is for prompts up to 200K tokens; over 200K it is $4 / $18. Context-cache storage billed separately ($4.50 per 1M tokens per hour). source (checked 2026-10-08)
- Gemini 3.8 Flash — Google's most intelligent Flash model. Promotional price through 2026-12-31; from 2027-01-01 it is $1.50 / $7.50 (cached $0.15). Cache storage $0.50 per 1M tokens per hour. source (checked 2026-10-08)
- DeepSeek V4 Pro — Peak-hour price (01:00-04:00 and 06:00-10:00 UTC, Mon-Fri). Off-peak is 50% lower. 384K max output. source (checked 2026-10-08)
- DeepSeek V4.1 Flash — API model name `deepseek-flash`. Peak-hour price; off-peak is 50% lower. Cache hits cost 2% of input - the deepest cache discount listed. source (checked 2026-10-08)
- Grok 4.7 — xAI flagship for programming and agent tasks. Price shown is for prompts under 200K tokens; at 200K+ it doubles ($4 / $12, cached $1). source (checked 2026-10-08)
Side-by-side comparison
Differences are shown for A relative to B. Green means A is cheaper.
InsightClaude Sonnet 5.5 and GPT-6.1 Sol are priced identically on input, output and at your workload.
| Metric | Claude Sonnet 5.5 | GPT-6.1 Sol | A vs B |
|---|---|---|---|
| Input / 1M | $2.00 | $2.00 | same |
| Cached input / 1M | $0.10 | $0.10 | same |
| Output / 1M | $10.00 | $10.00 | same |
| Monthly input cost | $21.00 | $21.00 | same |
| Monthly output cost | $50.00 | $50.00 | same |
| Monthly total | $71.00 | $71.00 | same |
| Context window | 1M | 1.05M | — |
- Claude Sonnet 5.5
- Everyday coding and agent workhorse. Cache reads 0.05x input. 128K max output.
- GPT-6.1 Sol
- Near-Astra performance at a fifth of the price. Cache reads 0.05x input. Price shown is for prompts under 272K tokens.
FAQ
Prompt caching, token estimation, and how the numbers are computed.
What is prompt caching and why does it change the ranking?
Prompt caching lets a provider reuse the processed form of a prompt prefix (system prompt, tool definitions, a long document) across requests instead of re-reading it every time. Cached input tokens are billed at a steep discount: 10% of the input price on most models, 5% on Claude Opus 5.5, Claude Sonnet 5.5 and GPT-6.1 Sol, 2.5% on Claude Fable 5.1, and 2% on DeepSeek. Because the discount differs per model, a model that looks expensive at 0% cache hits can become the cheapest at 80% hits. The cache-hit slider above applies the model's cached price to that share of your input tokens.
Which LLM is most cost-effective for coding?
It depends on your output volume. Coding agents are output-heavy and cache-heavy: they re-send the same repository context many times and generate a lot of code. Set the output slider to match your real usage and the cache-hit ratio to 60-80% - that is typical for agentic loops - and read the ranking. As of the latest check, Claude Sonnet 5.5 and GPT-6.1 Sol sit in the same price band ($2 in / $10 out), Gemini 3.8 Flash and DeepSeek V4 Pro are cheaper, and Claude Opus 5.5 is the premium pick at $4 / $20. Price is only one axis; benchmark quality on your own tasks before choosing.
How do I estimate my monthly token usage?
Start from requests per day. A typical chat turn is 500-3,000 input tokens (including history) and 100-800 output tokens; a RAG query with retrieved documents is 3,000-15,000 input tokens; an agent step with tools can be 10,000-50,000 input tokens. Multiply by daily requests and by 30. For English text, 1 token is roughly 4 characters or 0.75 words; Korean and code usually cost more tokens per character. Most providers expose a usage object on every API response - log input_tokens, output_tokens and cache_read tokens for a week and plug the monthly totals in here.
Are these prices exact? Where do they come from?
Every number is copied from the provider's official pricing page and the date of the last check is stored with each model (see the footer badge and the source links). The table shows the standard tier price for prompts within the base context tier. It does not apply batch discounts (usually 50%), DeepSeek off-peak pricing (50% lower), Gemini long-context surcharges above 200K tokens, OpenAI pricing above 272K tokens, xAI pricing above 200K tokens, Claude Haiku 5.5 pricing above 100K tokens, or context-cache storage fees. Treat the result as a list-price estimate, not an invoice.
How is the monthly cost calculated?
Monthly cost = input tokens × (1 − cache hit ratio) × input price + input tokens × cache hit ratio × cached input price + output tokens × output price. All prices are per 1M tokens, so the sliders are in millions. If a model has no cached price, the cache slider has no effect on it. The KRW figure is simply the USD total multiplied by the exchange rate you set (default 1,400).
Can I use this for GPT vs Claude vs Gemini comparisons?
Yes. Pick any two models in the VS panel (or tick two rows in the table). The panel highlights the percentage difference on input, cached input, output and total monthly cost for your workload, and generates a one-line summary you can paste into a cost review.
How to read this page
Input vs output. Output tokens cost 3-5× more than input on every model listed, so chatty assistants and code generators are dominated by the output price. Summarisation and classification are dominated by input.
Cached input. Anything you resend unchanged - system prompts, tool schemas, long documents, repository context - can be served from the prompt cache at 2-10% of the input price. The hit ratio slider models that.
Context window. A 1M window does not mean a 1M prompt is cheap: several providers charge more above 100K-272K tokens. Keep prompts inside the base tier unless you need the room.