Quick reference: current text-model list prices
| Model | Status | Input | Cached input | Output | Context / pricing breakpoint | Use it first for | Primary source |
|---|---|---|---|---|---|---|---|
| OpenAI GPT-5.6 Sol | Limited preview | $4.00 | $0.40 | $20.00 | 1.05M; >272K prompt has multiplier | Complex professional / agentic work | OpenAI model card |
| OpenAI GPT-5.6 Terra | Limited preview | $2.00 | $0.20 | $12.00 | 1.05M; >272K prompt has multiplier | Capability-cost balance | OpenAI model card |
| OpenAI GPT-5.6 Luna | Limited preview | $0.20 | $0.02 | $1.20 | 1.05M; >272K prompt has multiplier | Cost-sensitive, high-volume work | OpenAI model card |
| Anthropic Claude Opus 4.8 | Available | $5.00 | $0.50 | $25.00 | 1M standard pricing | High-capability reasoning/coding | Anthropic pricing |
| Anthropic Claude Sonnet 5 | Available | $2.00 | $0.20 | $10.00 | 1M standard pricing | Balanced agentic workloads | Anthropic pricing |
| Anthropic Claude Haiku 4.5 | Available | $1.00 | $0.10 | $5.00 | Read exact model context terms | Fast routing/classification | Anthropic pricing |
| Google Gemini 3.5 Flash | Available | $1.50 | $0.15 | $9.00 | Tool/search charges can apply | Fast multimodal/grounded work | Gemini pricing |
| Google Gemini 3.1 Flash-Lite | Available | $0.25 | $0.025 | $1.50 | Tools billed separately where used | High-volume, simple processing | Gemini pricing |
| Google Gemini 3.1 Pro | Preview | $2.00 ≤200K | $0.20 ≤200K | $12.00 ≤200K | Input/output step up above 200K | Multimodal/agent evaluation | Gemini pricing |
| Mistral Small 4 | Available; open weights | $0.15 | See live card | $0.60 | 256K | Low-cost hybrid reasoning/coding | Mistral model card |
The table compares standard text rates, not enterprise agreements, regional premiums, priority/flex, audio/image/video, storage, or third-party cloud resale.
Workload cost calculator
Cost vocabulary and a worked estimate
What can be billable?
- Input tokens: system instructions, user content, retrieved text, tool schemas, and conversation history.
- Output and reasoning tokens: generated text and, depending on the provider/model, hidden or intermediate reasoning consumed by an agent loop.
- Cached input: a reused prefix; cache reads are cheaper, while writes/storage can cost more than a normal input token.
- Tools and modalities: web search, grounding, code execution, containers, images, audio, video, retrieval, and storage can have separate meters.
- Quota: RPM, TPM, concurrent requests, and batch queue size can make an otherwise cheap row unusable.
- Resilience: retries, fallbacks, duplicate tool calls, and human review are deployment costs—not free edge cases.
GPT-5.6 Terra planning example
10,000 requests × 1,800 input + 450 output, with half of input read from cache:
cached input = 9.0M × $0.20 = $1.80
output = 4.5M × $12.00 = $54.00
token estimate = $73.80 / month
This excludes tools and assumes preview access. For >272K GPT-5.6 prompts, recalculate using the published long-context multipliers—not the base row.
Choose the operating mode before the model
| Constraint | Start with | Why | Do not assume |
|---|---|---|---|
| Best current OpenAI capability and approved preview access | GPT-5.6 Sol | Flagship 1.05M-context tier with tool support | That a preview is a universal production dependency |
| Capability/cost balance in the same preview family | GPT-5.6 Terra | Half Sol’s listed text rate | That longer prompts retain the base price |
| High-volume simple routing | Gemini 3.1 Flash-Lite or GPT-5.6 Luna | Low list rates, with different access posture | That quality, tool reliability, or quota are interchangeable |
| Known Claude production path | Sonnet 5, then Haiku 4.5 for cheaper tasks | Published model, cache, and dated rate schedule | That older tokenizer measurements still predict cost |
| Open-weight control | Mistral Small 4 or another evaluated open family | Weights may lower lock-in and enable self-hosting | That downloads eliminate GPU, ops, licence, or security cost |
Advanced cost traps and anti-patterns
Preview, aliases, and fallback design
A model card can publish a price before broad access exists. Pin a dated model identifier where the provider supports snapshots, keep a tested fallback in the same product path, and make your rate/quota behavior observable. Consumer plans are not API credits.
Context, caching, and agent loops
Long context is capacity, not a mandate to send the corpus. Cache a stable prefix only when the provider’s cache lifetime, write charge, and hit behavior work for your traffic. Cap agent turns, tool calls, retrieved documents, and output tokens; intermediate work can dominate the final answer.
Apples-to-apples evaluation
Price per token is not price per completed task. Compare models on identical prompts, tool schemas, retry rules, safety policy, output limits, region, and service tier. Re-tokenize when a vendor changes tokenizer: Anthropic documents that newer Claude models can produce about 30% more tokens for the same text.
Common mistakes
- Using a subscription price or an outdated screenshot as API pricing evidence.
- Comparing input-only price while ignoring output, cache writes, tools, and storage.
- Budgeting with an unapproved preview model and no fallback.
- Assuming published rate limits are sustained throughput guarantees.
- Using a giant context to avoid retrieval, ranking, or document selection.
- Calling self-hosting “free” because the weights are downloadable.
Continue the decision
Frontier AI labs · Open-weight model deployment · AI model picker · AI coding agents compared.
Sources: OpenAI models · GPT-5.6 preview terms · Anthropic pricing · Gemini pricing · Mistral Small 4.