API cost ledger

AI Model API Pricing: current models, real deployment constraints

Use this as a starting ledger, not a purchase order. Token prices are only one input: model availability, cache rules, long-context multipliers, tools, quota, data residency, and retries can dominate an application’s actual cost.

All dollar figures are USD per 1M text tokens unless noted. Preview rows are deliberately labeled; primary source pages win if they change.

Quick reference: current text-model list prices

Availability is a first-class pricing field. GPT-5.6 Sol, Terra, and Luna are in a limited API preview as of this verification date; do not make them the only production path unless your organization has access and a tested fallback.
ModelStatusInputCached inputOutputContext / pricing breakpointUse it first forPrimary source
OpenAI GPT-5.6 SolLimited preview$4.00$0.40$20.001.05M; >272K prompt has multiplierComplex professional / agentic workOpenAI model card
OpenAI GPT-5.6 TerraLimited preview$2.00$0.20$12.001.05M; >272K prompt has multiplierCapability-cost balanceOpenAI model card
OpenAI GPT-5.6 LunaLimited preview$0.20$0.02$1.201.05M; >272K prompt has multiplierCost-sensitive, high-volume workOpenAI model card
Anthropic Claude Opus 4.8Available$5.00$0.50$25.001M standard pricingHigh-capability reasoning/codingAnthropic pricing
Anthropic Claude Sonnet 5Available$2.00$0.20$10.001M standard pricingBalanced agentic workloadsAnthropic pricing
Anthropic Claude Haiku 4.5Available$1.00$0.10$5.00Read exact model context termsFast routing/classificationAnthropic pricing
Google Gemini 3.5 FlashAvailable$1.50$0.15$9.00Tool/search charges can applyFast multimodal/grounded workGemini pricing
Google Gemini 3.1 Flash-LiteAvailable$0.25$0.025$1.50Tools billed separately where usedHigh-volume, simple processingGemini pricing
Google Gemini 3.1 ProPreview$2.00 ≤200K$0.20 ≤200K$12.00 ≤200KInput/output step up above 200KMultimodal/agent evaluationGemini pricing
Mistral Small 4Available; open weights$0.15See live card$0.60256KLow-cost hybrid reasoning/codingMistral model card

The table compares standard text rates, not enterprise agreements, regional premiums, priority/flex, audio/image/video, storage, or third-party cloud resale.

Workload cost calculator

Cost vocabulary and a worked estimate

What can be billable?

  • Input tokens: system instructions, user content, retrieved text, tool schemas, and conversation history.
  • Output and reasoning tokens: generated text and, depending on the provider/model, hidden or intermediate reasoning consumed by an agent loop.
  • Cached input: a reused prefix; cache reads are cheaper, while writes/storage can cost more than a normal input token.
  • Tools and modalities: web search, grounding, code execution, containers, images, audio, video, retrieval, and storage can have separate meters.
  • Quota: RPM, TPM, concurrent requests, and batch queue size can make an otherwise cheap row unusable.
  • Resilience: retries, fallbacks, duplicate tool calls, and human review are deployment costs—not free edge cases.

GPT-5.6 Terra planning example

10,000 requests × 1,800 input + 450 output, with half of input read from cache:

uncached input = 9.0M × $2.00 = $18.00
cached input = 9.0M × $0.20 = $1.80
output = 4.5M × $12.00 = $54.00
token estimate = $73.80 / month

This excludes tools and assumes preview access. For >272K GPT-5.6 prompts, recalculate using the published long-context multipliers—not the base row.

Choose the operating mode before the model

ConstraintStart withWhyDo not assume
Best current OpenAI capability and approved preview accessGPT-5.6 SolFlagship 1.05M-context tier with tool supportThat a preview is a universal production dependency
Capability/cost balance in the same preview familyGPT-5.6 TerraHalf Sol’s listed text rateThat longer prompts retain the base price
High-volume simple routingGemini 3.1 Flash-Lite or GPT-5.6 LunaLow list rates, with different access postureThat quality, tool reliability, or quota are interchangeable
Known Claude production pathSonnet 5, then Haiku 4.5 for cheaper tasksPublished model, cache, and dated rate scheduleThat older tokenizer measurements still predict cost
Open-weight controlMistral Small 4 or another evaluated open familyWeights may lower lock-in and enable self-hostingThat downloads eliminate GPU, ops, licence, or security cost

Advanced cost traps and anti-patterns

Preview, aliases, and fallback design

A model card can publish a price before broad access exists. Pin a dated model identifier where the provider supports snapshots, keep a tested fallback in the same product path, and make your rate/quota behavior observable. Consumer plans are not API credits.

Context, caching, and agent loops

Long context is capacity, not a mandate to send the corpus. Cache a stable prefix only when the provider’s cache lifetime, write charge, and hit behavior work for your traffic. Cap agent turns, tool calls, retrieved documents, and output tokens; intermediate work can dominate the final answer.

Apples-to-apples evaluation

Price per token is not price per completed task. Compare models on identical prompts, tool schemas, retry rules, safety policy, output limits, region, and service tier. Re-tokenize when a vendor changes tokenizer: Anthropic documents that newer Claude models can produce about 30% more tokens for the same text.

Common mistakes

  • Using a subscription price or an outdated screenshot as API pricing evidence.
  • Comparing input-only price while ignoring output, cache writes, tools, and storage.
  • Budgeting with an unapproved preview model and no fallback.
  • Assuming published rate limits are sustained throughput guarantees.
  • Using a giant context to avoid retrieval, ranking, or document selection.
  • Calling self-hosting “free” because the weights are downloadable.

Continue the decision

Frontier AI labs · Open-weight model deployment · AI model picker · AI coding agents compared.

Sources: OpenAI models · GPT-5.6 preview terms · Anthropic pricing · Gemini pricing · Mistral Small 4.

David Veksler is a Principal AI Engineer in Denver. He leads agentic AI engineering at Antech, a Mars company, and builds AI platforms for regulated financial firms. This page was produced by a governed, multi-agent Claude Code pipeline with a git audit trail. How it's built Case studies