Models and API prices for agent builders
An agent pays for every step it takes, and most of what it pays for is input it has already sent. This table puts list prices, cache discounts and context limits side by side so you can pick a model by what a whole task costs, not by the headline rate.
Every model, one table
| Provider | ||||||
|---|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | $10* | $1 | $50 | 1.05M | 2026-09-08 |
| Claude Fable 5.1 | Anthropic | $10 | $0.25 | $50 | 1M | 2026-09-01 |
| Claude Opus 5.5 | Anthropic | $4 | $0.20 | $20 | 1M | 2026-09-22 |
| Kimi K3 | Moonshot AI (Kimi) | $3 | $0.30 | $15 | 1.05M | 2026-07-16 |
| Command A+third-party price | Cohere | $2.50 | — | $10 | 128K | 2026-05-20 |
| GPT-6.1 Sol | OpenAI | $2* | $0.10 | $10 | 1.05M | 2026-09-29 |
| Claude Sonnet 5.5 | Anthropic | $2 | $0.20 | $10 | 1M | 2026-09-28 |
| Gemini 3.1 Pro Previewpreview | $2* | $0.20 | $12 | 1.05M | 2026-02-19 | |
| Grok 4.7 | xAI | $2* | $0.50 | $6 | 500K | 2026-09-21 |
| Qwen3.8-Max | Alibaba Cloud (Qwen) | $1.65 | $0.21 | $4.95 | 1M | 2026-08-03 |
| Gemini 3.5 Flash | $1.50 | $0.15 | $9 | 1.05M | 2026-05-19 | |
| Mistral Medium 3.5 | Mistral AI | $1.50 | — | $7.50 | 256K | 2026-04-28 |
| GLM-5.3 | Z.ai (Zhipu) | $1.40 | $0.26 | $4.40 | 1M | 2026-08-14 |
| DeepSeek-V4-Pro-0813 | DeepSeek | $1.32 | $0.044 | $3.96 | 1M | 2026-08-13 |
| Grok 4.3 | xAI | $1.25* | $0.20 | $2.50 | 1M | — |
| Muse Spark 1.3 | Meta | $1.25 | $0.15 | $4.25 | 1M | 2026-09-02 |
| Amazon Nova 2 Propreviewthird-party price | Amazon | $1.25 | — | $10 | — | — |
| Claude Haiku 4.5 | Anthropic | $1 | $0.10 | $5 | 200K | — |
| Grok Build 0.1 | xAI | $1* | $0.20 | $2 | 256K | — |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | 1.05M | 2026-09-02 | |
| Gemini 3.5 Flash-Lite | $0.30 | $0.030 | $2.50 | 1.05M | 2026-07-21 | |
| DeepSeek-V4.1-Flash | DeepSeek | $0.30 | $0.006 | $1.20 | 1M | 2026-09-10 |
| Amazon Nova 2 Litethird-party price | Amazon | $0.30 | — | $2.50 | 1M | 2025-12-02 |
| Qwen3.7-Plus | Alibaba Cloud (Qwen) | $0.28* | $0.056 | $1.10 | 1M | 2026-05-26 |
| Mistral Small 4 | Mistral AI | $0.15 | — | $0.60 | 256K | 2026-03-16 |
| GLM-5.3-Flash | Z.ai (Zhipu) | $0.15 | $0.030 | $0.50 | 1M | — |
| GPT-6 Luna | OpenAI | $0.10* | $0.010 | $0.50 | 1.05M | 2026-09-22 |
| Qwen3.7-Flash | Alibaba Cloud (Qwen) | $0.028* | $0.006 | $0.11 | 1M | 2026-07-15 |
How to read these prices
Cached input is the number that matters
In an agent loop the system prompt, tool definitions and earlier turns are resent on every call. Providers charge a fraction of the normal input rate when that prefix is served from cache, often a tenth or less, so a model with a cheap cache can beat one with a cheaper headline price.
Long prompts can reprice the whole call
OpenAI, Google's Pro preview, xAI and Qwen charge a higher rate once a single prompt passes a threshold such as 200K or 272K tokens, and the higher rate applies to every token in that request. Anthropic currently bills its full million-token window at the standard rate.
Reasoning tokens are billed as output
Models that think before answering charge their hidden reasoning at the output rate. Lower the effort setting for routine steps and keep the high settings for planning or hard debugging.
Batch and off-peak discounts rarely fit agents
Batch endpoints halve the price but return results within hours, which suits evaluations and backfills rather than live agents. DeepSeek's off-peak rate is the exception that can work for scheduled jobs.
Questions builders ask
Which model is cheapest for an AI agent?
On our 30-step coding workload the cheapest models are small ones: GPT-6 Luna, Qwen3.7-Flash and GLM-5.3-Flash, each around 10 to 20 cents a task. Cheap per step does not mean cheap per task if the model needs more steps or retries, so test two or three candidates on your own traces.
Why does my agent cost far more than the price per token suggests?
Because every step resends the growing conversation. A 30-step task with a 25K-token starting prompt sends well over three million input tokens in total. Prompt caching, shorter tool outputs and summarising old turns cut that more than switching models.
Are these prices final?
They are list prices in US dollars from each provider's pricing page on 1 October 2026, before taxes, regional surcharges or negotiated discounts. Prices marked third-party could not be read from the vendor's own page.
What is the difference between cached input and cache writes?
A cache read is the discounted rate for tokens already in the cache. Some providers, Anthropic among them, also charge a premium the first time a prefix is written to the cache. The table shows read prices; write prices are noted on each model page where they apply.