dotsagent.io
Language:English
Build · prices checked 1 October 2026

Models and API prices for agent builders

An agent pays for every step it takes, and most of what it pays for is input it has already sent. This table puts list prices, cache discounts and context limits side by side so you can pick a model by what a whole task costs, not by the headline rate.

Every model, one table

API list prices of models used to build agents
Provider
GPT-6 Astra
gpt-6-astra
OpenAI⁨$10⁩*⁨$1⁩⁨$50⁩1.05M2026-09-08
Claude Fable 5.1
claude-fable-5-1
Anthropic⁨$10⁩⁨$0.25⁩⁨$50⁩1M2026-09-01
Claude Opus 5.5
claude-opus-5-5
Anthropic⁨$4⁩⁨$0.20⁩⁨$20⁩1M2026-09-22
Kimi K3
kimi-k3
Moonshot AI (Kimi)⁨$3⁩⁨$0.30⁩⁨$15⁩1.05M2026-07-16
Command A+third-party price
command-a-plus-05-2026
Cohere⁨$2.50⁩—⁨$10⁩128K2026-05-20
GPT-6.1 Sol
gpt-6.1-sol
OpenAI⁨$2⁩*⁨$0.10⁩⁨$10⁩1.05M2026-09-29
Claude Sonnet 5.5
claude-sonnet-5-5
Anthropic⁨$2⁩⁨$0.20⁩⁨$10⁩1M2026-09-28
Gemini 3.1 Pro Previewpreview
gemini-3.1-pro-preview
Google⁨$2⁩*⁨$0.20⁩⁨$12⁩1.05M2026-02-19
Grok 4.7
grok-4.7
xAI⁨$2⁩*⁨$0.50⁩⁨$6⁩500K2026-09-21
Qwen3.8-Max
qwen3.8-max
Alibaba Cloud (Qwen)⁨$1.65⁩⁨$0.21⁩⁨$4.95⁩1M2026-08-03
Gemini 3.5 Flash
gemini-3.5-flash
Google⁨$1.50⁩⁨$0.15⁩⁨$9⁩1.05M2026-05-19
Mistral Medium 3.5
mistral-medium-3-5
Mistral AI⁨$1.50⁩—⁨$7.50⁩256K2026-04-28
GLM-5.3
glm-5.3
Z.ai (Zhipu)⁨$1.40⁩⁨$0.26⁩⁨$4.40⁩1M2026-08-14
DeepSeek-V4-Pro-0813
deepseek-v4-pro
DeepSeek⁨$1.32⁩⁨$0.044⁩⁨$3.96⁩1M2026-08-13
Grok 4.3
grok-4.3
xAI⁨$1.25⁩*⁨$0.20⁩⁨$2.50⁩1M—
Muse Spark 1.3
muse-spark-1.3
Meta⁨$1.25⁩⁨$0.15⁩⁨$4.25⁩1M2026-09-02
Amazon Nova 2 Propreviewthird-party priceAmazon⁨$1.25⁩—⁨$10⁩——
Claude Haiku 4.5
claude-haiku-4-5-20251001
Anthropic⁨$1⁩⁨$0.10⁩⁨$5⁩200K—
Grok Build 0.1
grok-build-0.1
xAI⁨$1⁩*⁨$0.20⁩⁨$2⁩256K—
Gemini 3.8 Flash
gemini-3.8-flash
Google⁨$0.75⁩⁨$0.075⁩⁨$3.75⁩1.05M2026-09-02
Gemini 3.5 Flash-Lite
gemini-3.5-flash-lite
Google⁨$0.30⁩⁨$0.030⁩⁨$2.50⁩1.05M2026-07-21
DeepSeek-V4.1-Flash
deepseek-flash
DeepSeek⁨$0.30⁩⁨$0.006⁩⁨$1.20⁩1M2026-09-10
Amazon Nova 2 Litethird-party priceAmazon⁨$0.30⁩—⁨$2.50⁩1M2025-12-02
Qwen3.7-Plus
qwen3.7-plus
Alibaba Cloud (Qwen)⁨$0.28⁩*⁨$0.056⁩⁨$1.10⁩1M2026-05-26
Mistral Small 4
mistral-small-2603
Mistral AI⁨$0.15⁩—⁨$0.60⁩256K2026-03-16
GLM-5.3-Flash
glm-5.3-flash
Z.ai (Zhipu)⁨$0.15⁩⁨$0.030⁩⁨$0.50⁩1M—
GPT-6 Luna
gpt-6-luna
OpenAI⁨$0.10⁩*⁨$0.010⁩⁨$0.50⁩1.05M2026-09-22
Qwen3.7-Flash
qwen3.7-flash
Alibaba Cloud (Qwen)⁨$0.028⁩*⁨$0.006⁩⁨$0.11⁩1M2026-07-15

US dollars per million tokens, standard tier, short-prompt rate · * long prompts cost more

Price your own workload

How to read these prices

Cached input is the number that matters

In an agent loop the system prompt, tool definitions and earlier turns are resent on every call. Providers charge a fraction of the normal input rate when that prefix is served from cache, often a tenth or less, so a model with a cheap cache can beat one with a cheaper headline price.

Long prompts can reprice the whole call

OpenAI, Google's Pro preview, xAI and Qwen charge a higher rate once a single prompt passes a threshold such as 200K or 272K tokens, and the higher rate applies to every token in that request. Anthropic currently bills its full million-token window at the standard rate.

Reasoning tokens are billed as output

Models that think before answering charge their hidden reasoning at the output rate. Lower the effort setting for routine steps and keep the high settings for planning or hard debugging.

Batch and off-peak discounts rarely fit agents

Batch endpoints halve the price but return results within hours, which suits evaluations and backfills rather than live agents. DeepSeek's off-peak rate is the exception that can work for scheduled jobs.

Questions builders ask

Which model is cheapest for an AI agent?

On our 30-step coding workload the cheapest models are small ones: GPT-6 Luna, Qwen3.7-Flash and GLM-5.3-Flash, each around 10 to 20 cents a task. Cheap per step does not mean cheap per task if the model needs more steps or retries, so test two or three candidates on your own traces.

Why does my agent cost far more than the price per token suggests?

Because every step resends the growing conversation. A 30-step task with a 25K-token starting prompt sends well over three million input tokens in total. Prompt caching, shorter tool outputs and summarising old turns cut that more than switching models.

Are these prices final?

They are list prices in US dollars from each provider's pricing page on 1 October 2026, before taxes, regional surcharges or negotiated discounts. Prices marked third-party could not be read from the vendor's own page.

What is the difference between cached input and cache writes?

A cache read is the discounted rate for tokens already in the cache. Some providers, Anthropic among them, also charge a premium the first time a prefix is written to the cache. The table shows read prices; write prices are noted on each model page where they apply.

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026