---
title: "LLM API pricing for AI agents: 28 models compared · DotsAgent"
description: "Input, cached and output prices per million tokens for 28 models from OpenAI, Anthropic, Google, xAI, DeepSeek, Qwen and more, with context sizes and agent features."
url: https://dotsagent.io/models
---

Build · prices checked 1 October 2026

# Models and API prices for agent builders

An agent pays for every step it takes, and most of what it pays for is input it has already sent. This table puts list prices, cache discounts and context limits side by side so you can pick a model by what a whole task costs, not by the headline rate.

- [DeepSeek-V4.1-Flash](https://dotsagent.io/models/deepseek-v4-1-flash): $0.22

- [Qwen3.7-Plus](https://dotsagent.io/models/qwen3-7-plus): $0.34

- [GPT-6 Luna](https://dotsagent.io/models/gpt-6-luna): $0.10

## Every model, one table

| — | Provider | — | — | — | — | — |
| --- | --- | --- | --- | --- | --- | --- |
| [GPT-6 Astra](https://dotsagent.io/models/gpt-6-astra)gpt-6-astra | [OpenAI](https://dotsagent.io/providers/openai) | $10* | $1 | $50 | 1.05M | 2026-09-08 |
| [Claude Fable 5.1](https://dotsagent.io/models/claude-fable-5-1)claude-fable-5-1 | [Anthropic](https://dotsagent.io/providers/anthropic) | $10 | $0.25 | $50 | 1M | 2026-09-01 |
| [Claude Opus 5.5](https://dotsagent.io/models/claude-opus-5-5)claude-opus-5-5 | [Anthropic](https://dotsagent.io/providers/anthropic) | $4 | $0.20 | $20 | 1M | 2026-09-22 |
| [Kimi K3](https://dotsagent.io/models/kimi-k3)kimi-k3 | [Moonshot AI (Kimi)](https://dotsagent.io/providers/moonshot) | $3 | $0.30 | $15 | 1.05M | 2026-07-16 |
| [Command A+](https://dotsagent.io/models/command-a-plus)third-party pricecommand-a-plus-05-2026 | [Cohere](https://dotsagent.io/providers/cohere) | $2.50 | — | $10 | 128K | 2026-05-20 |
| [GPT-6.1 Sol](https://dotsagent.io/models/gpt-6-1-sol)gpt-6.1-sol | [OpenAI](https://dotsagent.io/providers/openai) | $2* | $0.10 | $10 | 1.05M | 2026-09-29 |
| [Claude Sonnet 5.5](https://dotsagent.io/models/claude-sonnet-5-5)claude-sonnet-5-5 | [Anthropic](https://dotsagent.io/providers/anthropic) | $2 | $0.20 | $10 | 1M | 2026-09-28 |
| [Gemini 3.1 Pro Preview](https://dotsagent.io/models/gemini-3-1-pro-preview)previewgemini-3.1-pro-preview | [Google](https://dotsagent.io/providers/google) | $2* | $0.20 | $12 | 1.05M | 2026-02-19 |
| [Grok 4.7](https://dotsagent.io/models/grok-4-7)grok-4.7 | [xAI](https://dotsagent.io/providers/xai) | $2* | $0.50 | $6 | 500K | 2026-09-21 |
| [Qwen3.8-Max](https://dotsagent.io/models/qwen3-8-max)qwen3.8-max | [Alibaba Cloud (Qwen)](https://dotsagent.io/providers/qwen) | $1.65 | $0.21 | $4.95 | 1M | 2026-08-03 |
| [Gemini 3.5 Flash](https://dotsagent.io/models/gemini-3-5-flash)gemini-3.5-flash | [Google](https://dotsagent.io/providers/google) | $1.50 | $0.15 | $9 | 1.05M | 2026-05-19 |
| [Mistral Medium 3.5](https://dotsagent.io/models/mistral-medium-3-5)mistral-medium-3-5 | [Mistral AI](https://dotsagent.io/providers/mistral) | $1.50 | — | $7.50 | 256K | 2026-04-28 |
| [GLM-5.3](https://dotsagent.io/models/glm-5-3)glm-5.3 | [Z.ai (Zhipu)](https://dotsagent.io/providers/zhipu) | $1.40 | $0.26 | $4.40 | 1M | 2026-08-14 |
| [DeepSeek-V4-Pro-0813](https://dotsagent.io/models/deepseek-v4-pro-0813)deepseek-v4-pro | [DeepSeek](https://dotsagent.io/providers/deepseek) | $1.32 | $0.044 | $3.96 | 1M | 2026-08-13 |
| [Grok 4.3](https://dotsagent.io/models/grok-4-3)grok-4.3 | [xAI](https://dotsagent.io/providers/xai) | $1.25* | $0.20 | $2.50 | 1M | — |
| [Muse Spark 1.3](https://dotsagent.io/models/muse-spark-1-3)muse-spark-1.3 | [Meta](https://dotsagent.io/providers/meta) | $1.25 | $0.15 | $4.25 | 1M | 2026-09-02 |
| [Amazon Nova 2 Pro](https://dotsagent.io/models/amazon-nova-2-pro)previewthird-party price | [Amazon](https://dotsagent.io/providers/amazon) | $1.25 | — | $10 | — | — |
| [Claude Haiku 4.5](https://dotsagent.io/models/claude-haiku-4-5)claude-haiku-4-5-20251001 | [Anthropic](https://dotsagent.io/providers/anthropic) | $1 | $0.10 | $5 | 200K | — |
| [Grok Build 0.1](https://dotsagent.io/models/grok-build-0-1)grok-build-0.1 | [xAI](https://dotsagent.io/providers/xai) | $1* | $0.20 | $2 | 256K | — |
| [Gemini 3.8 Flash](https://dotsagent.io/models/gemini-3-8-flash)gemini-3.8-flash | [Google](https://dotsagent.io/providers/google) | $0.75 | $0.075 | $3.75 | 1.05M | 2026-09-02 |
| [Gemini 3.5 Flash-Lite](https://dotsagent.io/models/gemini-3-5-flash-lite)gemini-3.5-flash-lite | [Google](https://dotsagent.io/providers/google) | $0.30 | $0.030 | $2.50 | 1.05M | 2026-07-21 |
| [DeepSeek-V4.1-Flash](https://dotsagent.io/models/deepseek-v4-1-flash)deepseek-flash | [DeepSeek](https://dotsagent.io/providers/deepseek) | $0.30 | $0.006 | $1.20 | 1M | 2026-09-10 |
| [Amazon Nova 2 Lite](https://dotsagent.io/models/amazon-nova-2-lite)third-party price | [Amazon](https://dotsagent.io/providers/amazon) | $0.30 | — | $2.50 | 1M | 2025-12-02 |
| [Qwen3.7-Plus](https://dotsagent.io/models/qwen3-7-plus)qwen3.7-plus | [Alibaba Cloud (Qwen)](https://dotsagent.io/providers/qwen) | $0.28* | $0.056 | $1.10 | 1M | 2026-05-26 |
| [Mistral Small 4](https://dotsagent.io/models/mistral-small-4)mistral-small-2603 | [Mistral AI](https://dotsagent.io/providers/mistral) | $0.15 | — | $0.60 | 256K | 2026-03-16 |
| [GLM-5.3-Flash](https://dotsagent.io/models/glm-5-3-flash)glm-5.3-flash | [Z.ai (Zhipu)](https://dotsagent.io/providers/zhipu) | $0.15 | $0.030 | $0.50 | 1M | — |
| [GPT-6 Luna](https://dotsagent.io/models/gpt-6-luna)gpt-6-luna | [OpenAI](https://dotsagent.io/providers/openai) | $0.10* | $0.010 | $0.50 | 1.05M | 2026-09-22 |
| [Qwen3.7-Flash](https://dotsagent.io/models/qwen3-7-flash)qwen3.7-flash | [Alibaba Cloud (Qwen)](https://dotsagent.io/providers/qwen) | $0.028* | $0.006 | $0.11 | 1M | 2026-07-15 |

US dollars per million tokens, standard tier, short-prompt rate · * long prompts cost more

[Price your own workload](https://dotsagent.io/cost)

## How to read these prices

### Cached input is the number that matters

In an agent loop the system prompt, tool definitions and earlier turns are resent on every call. Providers charge a fraction of the normal input rate when that prefix is served from cache, often a tenth or less, so a model with a cheap cache can beat one with a cheaper headline price.

### Long prompts can reprice the whole call

OpenAI, Google's Pro preview, xAI and Qwen charge a higher rate once a single prompt passes a threshold such as 200K or 272K tokens, and the higher rate applies to every token in that request. Anthropic currently bills its full million-token window at the standard rate.

### Reasoning tokens are billed as output

Models that think before answering charge their hidden reasoning at the output rate. Lower the effort setting for routine steps and keep the high settings for planning or hard debugging.

### Batch and off-peak discounts rarely fit agents

Batch endpoints halve the price but return results within hours, which suits evaluations and backfills rather than live agents. DeepSeek's off-peak rate is the exception that can work for scheduled jobs.

## Questions builders ask

### Which model is cheapest for an AI agent?

On our 30-step coding workload the cheapest models are small ones: GPT-6 Luna, Qwen3.7-Flash and GLM-5.3-Flash, each around 10 to 20 cents a task. Cheap per step does not mean cheap per task if the model needs more steps or retries, so test two or three candidates on your own traces.

### Why does my agent cost far more than the price per token suggests?

Because every step resends the growing conversation. A 30-step task with a 25K-token starting prompt sends well over three million input tokens in total. Prompt caching, shorter tool outputs and summarising old turns cut that more than switching models.

### Are these prices final?

They are list prices in US dollars from each provider's pricing page on 1 October 2026, before taxes, regional surcharges or negotiated discounts. Prices marked third-party could not be read from the vendor's own page.

### What is the difference between cached input and cache writes?

A cache read is the discounted rate for tokens already in the cache. Some providers, Anthropic among them, also charge a premium the first time a prefix is written to the cache. The table shows read prices; write prices are noted on each model page where they apply.

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026
