---
title: "AI agent cost calculator: price a task on 28 models · DotsAgent"
description: "Estimate what one agent task and a month of tasks cost on 28 LLM APIs. Set prompt size, context growth, steps, output and cache hit rate; long-prompt rates included."
url: https://dotsagent.io/cost
---

Build · tool

# Agent cost calculator

An agent is a loop: each call resends everything so far plus the newest tool result. Describe that loop once and see what a task, and a month of tasks, costs on every model in our price table.

| Model | Per task | 30 days at 50 a day |
| --- | --- | --- |
| [GPT-6 Luna](https://dotsagent.io/models/gpt-6-luna) · OpenAI | $0.099 | $148 |
| [Qwen3.7-Flash](https://dotsagent.io/models/qwen3-7-flash) · Alibaba Cloud (Qwen)28 calls at long-prompt rate | $0.101 | $152 |
| [GLM-5.3-Flash](https://dotsagent.io/models/glm-5-3-flash) · Z.ai (Zhipu) | $0.182 | $273 |
| [DeepSeek-V4.1-Flash](https://dotsagent.io/models/deepseek-v4-1-flash) · DeepSeek | $0.218 | $327 |
| [Gemini 3.5 Flash-Lite](https://dotsagent.io/models/gemini-3-5-flash-lite) · Google | $0.333 | $499 |
| [Qwen3.7-Plus](https://dotsagent.io/models/qwen3-7-plus) · Alibaba Cloud (Qwen) | $0.343 | $515 |
| [Mistral Small 4](https://dotsagent.io/models/mistral-small-4) · Mistral AI | $0.526 | $788 |
| [Gemini 3.8 Flash](https://dotsagent.io/models/gemini-3-8-flash) · Google | $0.742 | $1,112 |
| [DeepSeek-V4-Pro-0813](https://dotsagent.io/models/deepseek-v4-pro-0813) · DeepSeek | $0.961 | $1,441 |
| [Claude Haiku 4.5](https://dotsagent.io/models/claude-haiku-4-5) · Anthropic | $0.989 | $1,483 |
| [Amazon Nova 2 Lite](https://dotsagent.io/models/amazon-nova-2-lite) · Amazonthird-party price | $1.10 | $1,647 |
| [Grok Build 0.1](https://dotsagent.io/models/grok-build-0-1) · xAI | $1.16 | $1,746 |
| [Muse Spark 1.3](https://dotsagent.io/models/muse-spark-1-3) · Meta | $1.23 | $1,852 |
| [Grok 4.3](https://dotsagent.io/models/grok-4-3) · xAI | $1.31 | $1,970 |
| [Gemini 3.5 Flash](https://dotsagent.io/models/gemini-3-5-flash) · Google | $1.54 | $2,306 |
| [Qwen3.8-Max](https://dotsagent.io/models/qwen3-8-max) · Alibaba Cloud (Qwen) | $1.63 | $2,443 |
| [GLM-5.3](https://dotsagent.io/models/glm-5-3) · Z.ai (Zhipu) | $1.63 | $2,446 |
| [GPT-6.1 Sol](https://dotsagent.io/models/gpt-6-1-sol) · OpenAI | $1.69 | $2,541 |
| [Claude Sonnet 5.5](https://dotsagent.io/models/claude-sonnet-5-5) · Anthropic | $1.98 | $2,966 |
| [Gemini 3.1 Pro Preview](https://dotsagent.io/models/gemini-3-1-pro-preview) · Google | $2.05 | $3,074 |
| [Grok 4.7](https://dotsagent.io/models/grok-4-7) · xAI | $2.68 | $4,026 |
| [Kimi K3](https://dotsagent.io/models/kimi-k3) · Moonshot AI (Kimi) | $2.97 | $4,449 |
| [Claude Opus 5.5](https://dotsagent.io/models/claude-opus-5-5) · Anthropic | $3.39 | $5,082 |
| [Amazon Nova 2 Pro](https://dotsagent.io/models/amazon-nova-2-pro) · Amazonthird-party price | $4.56 | $6,840 |
| [Mistral Medium 3.5](https://dotsagent.io/models/mistral-medium-3-5) · Mistral AI | $5.31 | $7,965 |
| [Claude Fable 5.1](https://dotsagent.io/models/claude-fable-5-1) · Anthropic | $7.76 | $11,642 |
| [Command A+](https://dotsagent.io/models/command-a-plus) · Cohereexceeds context windowthird-party price | $8.76 | $13,140 |
| [GPT-6 Astra](https://dotsagent.io/models/gpt-6-astra) · OpenAI | $9.89 | $14,831 |

List prices, standard tier, US dollars. Cache writes, taxes, tool fees such as web search and regional surcharges are not included.

## How the estimate works

For each call in the loop we take the prompt size, the first prompt plus the growth from every earlier step, and split it into a cached share and a fresh share. The first call is always fresh. Fresh tokens are billed at the input rate, cached tokens at the cached rate, and the output at the output rate.

If a single prompt passes a provider's long-context threshold, that whole call is billed at the long-prompt rate, as OpenAI, Google, xAI and Qwen do. If the largest prompt is bigger than the model's context window, the row is flagged, because a real agent would have to trim or summarise its history first.

Real agents vary from task to task, retry failed steps and sometimes stop early. Treat these figures as a way to compare models on the same workload, then measure the actual token counts from your own traces before you budget.

## Questions about agent costs

### How many tokens does an AI agent use per task?

It depends mostly on the number of steps and how much each step adds. A short support exchange may use 30K input tokens; a 30-step coding task with a 25K-token starting prompt sends over three million, because the whole history is resent every call.

### How do I lower the cost of an agent?

Keep the start of the prompt identical between calls so it caches, trim tool results to what the model needs, summarise or drop old turns, use a small model for routine steps and a large one only for planning, and lower the reasoning effort where quality allows.

### Does prompt caching happen automatically?

With OpenAI, DeepSeek, Kimi and Gemini's implicit caching it does when prompts share a long prefix. Anthropic and some Qwen endpoints need explicit cache markers, and Anthropic charges extra for the first write.

### Why is the cheapest model per token not the cheapest per task?

Cache discounts differ by model, some providers bill long prompts at a higher rate, and reasoning models spend output tokens thinking. A model that needs fewer steps or fewer retries can also come out ahead, which this calculator cannot see.

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026
