---
title: "AI Agent 的 LLM API 定价：28 款模型对比 · DotsAgent"
description: "对比 OpenAI、Anthropic、Google、xAI、DeepSeek、Qwen 等厂商的 28 款模型，列出每百万 token 的输入、缓存和输出价格，以及上下文长度和 Agent 功能。"
url: https://dotsagent.io/zh/models
---

构建 · 价格核验于 2026 年 10 月 1 日

# 面向 Agent 开发者的模型与 API 价格

Agent 每执行一步都要付费，而大部分费用都花在之前已发送过的输入上。本表并列展示公开定价、缓存折扣和上下文限制，帮助你根据整个任务的成本选择模型，而不是只看醒目的单价。

- [DeepSeek-V4.1-Flash](https://dotsagent.io/zh/models/deepseek-v4-1-flash): US$0.22

- [Qwen3.7-Plus](https://dotsagent.io/zh/models/qwen3-7-plus): US$0.34

- [GPT-6 Luna](https://dotsagent.io/zh/models/gpt-6-luna): US$0.10

## 所有模型，一览无余

| — | 提供商 | — | — | — | — | — |
| --- | --- | --- | --- | --- | --- | --- |
| [GPT-6 Astra](https://dotsagent.io/zh/models/gpt-6-astra)gpt-6-astra | [OpenAI](https://dotsagent.io/zh/providers/openai) | US$10* | US$1 | US$50 | 1.05M | 2026-09-08 |
| [Claude Fable 5.1](https://dotsagent.io/zh/models/claude-fable-5-1)claude-fable-5-1 | [Anthropic](https://dotsagent.io/zh/providers/anthropic) | US$10 | US$0.25 | US$50 | 1M | 2026-09-01 |
| [Claude Opus 5.5](https://dotsagent.io/zh/models/claude-opus-5-5)claude-opus-5-5 | [Anthropic](https://dotsagent.io/zh/providers/anthropic) | US$4 | US$0.20 | US$20 | 1M | 2026-09-22 |
| [Kimi K3](https://dotsagent.io/zh/models/kimi-k3)kimi-k3 | [Moonshot AI (Kimi)](https://dotsagent.io/zh/providers/moonshot) | US$3 | US$0.30 | US$15 | 1.05M | 2026-07-16 |
| [Command A+](https://dotsagent.io/zh/models/command-a-plus)第三方价格command-a-plus-05-2026 | [Cohere](https://dotsagent.io/zh/providers/cohere) | US$2.50 | — | US$10 | 128K | 2026-05-20 |
| [GPT-6.1 Sol](https://dotsagent.io/zh/models/gpt-6-1-sol)gpt-6.1-sol | [OpenAI](https://dotsagent.io/zh/providers/openai) | US$2* | US$0.10 | US$10 | 1.05M | 2026-09-29 |
| [Claude Sonnet 5.5](https://dotsagent.io/zh/models/claude-sonnet-5-5)claude-sonnet-5-5 | [Anthropic](https://dotsagent.io/zh/providers/anthropic) | US$2 | US$0.20 | US$10 | 1M | 2026-09-28 |
| [Gemini 3.1 Pro Preview](https://dotsagent.io/zh/models/gemini-3-1-pro-preview)预览版gemini-3.1-pro-preview | [Google](https://dotsagent.io/zh/providers/google) | US$2* | US$0.20 | US$12 | 1.05M | 2026-02-19 |
| [Grok 4.7](https://dotsagent.io/zh/models/grok-4-7)grok-4.7 | [xAI](https://dotsagent.io/zh/providers/xai) | US$2* | US$0.50 | US$6 | 500K | 2026-09-21 |
| [Qwen3.8-Max](https://dotsagent.io/zh/models/qwen3-8-max)qwen3.8-max | [Alibaba Cloud (Qwen)](https://dotsagent.io/zh/providers/qwen) | US$1.65 | US$0.21 | US$4.95 | 1M | 2026-08-03 |
| [Gemini 3.5 Flash](https://dotsagent.io/zh/models/gemini-3-5-flash)gemini-3.5-flash | [Google](https://dotsagent.io/zh/providers/google) | US$1.50 | US$0.15 | US$9 | 1.05M | 2026-05-19 |
| [Mistral Medium 3.5](https://dotsagent.io/zh/models/mistral-medium-3-5)mistral-medium-3-5 | [Mistral AI](https://dotsagent.io/zh/providers/mistral) | US$1.50 | — | US$7.50 | 256K | 2026-04-28 |
| [GLM-5.3](https://dotsagent.io/zh/models/glm-5-3)glm-5.3 | [Z.ai (Zhipu)](https://dotsagent.io/zh/providers/zhipu) | US$1.40 | US$0.26 | US$4.40 | 1M | 2026-08-14 |
| [DeepSeek-V4-Pro-0813](https://dotsagent.io/zh/models/deepseek-v4-pro-0813)deepseek-v4-pro | [DeepSeek](https://dotsagent.io/zh/providers/deepseek) | US$1.32 | US$0.044 | US$3.96 | 1M | 2026-08-13 |
| [Grok 4.3](https://dotsagent.io/zh/models/grok-4-3)grok-4.3 | [xAI](https://dotsagent.io/zh/providers/xai) | US$1.25* | US$0.20 | US$2.50 | 1M | — |
| [Muse Spark 1.3](https://dotsagent.io/zh/models/muse-spark-1-3)muse-spark-1.3 | [Meta](https://dotsagent.io/zh/providers/meta) | US$1.25 | US$0.15 | US$4.25 | 1M | 2026-09-02 |
| [Amazon Nova 2 Pro](https://dotsagent.io/zh/models/amazon-nova-2-pro)预览版第三方价格 | [Amazon](https://dotsagent.io/zh/providers/amazon) | US$1.25 | — | US$10 | — | — |
| [Claude Haiku 4.5](https://dotsagent.io/zh/models/claude-haiku-4-5)claude-haiku-4-5-20251001 | [Anthropic](https://dotsagent.io/zh/providers/anthropic) | US$1 | US$0.10 | US$5 | 200K | — |
| [Grok Build 0.1](https://dotsagent.io/zh/models/grok-build-0-1)grok-build-0.1 | [xAI](https://dotsagent.io/zh/providers/xai) | US$1* | US$0.20 | US$2 | 256K | — |
| [Gemini 3.8 Flash](https://dotsagent.io/zh/models/gemini-3-8-flash)gemini-3.8-flash | [Google](https://dotsagent.io/zh/providers/google) | US$0.75 | US$0.075 | US$3.75 | 1.05M | 2026-09-02 |
| [Gemini 3.5 Flash-Lite](https://dotsagent.io/zh/models/gemini-3-5-flash-lite)gemini-3.5-flash-lite | [Google](https://dotsagent.io/zh/providers/google) | US$0.30 | US$0.030 | US$2.50 | 1.05M | 2026-07-21 |
| [DeepSeek-V4.1-Flash](https://dotsagent.io/zh/models/deepseek-v4-1-flash)deepseek-flash | [DeepSeek](https://dotsagent.io/zh/providers/deepseek) | US$0.30 | US$0.006 | US$1.20 | 1M | 2026-09-10 |
| [Amazon Nova 2 Lite](https://dotsagent.io/zh/models/amazon-nova-2-lite)第三方价格 | [Amazon](https://dotsagent.io/zh/providers/amazon) | US$0.30 | — | US$2.50 | 1M | 2025-12-02 |
| [Qwen3.7-Plus](https://dotsagent.io/zh/models/qwen3-7-plus)qwen3.7-plus | [Alibaba Cloud (Qwen)](https://dotsagent.io/zh/providers/qwen) | US$0.28* | US$0.056 | US$1.10 | 1M | 2026-05-26 |
| [Mistral Small 4](https://dotsagent.io/zh/models/mistral-small-4)mistral-small-2603 | [Mistral AI](https://dotsagent.io/zh/providers/mistral) | US$0.15 | — | US$0.60 | 256K | 2026-03-16 |
| [GLM-5.3-Flash](https://dotsagent.io/zh/models/glm-5-3-flash)glm-5.3-flash | [Z.ai (Zhipu)](https://dotsagent.io/zh/providers/zhipu) | US$0.15 | US$0.030 | US$0.50 | 1M | — |
| [GPT-6 Luna](https://dotsagent.io/zh/models/gpt-6-luna)gpt-6-luna | [OpenAI](https://dotsagent.io/zh/providers/openai) | US$0.10* | US$0.010 | US$0.50 | 1.05M | 2026-09-22 |
| [Qwen3.7-Flash](https://dotsagent.io/zh/models/qwen3-7-flash)qwen3.7-flash | [Alibaba Cloud (Qwen)](https://dotsagent.io/zh/providers/qwen) | US$0.028* | US$0.006 | US$0.11 | 1M | 2026-07-15 |

每百万 token 的美元价格；标准级别、短 prompt 费率 · * 长 prompt 费用更高

[估算你自己的工作负载](https://dotsagent.io/zh/cost)

## 如何看懂这些价格

### 缓存输入价格最值得关注

在 Agent 循环中，每次调用都会重新发送 system prompt、工具定义和之前的对话。当前缀从缓存中读取时，服务商只收取常规输入价格的一部分，通常不到十分之一。因此，缓存价格低的模型可能比标价更低的模型更划算。

### 长 prompt 可能让整次调用涨价

OpenAI、Google 的 Pro 预览版、xAI 和 Qwen 会在单个 prompt 超过 200K 或 272K token 等阈值后提高价格，而且这次请求中的每个 token 都按更高价格计费。目前，Anthropic 对其完整的百万 token 上下文窗口按标准价格收费。

### 推理 token 按输出计费

回答前会进行思考的模型，其隐藏推理过程按输出价格计费。常规步骤可调低推理强度设置，把高强度设置留给规划或复杂调试。

### 批处理和非高峰时段折扣通常不适合 Agent

批处理 endpoint 价格减半，但结果可能要几小时才能返回，因此更适合评测和历史数据回填，而非实时 Agent。DeepSeek 的非高峰时段价格是个例外，适用于定时任务。

## 开发者常见问题

### 哪个模型最适合低成本运行 AI Agent？

在我们的 30 步编码工作负载中，成本最低的是小型模型：GPT-6 Luna、Qwen3.7-Flash 和 GLM-5.3-Flash，每个任务约需 10 到 20 美分。单步便宜不代表整个任务便宜：如果模型需要更多步骤或重试，成本可能更高。建议用自己的调用记录测试两三个候选模型。

### 为什么我的 Agent 费用远高于按 token 单价估算的金额？

因为每一步都会重新发送不断增长的对话内容。一个起始 prompt 有 25K token 的 30 步任务，总输入量会远超 300 万 token。与更换模型相比，使用 prompt 缓存、缩短工具输出和总结较早的对话轮次更能降低成本。

### 这些是最终价格吗？

这些是各服务商定价页面在 2026 年 10 月 1 日公布的美元标价，不含税费、地区附加费或协商折扣。标注为第三方数据的价格无法从厂商自己的页面获取。

### 缓存读取和缓存写入有什么区别？

缓存读取是读取缓存中已有 token 时享受的折扣价格。一些服务商（包括 Anthropic）还会在前缀首次写入缓存时收取额外费用。表格列出的是读取价格；适用的写入价格会在各模型页面中注明。

为 AI agent 开发者提供的独立参考资料。与此处提及的任何厂商均无关联。

© 2026 DotsAgent · 事实核查日期：2026年10月1日
