---
title: "Giá API LLM cho AI agent: so sánh 28 model · DotsAgent"
description: "Giá input, input từ cache và output trên mỗi triệu token của 28 model từ OpenAI, Anthropic, Google, xAI, DeepSeek, Qwen và nhiều hãng khác, kèm giới hạn ngữ cảnh và tính năng cho agent."
url: https://dotsagent.io/vi/models
---

Xây dựng · giá được kiểm tra ngày 1 October 2026

# Model và giá API dành cho nhà phát triển agent

Agent trả phí cho từng bước nó thực hiện, và phần lớn chi phí là cho input đã gửi trước đó. Bảng này đặt giá niêm yết, mức giảm giá từ cache và giới hạn ngữ cảnh cạnh nhau để bạn chọn model dựa trên chi phí của cả tác vụ, thay vì chỉ nhìn mức giá quảng cáo.

- [DeepSeek-V4.1-Flash](https://dotsagent.io/vi/models/deepseek-v4-1-flash): 0,22 US$

- [Qwen3.7-Plus](https://dotsagent.io/vi/models/qwen3-7-plus): 0,34 US$

- [GPT-6 Luna](https://dotsagent.io/vi/models/gpt-6-luna): 0,10 US$

## Tất cả model trong một bảng

| — | Nhà cung cấp | — | — | — | — | — |
| --- | --- | --- | --- | --- | --- | --- |
| [GPT-6 Astra](https://dotsagent.io/vi/models/gpt-6-astra)gpt-6-astra | [OpenAI](https://dotsagent.io/vi/providers/openai) | 10 US$* | 1 US$ | 50 US$ | 1,05M | 2026-09-08 |
| [Claude Fable 5.1](https://dotsagent.io/vi/models/claude-fable-5-1)claude-fable-5-1 | [Anthropic](https://dotsagent.io/vi/providers/anthropic) | 10 US$ | 0,25 US$ | 50 US$ | 1M | 2026-09-01 |
| [Claude Opus 5.5](https://dotsagent.io/vi/models/claude-opus-5-5)claude-opus-5-5 | [Anthropic](https://dotsagent.io/vi/providers/anthropic) | 4 US$ | 0,20 US$ | 20 US$ | 1M | 2026-09-22 |
| [Kimi K3](https://dotsagent.io/vi/models/kimi-k3)kimi-k3 | [Moonshot AI (Kimi)](https://dotsagent.io/vi/providers/moonshot) | 3 US$ | 0,30 US$ | 15 US$ | 1,05M | 2026-07-16 |
| [Command A+](https://dotsagent.io/vi/models/command-a-plus)giá của bên thứ bacommand-a-plus-05-2026 | [Cohere](https://dotsagent.io/vi/providers/cohere) | 2,50 US$ | — | 10 US$ | 128K | 2026-05-20 |
| [GPT-6.1 Sol](https://dotsagent.io/vi/models/gpt-6-1-sol)gpt-6.1-sol | [OpenAI](https://dotsagent.io/vi/providers/openai) | 2 US$* | 0,10 US$ | 10 US$ | 1,05M | 2026-09-29 |
| [Claude Sonnet 5.5](https://dotsagent.io/vi/models/claude-sonnet-5-5)claude-sonnet-5-5 | [Anthropic](https://dotsagent.io/vi/providers/anthropic) | 2 US$ | 0,20 US$ | 10 US$ | 1M | 2026-09-28 |
| [Gemini 3.1 Pro Preview](https://dotsagent.io/vi/models/gemini-3-1-pro-preview)bản xem trướcgemini-3.1-pro-preview | [Google](https://dotsagent.io/vi/providers/google) | 2 US$* | 0,20 US$ | 12 US$ | 1,05M | 2026-02-19 |
| [Grok 4.7](https://dotsagent.io/vi/models/grok-4-7)grok-4.7 | [xAI](https://dotsagent.io/vi/providers/xai) | 2 US$* | 0,50 US$ | 6 US$ | 500K | 2026-09-21 |
| [Qwen3.8-Max](https://dotsagent.io/vi/models/qwen3-8-max)qwen3.8-max | [Alibaba Cloud (Qwen)](https://dotsagent.io/vi/providers/qwen) | 1,65 US$ | 0,21 US$ | 4,95 US$ | 1M | 2026-08-03 |
| [Gemini 3.5 Flash](https://dotsagent.io/vi/models/gemini-3-5-flash)gemini-3.5-flash | [Google](https://dotsagent.io/vi/providers/google) | 1,50 US$ | 0,15 US$ | 9 US$ | 1,05M | 2026-05-19 |
| [Mistral Medium 3.5](https://dotsagent.io/vi/models/mistral-medium-3-5)mistral-medium-3-5 | [Mistral AI](https://dotsagent.io/vi/providers/mistral) | 1,50 US$ | — | 7,50 US$ | 256K | 2026-04-28 |
| [GLM-5.3](https://dotsagent.io/vi/models/glm-5-3)glm-5.3 | [Z.ai (Zhipu)](https://dotsagent.io/vi/providers/zhipu) | 1,40 US$ | 0,26 US$ | 4,40 US$ | 1M | 2026-08-14 |
| [DeepSeek-V4-Pro-0813](https://dotsagent.io/vi/models/deepseek-v4-pro-0813)deepseek-v4-pro | [DeepSeek](https://dotsagent.io/vi/providers/deepseek) | 1,32 US$ | 0,044 US$ | 3,96 US$ | 1M | 2026-08-13 |
| [Grok 4.3](https://dotsagent.io/vi/models/grok-4-3)grok-4.3 | [xAI](https://dotsagent.io/vi/providers/xai) | 1,25 US$* | 0,20 US$ | 2,50 US$ | 1M | — |
| [Muse Spark 1.3](https://dotsagent.io/vi/models/muse-spark-1-3)muse-spark-1.3 | [Meta](https://dotsagent.io/vi/providers/meta) | 1,25 US$ | 0,15 US$ | 4,25 US$ | 1M | 2026-09-02 |
| [Amazon Nova 2 Pro](https://dotsagent.io/vi/models/amazon-nova-2-pro)bản xem trướcgiá của bên thứ ba | [Amazon](https://dotsagent.io/vi/providers/amazon) | 1,25 US$ | — | 10 US$ | — | — |
| [Claude Haiku 4.5](https://dotsagent.io/vi/models/claude-haiku-4-5)claude-haiku-4-5-20251001 | [Anthropic](https://dotsagent.io/vi/providers/anthropic) | 1 US$ | 0,10 US$ | 5 US$ | 200K | — |
| [Grok Build 0.1](https://dotsagent.io/vi/models/grok-build-0-1)grok-build-0.1 | [xAI](https://dotsagent.io/vi/providers/xai) | 1 US$* | 0,20 US$ | 2 US$ | 256K | — |
| [Gemini 3.8 Flash](https://dotsagent.io/vi/models/gemini-3-8-flash)gemini-3.8-flash | [Google](https://dotsagent.io/vi/providers/google) | 0,75 US$ | 0,075 US$ | 3,75 US$ | 1,05M | 2026-09-02 |
| [Gemini 3.5 Flash-Lite](https://dotsagent.io/vi/models/gemini-3-5-flash-lite)gemini-3.5-flash-lite | [Google](https://dotsagent.io/vi/providers/google) | 0,30 US$ | 0,030 US$ | 2,50 US$ | 1,05M | 2026-07-21 |
| [DeepSeek-V4.1-Flash](https://dotsagent.io/vi/models/deepseek-v4-1-flash)deepseek-flash | [DeepSeek](https://dotsagent.io/vi/providers/deepseek) | 0,30 US$ | 0,006 US$ | 1,20 US$ | 1M | 2026-09-10 |
| [Amazon Nova 2 Lite](https://dotsagent.io/vi/models/amazon-nova-2-lite)giá của bên thứ ba | [Amazon](https://dotsagent.io/vi/providers/amazon) | 0,30 US$ | — | 2,50 US$ | 1M | 2025-12-02 |
| [Qwen3.7-Plus](https://dotsagent.io/vi/models/qwen3-7-plus)qwen3.7-plus | [Alibaba Cloud (Qwen)](https://dotsagent.io/vi/providers/qwen) | 0,28 US$* | 0,056 US$ | 1,10 US$ | 1M | 2026-05-26 |
| [Mistral Small 4](https://dotsagent.io/vi/models/mistral-small-4)mistral-small-2603 | [Mistral AI](https://dotsagent.io/vi/providers/mistral) | 0,15 US$ | — | 0,60 US$ | 256K | 2026-03-16 |
| [GLM-5.3-Flash](https://dotsagent.io/vi/models/glm-5-3-flash)glm-5.3-flash | [Z.ai (Zhipu)](https://dotsagent.io/vi/providers/zhipu) | 0,15 US$ | 0,030 US$ | 0,50 US$ | 1M | — |
| [GPT-6 Luna](https://dotsagent.io/vi/models/gpt-6-luna)gpt-6-luna | [OpenAI](https://dotsagent.io/vi/providers/openai) | 0,10 US$* | 0,010 US$ | 0,50 US$ | 1,05M | 2026-09-22 |
| [Qwen3.7-Flash](https://dotsagent.io/vi/models/qwen3-7-flash)qwen3.7-flash | [Alibaba Cloud (Qwen)](https://dotsagent.io/vi/providers/qwen) | 0,028 US$* | 0,006 US$ | 0,11 US$ | 1M | 2026-07-15 |

Đô la Mỹ trên mỗi triệu token, phân khúc tiêu chuẩn, mức giá cho prompt ngắn · * prompt dài có giá cao hơn

[Tính giá cho khối lượng công việc của bạn](https://dotsagent.io/vi/cost)

## Cách đọc các mức giá này

### Giá input từ cache là con số đáng chú ý

Trong vòng lặp agent, system prompt, định nghĩa tool và các lượt hội thoại trước được gửi lại trong mỗi lần gọi. Nhà cung cấp tính một phần nhỏ so với giá input thông thường khi phần tiền tố đó được phục vụ từ cache, thường chỉ bằng một phần mười hoặc thấp hơn. Vì vậy, model có giá cache thấp có thể rẻ hơn model có giá niêm yết thấp hơn.

### Prompt dài có thể làm thay đổi giá toàn bộ lần gọi

OpenAI, bản xem trước Pro của Google, xAI và Qwen tính mức giá cao hơn khi một prompt vượt ngưỡng như 200K hoặc 272K token; mức giá cao hơn áp dụng cho mọi token trong yêu cầu đó. Hiện Anthropic tính phí toàn bộ cửa sổ một triệu token theo mức giá tiêu chuẩn.

### Token suy luận được tính phí như output

Các model suy nghĩ trước khi trả lời tính phí phần suy luận ẩn theo mức giá output. Hạ mức nỗ lực cho các bước thông thường và chỉ dùng mức cao cho việc lập kế hoạch hoặc gỡ lỗi khó.

### Giảm giá theo lô và ngoài giờ cao điểm hiếm khi phù hợp với agent

Endpoint xử lý theo lô giảm một nửa giá nhưng trả kết quả trong vài giờ, phù hợp với đánh giá và xử lý dữ liệu tồn đọng hơn là agent đang hoạt động trực tiếp. Mức giá ngoài giờ cao điểm của DeepSeek là ngoại lệ có thể phù hợp với các tác vụ được lên lịch.

## Câu hỏi thường gặp của nhà phát triển

### Model nào rẻ nhất cho AI agent?

Trong khối lượng công việc lập trình gồm 30 bước của chúng tôi, các model nhỏ có giá thấp nhất: GPT-6 Luna, Qwen3.7-Flash và GLM-5.3-Flash, mỗi tác vụ tốn khoảng 10 đến 20 xu. Chi phí mỗi bước thấp không đồng nghĩa chi phí mỗi tác vụ thấp nếu model cần nhiều bước hơn hoặc phải thử lại. Vì vậy, hãy thử hai hoặc ba ứng viên trên các trace của riêng bạn.

### Tại sao agent của tôi tốn nhiều hơn hẳn so với mức giá mỗi token?

Vì mỗi bước đều gửi lại cuộc hội thoại đang dài thêm. Một tác vụ gồm 30 bước với prompt ban đầu dài 25K token sẽ gửi tổng cộng hơn ba triệu token input. Dùng prompt caching, rút gọn output của tool và tóm tắt các lượt hội thoại cũ giúp giảm chi phí nhiều hơn so với đổi model.

### Đây có phải giá cuối cùng không?

Đây là giá niêm yết bằng đô la Mỹ trên trang giá của từng nhà cung cấp vào ngày 1 October 2026, chưa bao gồm thuế, phụ phí theo khu vực hoặc chiết khấu thương lượng riêng. Giá được đánh dấu là của bên thứ ba là những mức giá không thể tra cứu trên trang của chính nhà cung cấp.

### Input từ cache khác gì với ghi cache?

Đọc cache là mức giá đã giảm cho các token có sẵn trong cache. Một số nhà cung cấp, trong đó có Anthropic, cũng tính thêm phí khi ghi một tiền tố vào cache lần đầu. Bảng này hiển thị giá đọc; giá ghi được ghi chú trên trang của từng model nếu có.

Tài liệu tham khảo độc lập dành cho những người xây dựng AI agent. Không liên kết với bất kỳ nhà cung cấp nào được nêu ở đây.

© 2026 DotsAgent · Dữ liệu được kiểm tra ngày 1 tháng 10, 2026
