dotsagent.io
Language:English

Qwen3.8-Max API pricing and specs

The first open-weight Qwen Max model, with thinking and non-thinking modes and up to 262K tokens of reasoning.

Input
⁨$1.65⁩
Cached input
⁨$0.21⁩
Output
⁨$4.95⁩

US dollars per million tokens, standard tier, short-prompt rate

Explicit cache reads cost $0.137 per million and cache creation $2.063. The Singapore region charges $2 input and $6 output. Batch pricing is offered only in the Beijing region.

Spec sheet

API model string
qwen3.8-max
Provider
Alibaba Cloud (Qwen)
Context
1M tokens
Max output
131K tokens
Released
August 3, 2026
Batch API
No batch discount listed
Documented agent features
Tool callingStructured outputBuilt-in web searchAdjustable reasoningImage input

What a task costs

Four reference workloads priced at this model's list rates, with caching applied after the first call. Your numbers will differ; the calculator lets you change every assumption.

WorkloadPer taskPer 1,000 tasks
Support bot
4 calls, short tool lookups
⁨$0.039⁩⁨$39⁩
Coding agent
30 calls, files and test output
⁨$1.63⁩⁨$1,629⁩
Browser agent
25 calls, screenshots and page text
⁨$1.15⁩⁨$1,153⁩
Research agent
15 calls, long search results
⁨$1.04⁩⁨$1,037⁩

Change the assumptions

Other Alibaba Cloud (Qwen) models

Similar cost elsewhere

Models from other providers whose cost on the 30-step coding workload is closest to this one.

Sources

  1. alibabacloud.com/help/en/model-studio/qwen3-8-max
  2. alibabacloud.com/en/press-room/alibaba-unveils-qwen3-8-max
  3. qwen.ai/blog?id=qwen3.8

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026