dotsagent.io
Language:English
Alibaba Cloud (Qwen) · Small and fast

Qwen3.7-Flash API pricing and specs

The lowest-priced model in this table, with thinking and non-thinking modes and context caching.

Input
⁨$0.028⁩
Cached input
⁨$0.006⁩
Output
⁨$0.11⁩

US dollars per million tokens, standard tier, short-prompt rate

Pricing steps up with prompt length: above 32K tokens the rate triples, and above 256K it rises again to $0.165 input and $0.66 output per million.

Prompts over 32K tokens are billed at ⁨$0.083⁩ input, ⁨$0.017⁩ cached and ⁨$0.33⁩ output per million, and the higher rate applies to the whole request.

Spec sheet

API model string
qwen3.7-flash
Provider
Alibaba Cloud (Qwen)
Context
1M tokens
Max output
131K tokens
Released
July 15, 2026
Batch API
50% off, results within hours
Documented agent features
Tool callingAdjustable reasoning

What a task costs

Four reference workloads priced at this model's list rates, with caching applied after the first call. Your numbers will differ; the calculator lets you change every assumption.

WorkloadPer taskPer 1,000 tasks
Support bot
4 calls, short tool lookups
⁨$0.001⁩⁨$1⁩
Coding agent
30 calls, files and test output
28 calls at the long-prompt rate
⁨$0.101⁩⁨$101⁩
Browser agent
25 calls, screenshots and page text
18 calls at the long-prompt rate
⁨$0.060⁩⁨$60⁩
Research agent
15 calls, long search results
12 calls at the long-prompt rate
⁨$0.054⁩⁨$54⁩

Change the assumptions

Other Alibaba Cloud (Qwen) models

Similar cost elsewhere

Models from other providers whose cost on the 30-step coding workload is closest to this one.

Sources

  1. alibabacloud.com/help/en/model-studio/qwen3-7-flash
  2. help.aliyun.com/en/model-studio/model-pricing

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026