Alibaba Cloud (Qwen) · Small and fast
Qwen3.7-Flash API pricing and specs
The lowest-priced model in this table, with thinking and non-thinking modes and context caching.
$0.028
$0.006
$0.11
Pricing steps up with prompt length: above 32K tokens the rate triples, and above 256K it rises again to $0.165 input and $0.66 output per million.
Prompts over 32K tokens are billed at $0.083 input, $0.017 cached and $0.33 output per million, and the higher rate applies to the whole request.
Spec sheet
- API model string
qwen3.7-flash- Provider
- Alibaba Cloud (Qwen)
- Context
- 1M tokens
- Max output
- 131K tokens
- Released
- July 15, 2026
- Batch API
- 50% off, results within hours
- Documented agent features
- Tool callingAdjustable reasoning
What a task costs
Four reference workloads priced at this model's list rates, with caching applied after the first call. Your numbers will differ; the calculator lets you change every assumption.
| Workload | Per task | Per 1,000 tasks |
|---|---|---|
| Support bot | $0.001 | $1 |
| Coding agent 28 calls at the long-prompt rate | $0.101 | $101 |
| Browser agent 18 calls at the long-prompt rate | $0.060 | $60 |
| Research agent 12 calls at the long-prompt rate | $0.054 | $54 |
Other Alibaba Cloud (Qwen) models
Similar cost elsewhere
Models from other providers whose cost on the 30-step coding workload is closest to this one.