dotsagent.io
Language:English
Z.ai (Zhipu) · Small and fast

GLM-5.3-Flash API pricing and specs

Z.ai's small, low-cost GLM model.

Input
⁨$0.15⁩
Cached input
⁨$0.030⁩
Output
⁨$0.50⁩

US dollars per million tokens, standard tier, short-prompt rate

The launch discount ended on 9 September 2026. A faster FlashX variant costs $0.37 input and $1.25 output per million.

Spec sheet

API model string
glm-5.3-flash
Provider
Z.ai (Zhipu)
Context
1M tokens
Max output
—
Released
—
Batch API
No batch discount listed
Documented agent features
The vendor's page lists none for this model.

What a task costs

Four reference workloads priced at this model's list rates, with caching applied after the first call. Your numbers will differ; the calculator lets you change every assumption.

WorkloadPer taskPer 1,000 tasks
Support bot
4 calls, short tool lookups
⁨$0.004⁩⁨$4⁩
Coding agent
30 calls, files and test output
⁨$0.182⁩⁨$182⁩
Browser agent
25 calls, screenshots and page text
⁨$0.115⁩⁨$115⁩
Research agent
15 calls, long search results
⁨$0.101⁩⁨$101⁩

Change the assumptions

Other Z.ai (Zhipu) models

Similar cost elsewhere

Models from other providers whose cost on the 30-step coding workload is closest to this one.

Sources

  1. docs.z.ai/guides/overview/pricing
  2. docs.z.ai/guides/overview/overview

Independent reference for people who build AI agents. Not affiliated with any vendor named here.

© 2026 DotsAgent · Facts checked October 1, 2026