Z.ai

Glm 5 API Pricing

Glm 5 costs $1.00 per 1M input tokens and $3.20 per 1M output tokens on Z.ai's API. Cached input is billed at $0.2 per 1M tokens. Prices are provider list prices in USD, refreshed as sources update; see the Glm 5 model profile for capability scores.

Input / 1M
$1.00
Output / 1M
$3.20
Effective / 1M I/O
$4.88
Cost Rank
#64

Price List

RatePrice (USD)Notes
Input$1.00per 1M tokens
Output$3.20per 1M tokens
Cached input (read)$0.2per 1M tokens
Blended input + output$4.201M in + 1M out at list price

What a Workload Costs

Computed straight from the list prices above — token counts are illustrative workload sizes, not measurements of Glm 5.

WorkloadCostTokens
Short chat turn$0.002601K in / 500 out
Summarize a long document$0.106100K in / 2K out
Agentic coding session$0.820500K in / 100K out
1M input + 1M output tokens$4.201M in / 1M out

Effective Cost

In AI IQ's scoring, Glm 5's list price is adjusted by a token-usage multiplier of 1.163 (measured from real benchmark runs), giving an effective cost of $4.88 per 1M input + output tokens. That places it #64 of 117 models on the effective-cost ranking (lower is cheaper). Some models spend far more tokens than others on the same task, so effective cost compares what a unit of work really costs. Method details are on the methodology page; see all models on the cost charts and the falling cost of intelligence over time.

Cheaper Alternatives at Similar Capability

Models with a lower effective cost that score within a few IQ points of Glm 5 (IQ 106), or better.

ModelProviderIQEffective Cost / 1M
Muse Spark 1.2Meta122$4.19
muse-spark-1.1Meta119$4.19
kimi-k2.6Kimi119$4.50
kimi-k2.7-codeKimi118$3.73
mimo-v2.5-proXiaomi116$1.80
gemini-3-flashGoogle116$2.74

Compare