Z.ai launched a surprise low-cost model that tightens an ongoing race to undercut inference prices in China.
The Information reported the release and framed it as heightening pressure on rivals, and Z.ai, a Chinese developer of GLM/ChatGLM series models, has spent the last year pushing cheaper tiers to win developers The Information.
That strategy now reads as a full ladder from free-tier flash variants through aggressively discounted paid models to higher-priced flagships, a move that has driven token prices down across the market and shifted comparisons from model size to cost and real-world coding and reasoning performance.
Z.ai’s lineup includes flash versions such as GLM-5.3-Flash and a free GLM-4.7-Flash variant aimed at coding and high-throughput tasks, and the company has highlighted running newer models on domestically produced semiconductors to cut infrastructure costs.
The effect is tactical: peers such as DeepSeek, Moonshot and MiniMax are responding with their own cuts, and benchmarking and usage trackers now play a larger role in measuring competitiveness than raw top-line benchmarks.
Z.ai is also running limited-time pricing promotions on some GLM-5.x flash builds through early September, keeping the pricing pressure on the rest of the field.