Zhipu's GLM-5.3-Flash Undercuts Claude and GPT on Price, Not on Hardware

Zhipu's GLM-5.3-Flash Undercuts Claude and GPT on Price, Not on Hardware
View on original source
Category: SciTech
Share
Archive
Like
Zhipu AI's new open-weight model, GLM-5.3-Flash, undercuts Claude Opus and GPT-5.6 on price and rivals them on coding benchmarks, but running it yourself takes a lot more than one gaming GPU. For a week in August, developers on OpenRouter and OpenCode could not stop talking about an anonymous model called Ox Alpha. It matched Claude Opus 4.8 on agentic coding tasks. It cost almost nothing to call. And nobody knew who had built it. On August 26, 2026, Zhipu AI ended the guessing game: Ox Alpha was GLM-5.3-Flash, and the company published the full model weights on Hugging Face under an MIT license the same day. Cheap, and it can code The model is a 320-billion-parameter mixture-of-experts system that activates only 18 billion parameters per token. That's the whole trick behind its price. Zhipu, which also goes by Z.ai, set API pricing at 15 cents per million input tokens and 50 cents per million output tokens, with a launch promo cutting input costs to 7.5 cents through September 9. According to VentureBeat, a full million-token round trip on GLM-5.3-Flash runs about $5.80, against $30 for Claude Opus 5 and $35 for GPT-5.6 Sol. The benchmarks back up the hype, at least the ones Zhipu chose to publish. On its own software-engineering agent benchmark, GLM-5.3-Flash scored 63.4, ahead of Claude Opus 4.8 at 58.0 and DeepSeek-V4-Vision-Exp at 59.3. On the separate KingBench evaluation, it scored 63 out of 80, just behind Opus 4.8 but ahead of Opus 5 and DeepSeek V4 Pro. Developers noticed. OpenRouter says the anonymous Ox Alpha endpoint pulled in more than 11 trillion tokens in its first three days: the platform's biggest launch ever, before anyone knew whose model it was. The catch: it doesn't run on your gaming GPU It is also the first natively multimodal release in the GLM-5 line, handling image and video input alongside a context window that stretches to just over a million tokens. Zhipu says the whole stealth run, all 62 trillion tokens of it, was served on a cluster of 100,000 domestically produced chips: a claim meant to show that Chinese inference hardware can carry a global-scale launch without Nvidia. CNBC and other outlets could not independently verify that chip count, and Zhipu has not named the hardware. Treat that number as a marketing beat, not a receipt. The pitch that made GLM-5.3-Flash spread among founders was not really the price. It was the idea that a frontier-grade model could run on hardware you already own. That's not quite right. Full BF16 precision needs around 740GB of memory. Even the most aggressive quantization, a 1-bit build, still needs 90 to 128GB of combined RAM and VRAM, according to hardware guides from Unsloth and Zima Space. A practical 4-bit version, the kind most teams would actually deploy, needs roughly 190GB. None of that fits on a single consumer GPU, no matter how the marketing reads. What it actually threatens What GLM-5.3-Flash actually threatens is not the GPU aisle. It is the API subscription. A startup that has been paying Anthropic or OpenAI by the token can now get comparable coding performance for a sixth of the price, hosted, without touching a rack of hardware at all. Frankly, that is the real disruption here, and it is not do-it-yourself inference for solo builders. It is a credible discount API sitting one key away from every product currently wired to GPT or Claude. Zhipu is not alone in making that pitch. Tencent open-sourced a 770-billion-parameter model earlier this month aimed at the same crowd of builders wary of vendor lock-in, and the MIT license on GLM-5.3-Flash goes further than most closed-weight rivals are willing to go. You don't need Zhipu's permission to fine-tune it, resell access to it, or run it behind your own product with no attribution required. Zhipu's Hong Kong-listed shares closed more than 12% higher, at HK$1,160, the day the Ox Alpha reveal went public. Whether that rally holds depends less on the chip story than on whether GLM-5.3-Flash keeps beating its price tag on real coding work, not on a benchmark deck Zhipu wrote about itself. That's the real test. Also read: SK Hynix Is Weighing a Japan Memory Fab to Keep Up With AI Demand • OpenAI Is Buying So Many Mac Minis and Studios That Apple Can't Keep Up • Anthropic Sued Over Claude Max Plans That Deliver Far Less Than Advertised

(0)Comments

 

A note on cookies

Newshunt uses essential cookies to keep you signed in and to remember your language and country, so the site works the way you expect. With your permission, we'd also like to use analytics cookies to understand how people use Newshunt and improve it over time.

Accepting only affects analytics. To learn more, view our Privacy Policy or Terms & Conditions.