AI explained. Tools tested. Ideas worth trying.By Sam Taylor & Samwise ↗

Ox Alpha was GLM-5.3-Flash the whole time. Z.ai just dropped the MIT weights.

Jump to a section
See the story at a glance
$0.075
per million input tokens
MIT license · Z.ai GLM-5.3-Flash

Z.ai released GLM-5.3-Flash on August 26. If the name doesn't immediately ring a bell, you may know it as Ox Alpha — the model Z.ai had been running on OpenRouter for several weeks without telling anyone what it actually was.

Now the name is official. The weights are on Hugging Face under MIT. The price is $0.075 per million input tokens, $0.25 per million output, with a 50% launch discount that runs through September 9.

The architecture: 320 billion total parameters, 18 billion active per token, mixture-of-experts, natively multimodal — text, images, video — and a 1,048,576-token context window. Trained on a 30-trillion-token corpus.

Source spread

Pros & cons

What's genuinely interesting:

  • MIT license. Not Apache 2.0, not a custom "commercial-friendly" license with restrictions that only matter when you read the fine print — MIT. For a 320B model, that's as permissive as it gets.
  • 18B active parameters is efficient for the scale claimed. A 320B MoE activating 18B per token should run meaningfully cheaper than a 320B dense model at the same pricing tier.
  • The Ox Alpha stealth period generated real signal before the brand dropped. Builders used it for weeks without knowing who made it. No preconceptions attached to the GLM name during that window.
  • $0.075/M input is at the commodity flash tier, competing against models that cost 10-40x more per token in the quality tier they're claiming to match.

What deserves scrutiny:

  • Benchmark comparisons to Claude Opus 4.8 and GPT-5.6 Terra are first-party. Z.ai ran the evals. Until independent reproduction exists, this is a marketing comparison, not a verified capability claim.
  • The 50% launch discount expires September 9. Post-discount pricing (roughly $0.15/M input, $0.50/M output) is still cheap by Western-lab standards, but the economics look different from the launch numbers.
  • GLM-5.3 — the main model — remains held back from open weights pending a cybersecurity safety review. The flash variant clearing while the full model waits suggests either the lighter model's offensive security capabilities are below the threshold that concerned Z.ai, or the full model's review is genuinely taking longer. Worth knowing which.
  • Context: GLM-5.3 was held after post-training produced unintended exploit-chain reasoning — finding 1,097 critical vulnerabilities in Linux, WebKit, and FreeBSD. The flash model's capability profile on that dimension is uncharacterized.
$0.075
GLM-5.3-Flash input price per million tokens — 50% promo through September 9, 2026

→ Source: OpenRouter

GLM-5.3-Flash vs comparable models (August 2026 pricing)
ModelInput ($/M)Output ($/M)ContextLicense
GLM-5.3-Flash (promo)$0.075$0.251M tokensMIT
GLM-5.3-Flash (post-Sept 9)~$0.15~$0.501M tokensMIT
GPT-5.6 Terra$1.00$4.00128K tokensProprietary
Claude Opus 4.8 Fast$3.00$15.00200K tokensProprietary
DeepSeek V4 Flash 0731$0.14$0.28128K tokensMIT

Samwise's take

What builders need to know

  • Run your own evals on your actual workload before assuming Claude Opus 4.8 parity. The comparison is first-party from Z.ai.
  • The MIT license means commercial use, modification, fine-tuning, and deployment without royalties. That's the most permissive option in this capability tier.
  • The promo discount expires September 9 — build your cost models against the real rate (~$0.15/M input post-discount), not the launch numbers.
  • Ox Alpha impressions from the stealth period are publicly available on OpenRouter forums and builder communities. Search for those before your own evals — real feedback from an unbiased evaluation window.
  • The full GLM-5.3 weights are still unreleased pending a safety review. If the flash variant works for your use case, you're good. If you need the fuller model's capabilities, you're waiting.

Further reading

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.

Keep following the thread

A little more to explore.

All articles ↗
Open source6 min read

Reflection's first model is competitive. It's fifth on Terminal-Bench. The compute-efficiency argument is the interesting part.

Reflection AI released Beam on October 5 — a 501B-parameter MoE with 23B active, pretrained on 23.8T tokens, with a 1M context window. Per Reflection's own benchmarks: Terminal Bench v2.1 at 80.1% (fifth on the September 2026 leaderboard), SWE-bench Verified at 80.9%, GPQA Diamond at 90.5%. Apache 2.0 weights are coming later in October. No date confirmed yet.

Read the story ↗
Open source5 min read

Mistral's 1T open-weight bet is real. Third on coding is the honest answer.

Mistral released Large 4 on October 6 — 1.05T total parameters, 49B active, natively multimodal, API-only for now. The coding benchmarks put it third among open-weight models on DeepSWE v1.1 at 61.7%, behind Kimi K3 and DeepSeek V4.1 Flash. Open weights land October 27 after a three-week government and cybersecurity review.

Read the story ↗
Open source5 min read

Tencent's Hy4 Preview is 770B parameters. Only 6% fire on any given token.

Tencent open-sourced Hy4 Preview on August 28: 770B total parameters, 49B active per token inference, 1M context window, Apache 2.0 license, $0.834/M input on OpenRouter. The model is real and the benchmarks are competitive. The first-party-only evaluation status is the only caveat worth carrying.

Read the story ↗