Vol. 1 · Edition 035Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

$0.075
per million input tokens
MIT license · Z.ai GLM-5.3-Flash
Open Source
By Sam Taylor with Samwise

On 320B-A18B MoE, the Ox Alpha stealth period, MIT weights on Hugging Face, and $0.075/M that doesn't make Western-lab pricing look great

Ox Alpha was GLM-5.3-Flash the whole time. Z.ai just dropped the MIT weights.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

02

Pro (hyped)

01

← Anti-AI · Pro-AI →

Z.ai released GLM-5.3-Flash on August 26. If the name doesn't immediately ring a bell, you may know it as Ox Alpha — the model Z.ai had been running on OpenRouter for several weeks without telling anyone what it actually was.

Now the name is official. The weights are on Hugging Face under MIT. The price is $0.075 per million input tokens, $0.25 per million output, with a 50% launch discount that runs through September 9.

The architecture: 320 billion total parameters, 18 billion active per token, mixture-of-experts, natively multimodal — text, images, video — and a 1,048,576-token context window. Trained on a 30-trillion-token corpus.

Source spread

Pros & cons

What's genuinely interesting:

  • MIT license. Not Apache 2.0, not a custom "commercial-friendly" license with restrictions that only matter when you read the fine print — MIT. For a 320B model, that's as permissive as it gets.
  • 18B active parameters is efficient for the scale claimed. A 320B MoE activating 18B per token should run meaningfully cheaper than a 320B dense model at the same pricing tier.
  • The Ox Alpha stealth period generated real signal before the brand dropped. Builders used it for weeks without knowing who made it. No preconceptions attached to the GLM name during that window.
  • $0.075/M input is at the commodity flash tier, competing against models that cost 10-40x more per token in the quality tier they're claiming to match.

What deserves scrutiny:

  • Benchmark comparisons to Claude Opus 4.8 and GPT-5.6 Terra are first-party. Z.ai ran the evals. Until independent reproduction exists, this is a marketing comparison, not a verified capability claim.
  • The 50% launch discount expires September 9. Post-discount pricing (roughly $0.15/M input, $0.50/M output) is still cheap by Western-lab standards, but the economics look different from the launch numbers.
  • GLM-5.3 — the main model — remains held back from open weights pending a cybersecurity safety review. The flash variant clearing while the full model waits suggests either the lighter model's offensive security capabilities are below the threshold that concerned Z.ai, or the full model's review is genuinely taking longer. Worth knowing which.
  • Context: GLM-5.3 was held after post-training produced unintended exploit-chain reasoning — finding 1,097 critical vulnerabilities in Linux, WebKit, and FreeBSD. The flash model's capability profile on that dimension is uncharacterized.
$0.075
GLM-5.3-Flash input price per million tokens — 50% promo through September 9, 2026

→ Source: OpenRouter

GLM-5.3-Flash vs comparable models (August 2026 pricing)
ModelInput ($/M)Output ($/M)ContextLicense
GLM-5.3-Flash (promo)$0.075$0.251M tokensMIT
GLM-5.3-Flash (post-Sept 9)~$0.15~$0.501M tokensMIT
GPT-5.6 Terra$1.00$4.00128K tokensProprietary
Claude Opus 4.8 Fast$3.00$15.00200K tokensProprietary
DeepSeek V4 Flash 0731$0.14$0.28128K tokensMIT

Samwise's take

What builders need to know

  • Run your own evals on your actual workload before assuming Claude Opus 4.8 parity. The comparison is first-party from Z.ai.
  • The MIT license means commercial use, modification, fine-tuning, and deployment without royalties. That's the most permissive option in this capability tier.
  • The promo discount expires September 9 — build your cost models against the real rate (~$0.15/M input post-discount), not the launch numbers.
  • Ox Alpha impressions from the stealth period are publicly available on OpenRouter forums and builder communities. Search for those before your own evals — real feedback from an unbiased evaluation window.
  • The full GLM-5.3 weights are still unreleased pending a safety review. If the flash variant works for your use case, you're good. If you need the fuller model's capabilities, you're waiting.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.