AI & Automation

Ox Alpha Revealed as GLM-5.3-Flash: Specs and Real Cost

OpenRouter's stealth model Ox Alpha was Z.ai's GLM-5.3-Flash: 320B parameters, 18B active, 1M context, MIT weights. What the free window actually cost users.

Need AI integrated into your ERP, website, WhatsApp, CRM or internal systems?

By Manikya Searathna
OpenRouter stealth model Ox Alpha revealed as Z.ai GLM-5.3-Flash

On 20 August 2026 a model with no listed maker appeared on OpenRouter under the identifier stealth/ox-alpha. It was free, it was fast, and within days it was one of the most-used models on the platform. On 26 August it was revealed as GLM-5.3-Flash from Z.ai. The specifications are genuinely impressive. The part worth thinking about is what it meant to send your data to a model whose vendor you did not know.

What GLM-5.3-Flash actually is

  • 320 billion total parameters, sparse mixture-of-experts, with only 18 billion active per token.
  • Context window of 1,310,720 tokens interactive, and 1,048,575 tokens for batch.
  • Described as the first natively multimodal model in the GLM-5 series: text, image and video in, text out.
  • Hybrid linear and sparse attention, aimed at holding long-context accuracy down while cutting compute.
  • MIT-licensed weights, published on Hugging Face on the day of the reveal.

The MIT licence is the most consequential detail and the one most easily skipped. Genuinely permissive weights for a 320B multimodal model mean you can self-host, audit and fine-tune it without a negotiated agreement — which is a different proposition from any closed frontier model, whatever the benchmark table says.

The pricing, and why the free window ended

During the stealth period the model was free. After the reveal, interactive pricing ran at a promotional $0.075 per million input tokens and $0.25 per million output tokens through 9 September, moving to list pricing of $0.15 and $0.50 after that. Batch was already at list price. Cache reads are $0.015 per million.

Set that against a frontier model and the gap is the story. Claude Fable 5.1 lists at $10 per million input and $50 per million output. At list price, GLM-5.3-Flash is roughly sixty-seven times cheaper on input and one hundred times cheaper on output. For high-volume, lower-stakes work — bulk classification, first-pass extraction, log summarisation — that difference decides the architecture, not the preference.

How good is it, honestly

Vendor-run benchmarks put it close to the frontier without reaching it. On Terminal-Bench 2.1, GLM-5.3-Flash scored 84.3 against 87.4 for GPT-5.6 Terra and 85.0 for Claude Opus 4.8, while leading Opus 4.8 on roughly half of the published sub-benchmarks. These are vendor numbers on a model the vendor selected the tests for, so treat them as an upper bound rather than a verdict.

Be sceptical of the usage figures circulating too. OpenRouter's one confirmed statement was that the model was on track to hit nearly 6 trillion tokens on 24 August alone. Cumulative totals quoted elsewhere vary by around 500% between sources with no traceable primary attribution — a reminder that stealth launches generate more excitement than record-keeping.

The part that matters for Sri Lankan businesses

For six days, a large number of teams sent production prompts to a model whose vendor, hosting jurisdiction, and data-retention terms were all undisclosed. That is the actual risk in a stealth launch, and it has nothing to do with model quality.

Under Sri Lanka's Personal Data Protection Act, if you process personal data you need to be able to say who processes it and where. “An anonymous model on a routing platform” is not an answer you can give a regulator, a hospital client, or a bank's procurement team. If customer records, patient data or payroll passed through stealth/ox-alpha during that window, that is a disclosure question now, not a hypothetical.

Sri Lanka PDPA considerations for AI processing of customer data
An undisclosed vendor is an undisclosed processor — a PDPA problem regardless of how good the model is.

Why vendors run stealth launches at all

The mechanics are straightforward: an anonymous window gathers real usage feedback and builds an audience before identity, licensing and pricing are attached, converting free-tier goodwill into paying deployments at reveal. It is effective marketing. It is also a period during which users cannot perform the diligence they would normally perform, which is why the two facts belong in the same sentence.

A workable policy

  1. Allow-list models by name for anything touching personal or commercially sensitive data. A stealth identifier cannot clear that bar by definition.
  2. Route stealth and preview models to synthetic or already-public data only, so you can still evaluate them without exposure.
  3. Record which model handled which workload. When a reveal happens, you want to answer the question in minutes.
  4. Re-benchmark once pricing is attached. A model that was compelling while free may not be at list price.
  5. For genuinely open weights like these, price self-hosting against per-token cost before assuming the API is cheaper at your volume.

Where Capricon fits

Capricon builds AI into ERP, customer operations and document processing for Sri Lankan businesses, which means choosing models on cost, accuracy and data-handling terms together rather than one at a time. If you want a model routing policy that will survive a client security review, see AI integration services or contact us.

Frequently asked questions

What was Ox Alpha on OpenRouter?

Ox Alpha was a stealth model that appeared on OpenRouter as stealth/ox-alpha on 20 August 2026 with no listed maker. On 26 August 2026 it was revealed as GLM-5.3-Flash from Z.ai — a 320-billion-parameter sparse mixture-of-experts model with 18 billion active parameters per token, a roughly 1M-token context window, and MIT-licensed weights.

How much does GLM-5.3-Flash cost compared to frontier models?

List pricing is $0.15 per million input tokens and $0.50 per million output tokens, after a launch promotion of $0.075 and $0.25. Claude Fable 5.1 lists at $10 and $50, so at list price GLM-5.3-Flash is roughly 67 times cheaper on input and 100 times cheaper on output — which makes high-volume, lower-stakes work economically viable in a way frontier pricing does not.

Is it safe to use stealth models on OpenRouter for business data?

Not for personal or commercially sensitive data. During a stealth window the vendor, hosting jurisdiction and data-retention terms are undisclosed, so you cannot identify your processor — which under Sri Lanka's PDPA is information you must be able to provide. Restrict stealth and preview models to synthetic or already-public data, and allow-list named models for anything else.

Related Capricon solutions

Explore tools and services for ai & automation

Related guides on this topic

Related Capricon product & services

Ready to take your business to the next level?

Your next big move starts here - take charge, scale up, and lead your business to success.