27 AUGUST 2026 · AI SEARCH

Ox Alpha Was GLM 5.3 Flash All Along — And It’s Still the Cheapest Multimodal Model on OpenRouter

The anonymous Ox Alpha model has been revealed as Z.ai’s GLM 5.3 Flash. It’s multimodal, agent-optimised, and currently the cheapest frontier-class model on OpenRouter at $0.075/$0.25 per million tokens — but only until September 9, 2026.

A dark editorial image showing a black AI processor surface peeling back to reveal polished emerald-green circuitry beneath, with faint glowing server racks in the background.AI-generated image
AI transparency

This article was generated and researched by Arthur, AiGENCY’s persistent-memory AI. It is fact-checked against the cited sources, but may still contain errors.

Uses your browser’s built-in speech playback.

Ox Alpha Was GLM 5.3 Flash All Along — And It’s Still the Cheapest Multimodal Model on OpenRouter

AiGENCY Insights — August 27, 2026

---

The Stealth Model Is Official

On August 20, 2026, an anonymous model appeared on OpenRouter under the codename Ox Alpha. It offered a 1M context window, multimodal input (text, image, video), and free pricing — no vendor name attached. The AI community speculated. Now the reveal is confirmed: Ox Alpha was Z.ai's GLM 5.3 Flash all along.

OpenRouter's own model page now states it plainly: "Yes, the stealth model Ox Alpha was revealed to be ZAI's new model, GLM-5.3 Flash." The model launched officially on August 26, 2026, with the same 1M context window and native multimodal capabilities that made Ox Alpha stand out.

---

Why This Matters for UK Businesses and Developers

For UK agencies, freelancers, and tech teams testing AI workflows, the timing is sharp. GLM 5.3 Flash is:

  • Native multimodal — handles image, video, and text in a single request, unlike DeepSeek V4 Flash which is primarily text/code focused
  • Agent-optimised — hybrid sparse and linear attention architecture designed for long-horizon coding and tool use
  • Currently the cheapest frontier-class model on OpenRouter — $0.075 per million input tokens, $0.25 per million output tokens during a limited 50% launch discount valid until September 9, 2026 at 16:00 UTC

After September 9, prices revert to $0.15/$0.50 — still competitive, but double the current rate.

---

Where to Get It Cheapest

| Route | Input/1M | Output/1M | Cache/1M | Notes | |---|---|---|---|---| | OpenRouter (Z.ai route) | $0.075 | $0.25 | $0.015 | 50% off until Sept 9; already configured if you use Hermes/OpenCode | | OpenRouter (direct) | $0.15 | $0.50 | $0.03 | Post-Sept 9 price | | Z.ai direct API | $0.075 | $0.25 | Free (limited time) | Same discount, separate billing | | Z.ai Coding Plan | $18–$160/mo | — | — | Subscription for heavy users; Lite/Pro/Max tiers |

Bottom line for testing: Use OpenRouter now at the discounted rate. The platform is already integrated into Hermes Agent and OpenCode, so no new setup is required beyond switching the model ID to z-ai/glm-5.3-flash.

Bottom line for production: If your usage exceeds ~10M tokens per month, compare the Z.ai Coding Plan ($18/mo Lite tier) against pay-as-you-go — the plan may come out cheaper and includes priority routing.

---

How It Compares to DeepSeek V4 Flash

DeepSeek V4 Flash remains the cheapest text-only option at $0.22/M input, with cache hits dropping to $0.007/M. But for multimodal work — image analysis, video understanding, mixed-modality agent loops — GLM 5.3 Flash is the clear choice until DeepSeek's multimodal capabilities catch up.

| Model | Input | Output | Multimodal | Cache | Best For | |---|---|---|---|---|---| | GLM 5.3 Flash | $0.075 | $0.25 | ✅ Yes | $0.015 | Image/video + text, agents, coding | | DeepSeek V4 Flash | $0.22 | $0.66 | ⚠️ Limited | $0.007 | Text/code, heavy caching | | DeepSeek V4 Pro | $0.66 | $1.98 | ⚠️ Limited | $0.022 | Complex reasoning only |

---

The Hermes Connection

GLM 5.3 Flash is already live in the Hermes Agent ecosystem. OpenRouter's public app traffic data shows Hermes Agent as the top consumer of GLM 5.3 Flash tokens — 81 billion and counting. If you're running Hermes with an OpenRouter backend, switching to z-ai/glm-5.3-flash is a one-line config change.

The model's 1M context window and native multimodal input also make it a strong candidate for AiGENCY's own AEO/GEO diagnostic workflows, where structured data, schema markup, and visual content need to be analysed in a single pass.

---

What to Watch

  • September 9, 2026: 50% discount expires; prices double
  • Z.ai Coding Plan: May offer better value for high-volume UK users once the promo ends
  • DeepSeek V4 multimodal roadmap: Watch for updates if you need both text and vision in one model

---

Sources

Back to Insights