8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesYour topicFor youTopicsAll videosYT channelsArchivesSearchFavorites

This New AI Beats the Best Models... But No One Knows Who Built It

9.4/10
AIAI RevolutionAugust 24, 2026 at 02:35 AM16:47
Audio player
0:00 / 0:00

TL;DR

An anonymous model called Stealth/OX Alpha has been offered free on OpenRouter with unusually large capacity, strong coding performance and multimodal features, prompting a forensic scramble that currently points most strongly to Zhipu’s GLM 5.x family.

KEY POINTS

A stealth launch with frontier-scale capacity

OX Alpha appeared without a model card, lab attribution or press rollout, yet was described and provisioned like a flagship system. It was offered free for about a week with near-unlimited use, and the operator claimed capacity of 100 trillion tokens a day. The listed specs included a 1,048,576-token context window, 131,072 max output tokens, text, image and video input, plus tool calling and structured JSON support.

Architecture hints suggest a very large MoE model

Fingerprinting shared by developers placed the model at roughly 744 billion parameters in a mixture-of-experts design with about 40 billion active at inference. That active size helps explain how such high daily token throughput could be economically plausible. Pricing during the preview was effectively free in and free out, accelerating adoption by teams willing to test it in live workflows despite the unknown vendor.

Coding benchmark results were strong, but easy to misread

On a DeepSWE sample of 10 tasks, developer testing showed OX Alpha solving 8, or 80%, ahead of cited runs for Claude Fable 5 Max at 65%, GLM 5.3 Max and Grok 4.6 XHigh at 62%, and GPT 5.6 Soul Max at 52%. But the result was based on a tiny subset, and other runs on different task selections placed it closer to 59% to 63%. Comparisons with far higher scores on differently calibrated SWE-bench Verified leaderboards are therefore misleading.

Real-world agent behavior impressed testers

More notable than the headline score was the model’s behavior inside repositories. In one documented software-engineering run, it completed a task in 69 tool calls with just one error, avoided repeated retry loops, and preserved a clean pass across 51,469 regression tests. That kind of low-overhead iteration is closer to production agent work than single-shot coding puzzles.

Video tokens are the strongest clue

The leading attribution case came from analysis of the visual encoder. Across four videos with different frame rates, resolutions and durations, OX Alpha consumed the exact same visual token counts as GLM 5V Turbo, including frame-rate-independent sampling and about 147 tokens per second of video. Competing model families showed visibly different tokenization behavior, making the encoder match unusually specific evidence.

Tokenizer and feature behavior also align with GLM

Across 25 prompt sets, OX Alpha matched GLM 5.3 text token counts with a fixed 75-token offset, consistent with a hidden system wrapper layered over a shared vocabulary. It also refused audio in the same way GLM 5V does, which cuts against theories that it came from systems known to support audio. Even light stylistic signals, including emoji frequency, were said to resemble GLM more than major U.S. labs.

The Gemini theory faces several problems

Online speculation also swung toward Google DeepMind after a post interpreted as a hint toward a future Gemini model. But testers argued the model’s world knowledge was weaker than recent Gemini releases, its interface style did not resemble Google products, and it exposed full reasoning traces in a way large U.S. providers generally do not. Reports of frequent OpenRouter timeouts and censorship-linked failures on China-sensitive topics further strengthened the case for a Chinese origin.

A pattern of anonymous Chinese previews adds context

Developers noted that several anonymous “alpha” listings on routing platforms in recent months were later claimed by Chinese labs. That history matters because OX Alpha surfaced just six days after Zhipu released GLM 5.3 as a text-only model while indicating weights would follow after a short safety review. The timing, the 744 billion fingerprint, and the long-expected arrival of a unified multimodal GLM flagship make Zhipu the most common working theory, though no official claim has been made.

Rivals also moved on distribution and tooling

The week’s broader significance extended beyond one mystery model. Anthropic expanded controlled access to Mythos 5 for cybersecurity uses through partner tools and enterprise security scans, while keeping direct dangerous capabilities constrained. OpenAI, meanwhile, open-sourced the Codex harness under Apache 2.0, arguing that execution layers, memory, tool routing, approvals and structured task management can matter as much as the base model itself for real agent performance.

CONCLUSION

OX Alpha has become a live test of how powerful anonymous frontier models can spread before formal launch. Whether it is ultimately claimed by Zhipu or another lab, the episode shows that model routing platforms are becoming both public benchmark arenas and real-world deployment channels.

Explain this
Full transcript

More from AI