A Disruptive Arrival in the AI Landscape
The artificial intelligence sector is currently grappling with the emergence of a mysterious model dubbed “Ox Alpha.” Released anonymously on platforms like OpenRouter and OpenCode on August 20, 2026, the model has captured industry attention by offering a staggering 100 trillion free tokens per day for a week. This aggressive resource allocation, combined with a 1-million-token multi-modal context window, has positioned Ox Alpha as a major point of interest for researchers and developers alike.
Positioned as a reasoning model capable of handling complex software engineering, sustained agentic work, and combined text-visual workflows, Ox Alpha’s technical specifications have fueled a rapid investigation into its origins. The model features a 131,072-token output limit, signaling high-performance capabilities that typically require significant compute infrastructure.
The Search for the Developer
The industry is currently divided on who is behind the project. Initial theories heavily favored the Chinese AI lab Zhipu, citing similarities in video encoding pipelines, tokenizer math, and character-for-character response patterns that mirror the GLM series. Zhipu’s recent infrastructure expansion, including the energizing of a 1GW data center and the management of multiple 10,000-GPU clusters, provides the necessary compute capacity to sustain such a massive free token offer.
However, the hypothesis remains contested. While Zhipu has a history of testing models under “Alpha” monikers—such as the previous “Pony Alpha” release for GLM-5—the use of OpenAI’s cl100k_base tokenizer in Ox Alpha presents a technical contradiction. Chinese-developed models typically avoid OpenAI’s proprietary encoding, leading some observers to suggest that the model could be an unreleased iteration of Microsoft’s upcoming “MAI 2” frontier model.
Alternative theories linking the model to DeepSeek have largely been dismissed. While DeepSeek’s “Hunter Alpha” launch in March 2026 established a precedent for such releases, the company has recently increased prices for its V4-Flash model, suggesting it lacks the current compute resources to sustain the throughput observed with Ox Alpha.

