Skip to content
Daily Edition · AI industry record
Live desk ●
LaunchNews Report1 min readUpdatedByAI Tools Daily

Grok 4 Launches: xAI's 200,000-GPU Bid for the Benchmark Crown, With a $300/mo Heavy Tier

xAI released Grok 4 in July 2025, setting records on GPQA and other hard benchmarks, plus a multi-agent Heavy edition.

xAI released Grok 4 on July 9, 2025, trained on the ~200,000-GPU Colossus cluster. It set then-records on hard benchmarks like GPQA Diamond (~88%) and topped the Artificial Analysis Intelligence Index.

The companion Grok 4 Heavy uses a multi-agent parallel reasoning architecture, offered via the $300/month SuperGrok Heavy subscription — pioneering an ultra-premium consumer AI tier.

Critics noted xAI published no training-data or safety-evaluation details, lagging peers on transparency; and Gemini 3 and Opus 4.5 soon reshuffled the leaderboards — the frontier lead window has shrunk to months.

Brute Force: Stress-Testing the Compute Entry Ticket

Grok 4 is the extreme case of the 'compute is destiny' playbook: xAI built the 200,000-GPU Colossus cluster in under two years, bypassing infrastructure moats other labs spent years assembling — trading capex for time. The path proved one thing: with enough compute and talent, a latecomer can crash the front rank within two years. But it also drew a new threshold — the entry ticket to frontier competition is now a 100,000-GPU-class cluster, a microcosm of the global compute arms race (see our Stargate coverage).

The $300/month Heavy tier was a second experiment: multi-agent parallel reasoning multiplies per-query costs, so vendors must test how much premium power users will pay for maximum intelligence. OpenAI and Anthropic followed with their own high-end subscription tiers — the consumer AI price ceiling was officially lifted.

Our Take

Grok 4's rise and eclipse together sketch the real shape of 2025's frontier race: the top of the leaderboard went from moat to billboard with a lease measured in months (see our Opus 4.5 coverage). xAI's true differentiation is not scores but resource-assembly speed — a self-built mega-cluster plus X's real-time data. And the transparency gap is a reminder: beyond benchmarks, safety disclosure and reproducibility are becoming the second yardstick in enterprise procurement — a leaderboard rank does not buy a contract.

This article aggregates official announcements and public reporting; original sources are linked below.

Tools in this story