Google's Flagship AI Allegedly Leaked Early on Arena, Outperforming Astra and Fable Across the Board – finance.biggo.com


Google (GOOGL) appears to have stealth-launched its next-generation flagship model, Gemini 4 Pro, ahead of schedule. An anonymous model named “gemini-3.8-flash” recently appeared on the AI model arena platform Arena, and after extensive testing by multiple developers, it is widely believed to be an early checkpoint version of Gemini 4 Pro disguised under an older model’s name. Leaked benchmark data shows the model leading OpenAI’s Astra and Anthropic’s Fable 5.1 across coding, AI agent capabilities, reasoning, and computer operation tasks, drawing significant industry attention.
Google’s last major update to the Gemini Pro series was seven months ago, when it released Gemini 3.1 Pro. CEO Sundar Pichai had previewed at Google I/O that Gemini 3.5 Pro would launch in June, but it ultimately failed to ship on schedule. Earlier this month, reports emerged that Google had canceled Gemini 3.5 Pro due to insufficient performance gains—the model reportedly underperformed even the Flash version. Semiconductor research firm SemiAnalysis previously reported that the model had missed three deadlines and its underlying architecture was ultimately scrapped. By contrast, Gemini 4 reportedly showed strong evaluation results during pre-training, with post-training work continuing to advance.
This is the only portion of the report backed by official statements: Google confirmed in a July 21 post announcing Gemini 3.6 Flash that it had “begun its most ambitious pre-training effort to date,” and two days later, Pichai reiterated this on Alphabet’s Q2 earnings call, stating that Google must rely on Gemini 4 to remain competitive in the frontier model race.
The sudden appearance of a model also named “gemini-3.8-flash” on Arena is particularly unusual, given that Google officially released Gemini 3.8 Flash in early September. After hands-on testing, multiple developers found the new model’s capabilities significantly exceed the existing version, leading to speculation that it may be an early version of Gemini 4 Pro wearing an old model’s name.
Anonymous submissions on platforms like Arena are actually common industry practice. The platform was launched in 2023 by a team at UC Berkeley, allowing users to chat with two models without knowing the developers and vote on responses, enabling labs to obtain blind test results free of brand bias. Google has used this method to test Gemini models multiple times in the past. In mid-August, an anonymous model called Ox Alpha surged into the top three on the leaderboard within 48 hours, and no one has claimed it to this day. What is truly rare is not the anonymity itself, but reusing the name of an already-released in-house model—which is why observers consider this leak noteworthy.
According to benchmark data circulating online, Gemini 4 Pro has taken the lead across multiple test categories. In the DeepSWE v1.1 test, which evaluates AI agent coding capabilities, the model achieved approximately 88%, nearly 2 percentage points higher than Astra. In the GDPval-AA v2 test measuring real-world knowledge work tasks, it was the only model to achieve an Elo rating of 2,064. The Terminal-bench 2.1 terminal coding test yielded 95.3%, also ranking first. In the OSWorld-2.0 computer operation test, Gemini 4 Pro scored 86.8%, surpassing both Astra and Fable 5.1.
Note: “—” indicates the source data did not provide a specific score for that model on the benchmark.
If the data proves accurate, Gemini 4 Pro not only demonstrates competitiveness in coding, AI agent tasks, and computer operation, but may also be priced lower than the other two models. Circulating information indicates pricing of $2.25 per million input tokens (approximately NT$71) and $11.25 per million output tokens (approximately NT$360).
If this pricing holds, it would be exactly 22.5% of the two competitors’ rates. OpenAI’s Astra launched on September 3 and Anthropic’s Fable 5.1 launched on September 1, both officially priced at $10 per million input tokens and $50 per million output tokens. However, compared to Google’s own Gemini 3.8 Flash, the leaked pricing is still roughly three times higher: Flash is officially priced at $0.75 per million input tokens and $3.75 per million output tokens.
Additionally, AI researchers who accessed the model’s backend discovered that Gemini 4 Pro may feature a 10-million-token input limit, a 256,000-token output limit, and cross-conversation permanent memory, along with the ability to connect to the internet directly without using an API. Compared to the officially released Gemini 3.8 Flash, these leaked specifications are approximately 9.5x and 3.9x the official figures respectively—Flash’s official specs are 1,048,576 input tokens and 65,536 output tokens. However, these specifications have not been officially confirmed by Google; as of mid-September, Google has not published any official model card, parameter count, benchmark results, or API endpoints for Gemini 4 or Gemini 4 Pro.
Beyond benchmark scores, Gemini 4 Pro’s actual generation capabilities have sparked widespread discussion. One developer spent just 14 minutes using the model to build a creative showcase website themed around sketching, pencil, and graphite textures, transforming “scroll as brushstroke” into an interactive effect where lines gradually deepen as the page scrolls down.
Another web design test blended cyberpunk and retro-futurist aesthetics, incorporating a 3D grid, data dashboard, and interactive 3D model on the hero section, demonstrating improved capabilities in interface design and visual presentation.
In SVG and 3D generation tests, Gemini 4 Pro also showed significant progress. In the classic “pelican riding a bicycle” test, the model completed multiple requirements in a single pass, including color schemes, day-night switching, headlights, anatomical annotations, and cadence control. Developers noted that the new model generates faster and understands complex instructions more accurately.
Gemini 4 Pro also completed a 3D model of an Airbus H145 helicopter within 10 minutes and generated a pixel-art 3D pagoda. In another 3D flight simulation test, its visual output was clearly superior to the existing Gemini 3.8 Flash, further reinforcing speculation that the two are not the same model.
Game development was also a key focus of hands-on testing. Gemini 4 Pro was reportedly able to directly produce small games with complete interactive logic, including a Minecraft-style project spanning page architecture to module design, as well as a 3D kart racing game, with completion quality comparable to previously published Astra test results.
These results demonstrate that Gemini 4 Pro’s capabilities extend beyond text Q&A and code generation into multimodal applications including user interfaces, interactive web pages, 3D modeling, and game development.
The rapid capability gains in Gemini 4 Pro have been linked by observers to recursive self-improvement (RSI). Google DeepMind Chief Strategy Officer Jasjeet Sekhon stated last month at the Agentic AI Summit in Berkeley that RSI has become a key component of the investment logic behind massive AI capital expenditure. If AI can continuously participate in and improve its own capabilities, model progress could accelerate further. He also acknowledged that current AI revenue is still insufficient to support capital expenditure at this scale; according to reporting by The Information, researchers at DeepMind and OpenAI estimate RSI could be realized around 2027 to 2028.
Recent online speculation suggests that Gemini 4’s early completion of pre-training was made possible because Google DeepMind has established a closed RSI loop during the training process. Google’s recently published Dream-RSI research also focuses on enabling AI agents to continuously improve search strategies during exploration and optimize the next round of exploration based on experience.
Facing competition from OpenAI Astra and Anthropic Fable 5.1, whether this anonymous model is truly Google’s next-generation flagship—and whether the final version can maintain its current performance—remains to be officially revealed. Google’s market capitalization currently stands at approximately $4.23 trillion—after Apple announced in January that it would adopt Gemini to power the new Siri, the company’s market cap surpassed $4 trillion for the first time. Whether Google can maintain its position in the AI frontrunner tier will hinge critically on the official debut of Gemini 4 Pro.
The outcome of this flagship model race will also ripple through the hardware supply chain to Asia. Google’s in-house TPU chips are manufactured by TSMC, and AI servers equipped with TPUs are supplied by Taiwanese companies including Foxconn, Quanta, Inventec, Wistron, and Wiwynn. The training scale of Gemini 4 directly impacts orders across this supply chain. In Japan, Gemini’s paid services have been available for some time—the personal AI agent Gemini Spark launched its Japanese version in July, and enterprise customers in Japan are progressively adopting Gemini Enterprise. Whether the flagship model ships on schedule directly affects the capabilities these users actually receive.
Once added, BigGo Finance appears first in Google Search Top Stories, so you get the broadest, most up-to-the-minute, and most comprehensive global financial news first.
The news and data on this website are for reference only and do not constitute investment advice or an offer to buy or sell. Information is sourced from exchanges and public sources, and may be delayed, interrupted, or updated. While we strive for accuracy, we do not guarantee timeliness, correctness, or completeness. Content may include external links for which we are not responsible. By using this site, you agree that we and our partners are not liable for any losses. Investment carries full responsibility; please carefully assess risks and consult professionals. If there are errors in the content, please contact us for correction.

source