
Google officially introduced its latest flagship AI model, Gemini 4 Argon, showing off test scores that beat OpenAI’s GPT-6 Astra across several key categories. But behind the launch numbers, internal reports indicate that Googlers are split on how well the model handles practical, everyday work.
According to Bloomberg, several employees with access to the model found that while Gemini 4 Argon crushes standardized benchmarks, it struggles when put to work on real-world coding tasks and front-end website design. Critics point to “benchmaxxing”—an industry trend where models are trained specifically to pass test suites rather than solve messy software problems. Edwin Chen, founder of Surge AI, compared it to a student getting top SAT scores without developing actual real-world skills.
A high-stakes race after shifting strategies and talent departures
The pressure on Gemini 4 Argon is huge. Google previously scrapped its planned Gemini 3.5 Pro model after promising a June release. Analysts at Bloomberg Intelligence estimate this decision may have cost up to $400 million in wasted training runs alone. The company has also seen high-profile talent departures, including legendary engineer Jeff Dean, Nobel winner John Jumper, and Noam Shazeer. Meanwhile, Demis Hassabis moved into a chairman role while longtime lieutenant Koray Kavukcuoglu took over day-to-day operations at Google DeepMind.
Despite those shifts, Google strongly denies that Gemini 4 Argon underperforms in coding. Kavukcuoglu expressed full confidence in the team, stating it is a certainty Google will stay at the AI frontier. Internal supporters note there is a large consensus that the model leads in safety, cybersecurity defense, natural conversation, and video metadata extraction. Google also emphasized that its consumer chatbot and AI Mode in Search have both passed 1 billion users. This shows strong momentum for its consumer AI products.
Tough competition, higher specs and lower introductory prices
Gemini 4 Argon has some pretty impressive specs. To start, there’s a massive output limit of up to 1 million tokens, on paper. Google is pricing the model at $4 per million input tokens and $20 per million output tokens. However, the firm is launching it with a temporary 50% discount to entice developers.
Google needs this flagship model to land cleanly. Gemini powers everything from Search and Maps to Gmail and Chrome, but competitors aren’t slowing down. OpenAI and Anthropic are advancing models like Fable and Astra, while Meta recently launched its Muse AI task agent. Offering high-context windows and half-price entry rates is Google’s bid to keep developers locked into its platform as the AI race speeds up.
The post Googlers Are Privately Questioning If Gemini 4 Argon Can Actually Deliver appeared first on Android Headlines.