Google Skips 3.5, Unveils Gemini 4 With Security Gains but Coding Doubts

Top-Tier Model Arrives After Seven-Month Gap Matches OpenAI's Astra in Security Testing Internal Voices Point to Benchmark-Focused Tuning

International|
|
By Kim Chang-youngkcy@sedaily.com
||
The Google logo. AFP/Yonhap - Seoul Economic Daily International News from South Korea
The Google logo. AFP/Yonhap

SILICON VALLEY — Google has released its most advanced artificial intelligence model after a seven-month gap, giving it an opening to catch up with OpenAI and Anthropic. Inside the company, however, employees are raising questions about how well the model actually performs.

Google unveiled its next-generation AI model, Gemini 4 Argon, on Sept. 30. The model succeeds Gemini 3, released last November, and is the company's first top-tier release in about seven months, following Gemini 3.1 Pro in February. At its developer conference in May, Google said it would launch Gemini 3.5 Pro in June, but it abandoned that release and moved straight to the next version.

Google said Argon delivers strong performance in knowledge-intensive professional work such as finance, law and taxation, as well as in areas including cybersecurity. To let the model tackle complex problems, Google expanded its output token limit to 1 million tokens, the highest in the industry, from 64,000 tokens.

The new model outperformed OpenAI and Anthropic models on coding and task-performance benchmarks, and matched OpenAI's most advanced model, GPT-6 Astra, in cybersecurity assessments. In an intelligence index compiled by AI model evaluator Artificial Analysis, Argon tied for third place with Anthropic's Claude Fable 5.1 and GPT-6 Astra.

Google stressed that it had trained the model to defend against cyberattacks, in an apparent nod to recent controversy over hacking of AI agents. Argon can autonomously identify and patch critical software vulnerabilities.

Unease persists within Google, however. Bloomberg, citing a person inside the company, reported that Argon's benchmark scores came out well but that employees are running into difficulties when using it for coding work. Experts said Google may be struggling with "benchmaxing," a focus on benchmark scores at the expense of real-world performance.

Original reporting by Kim Chang-young for Seoul Economic Daily.

AI-translated from Korean. Quotes from foreign sources are based on Korean-language reports and may not reflect exact original wording.

Watch · Seoul Economic Daily

More →
1:05

AI KEY

Preview
Korean Corporate Intelligence HubKOSPI · KOSDAQ · 12 sectors

A live, cap-weighted view of every KOSPI and KOSDAQ sector, with same-day Korean reporting distilled by company — built for foreign investors, correspondents and analysts who need to scan Korea before the next session.

Korea Company Atlas

Preview
Market Ontology · The Feedback LoopKFTC 2025 · 92 groups · 121,954 articles

An English ontology of the Korean market — how companies, the media, the government and the National Assembly move each other in a loop. Korea's named controlling persons and designated business groups are a mechanism, not a risk to be priced blind.

SIGNAL

Now live
English Edition · Capital MarketsM&A · IPO · PE · Fund Flows

SIGNAL English Edition is live — Korea's deal desk reporting in English. M&A, IPOs, private equity and fund flows, covered daily for global institutional investors. Browse free; subscriber-only scoops at the 50% intro rate.