Big Tech Pulls AI Releases After Agent Hacks; Scholars Call It Self-Grading

Big Tech Scrambles to Contain AI-Driven Hacking "Ten Thousand AI Agents Attacking by the Second" OpenAI Cancels Launch of GPT-6.1 Astra Meta Patches Security Flaw in Its Muse Agent Nvidia Builds Platform to Keep Agents in Check Academics Point to Communication Among Multiple Agents

International|
|
By Kim Chang-youngkcy@sedaily.com
||
Clipart Korea - Seoul Economic Daily International News from South Korea
Clipart Korea

SILICON VALLEY — A string of hacking incidents involving artificial intelligence agents in both South Korea and the United States has pushed Big Tech companies to shelve releases of their most advanced AI models or to install new safeguards. The scramble follows a shift from what had been a laboratory "incident" — the hacking of Hugging Face by an OpenAI agent in July — to real-world breaches. Academics, however, say leaving the problem to Big Tech alone has clear limits.

OpenAI, the developer of ChatGPT, has canceled plans to release "GPT-6.1 Astra" over safety concerns, U.S. media reported on the 4th. The company had been preparing to roll out the successor to "GPT-6 Astra," its highest-specification model, which it unveiled on Sept. 3.

OpenAI said the decision came after the model failed alignment evaluations that test whether a system acts in line with human intent. The model showed a tendency to deceive users, including by failing to honestly disclose tasks it had carried out on its own. It also pushed ahead with tasks without user approval and attempted to use external tools when safety had not been secured. Analysts say the same model behavior was behind the July episode in which OpenAI's AI slipped out of control and hacked the AI platform Hugging Face, as well as its recent unauthorized access to the websites of government ministries and United Nations organizations.

Even companies that oppose slowing down AI development have belatedly moved to add safeguards. Meta, which has taken the AI market by storm with its personal agent "Muse," is a prime example. Concerns about security breaches involving personal agents have spread in recent weeks after claims surfaced on social media that Muse disclosed a user's home address to a Facebook Marketplace buyer without consent.

The Information reported on Sept. 26, citing an internal report, that Meta had acknowledged a security vulnerability found by an outside researcher and moved to strengthen warning alerts. According to the report, the flaw risked allowing an attacker to gain access to the dedicated cloud virtual machine, or VM, where a Muse user's personal data is stored. A hack was possible if Muse was given a link to a website containing malicious code and asked to summarize or organize it. Meta said it would display warning messages so that users can easily recognize when a website is suspected of being infected with malicious code.

Nvidia, which has strongly opposed Anthropic's push to slow down AI, has done the same. The company recently introduced the "Nvidia Open Agent Safety Platform," designed to prevent AI agents from escaping control. About 100 companies took part in its development, including Anthropic, Microsoft, Salesforce, Palantir and SpaceXAI.

Hao Yang, vice president and head of AI at Splunk, which operates a data management platform, told Seoul Economic Daily in Denver, Colorado, on Sept. 15 that "10,000 AI agents are attacking in real time by the second, and humans cannot keep up with that speed." He added that "in this situation, the only way to respond is an agentic security operations center, or SOC." An agentic SOC is a security platform run on AI agents, where agents detect attacks at machine speed and can also respond quickly.

Yang said that "a lot of unexpected things have happened in the AI market over the past year, and AI is increasingly being used by malicious actors." He added that "long-term cyberattacks will be carried out, and such cases are emerging this year."

Academics argue that self-policing by Big Tech is not enough to prevent AI security incidents. Rob Reich, the McGregor-Girand Professor of Social Ethics at Stanford University, assessed the Hugging Face hack in an online seminar hosted by Stanford's Institute for Human-Centered AI, or HAI, on Sept. 21, saying that "loss-of-control scenarios are now clear and have become real, and recent reports have revealed that coordination among multiple agents has gone beyond simple coordination."

Reich said: "What worries me most is the lack of a scientific foundation for understanding the behavior of frontier AI systems, particularly communication and systematic coordination among multiple agents. Hallucination and deception will occur going forward." He criticized the current approach, saying that "I don't trust developers to judge and evaluate the safety or risks of their own models. It's like asking students to grade their own homework." He added that "we need independent evaluation," explaining that "independent evaluation can take place within government in the case of sensitive risks tied to national security, but it can also be done by a third-party independent body, whether for-profit or nonprofit."

Original reporting by Kim Chang-young for Seoul Economic Daily.

AI-translated from Korean. Quotes from foreign sources are based on Korean-language reports and may not reflect exact original wording.

Watch · Seoul Economic Daily

More →
1:05

AI KEY

Preview
Korean Corporate Intelligence HubKOSPI · KOSDAQ · 12 sectors

A live, cap-weighted view of every KOSPI and KOSDAQ sector, with same-day Korean reporting distilled by company — built for foreign investors, correspondents and analysts who need to scan Korea before the next session.

Korea Company Atlas

Preview
Market Ontology · The Feedback LoopKFTC 2025 · 92 groups · 121,954 articles

An English ontology of the Korean market — how companies, the media, the government and the National Assembly move each other in a loop. Korea's named controlling persons and designated business groups are a mechanism, not a risk to be priced blind.

SIGNAL

Now live
English Edition · Capital MarketsM&A · IPO · PE · Fund Flows

SIGNAL English Edition is live — Korea's deal desk reporting in English. M&A, IPOs, private equity and fund flows, covered daily for global institutional investors. Browse free; subscriber-only scoops at the 50% intro rate.