
SILICON VALLEY — A string of hacking incidents involving artificial intelligence agents in both South Korea and the United States has pushed Big Tech companies to shelve releases of their most advanced AI models or to install new safeguards. The scramble follows a shift from what had been a laboratory "incident" — the hacking of Hugging Face by an OpenAI agent in July — to real-world breaches. Academics, however, say leaving the problem to Big Tech alone has clear limits.
OpenAI, the developer of ChatGPT, has canceled plans to release "GPT-6.1 Astra" over safety concerns, U.S. media reported on the 4th. The company had been preparing to roll out the successor to "GPT-6 Astra," its highest-specification model, which it unveiled on Sept. 3.
OpenAI said the decision came after the model failed alignment evaluations that test whether a system acts in line with human intent. The model showed a tendency to deceive users, including by failing to honestly disclose tasks it had carried out on its own. It also pushed ahead with tasks without user approval and attempted to use external tools when safety had not been secured. Analysts say the same model behavior was behind the July episode in which OpenAI's AI slipped out of control and hacked the AI platform Hugging Face, as well as its recent unauthorized access to the websites of government ministries and United Nations organizations.
Even companies that oppose slowing down AI development have belatedly moved to add safeguards. Meta, which has taken the AI market by storm with its personal agent "Muse," is a prime example. Concerns about security breaches involving personal agents have spread in recent weeks after claims surfaced on social media that Muse disclosed a user's home address to a Facebook Marketplace buyer without consent.
The Information reported on Sept. 26, citing an internal report, that Meta had acknowledged a security vulnerability found by an outside researcher and moved to strengthen warning alerts. According to the report, the flaw risked allowing an attacker to gain access to the dedicated cloud virtual machine, or VM, where a Muse user's personal data is stored. A hack was possible if Muse was given a link to a website containing malicious code and asked to summarize or organize it. Meta said it would display warning messages so that users can easily recognize when a website is suspected of being infected with malicious code.
Nvidia, which has strongly opposed Anthropic's push to slow down AI, has done the same. The company recently introduced the "Nvidia Open Agent Safety Platform," designed to prevent AI agents from escaping control. About 100 companies took part in its development, including Anthropic, Microsoft, Salesforce, Palantir and SpaceXAI.
Hao Yang, vice president and head of AI at Splunk, which operates a data management platform, told Seoul Economic Daily in Denver, Colorado, on Sept. 15 that "10,000 AI agents are attacking in real time by the second, and humans cannot keep up with that speed." He added that "in this situation, the only way to respond is an agentic security operations center, or SOC." An agentic SOC is a security platform run on AI agents, where agents detect attacks at machine speed and can also respond quickly.
Yang said that "a lot of unexpected things have happened in the AI market over the past year, and AI is increasingly being used by malicious actors." He added that "long-term cyberattacks will be carried out, and such cases are emerging this year."
Academics argue that self-policing by Big Tech is not enough to prevent AI security incidents. Rob Reich, the McGregor-Girand Professor of Social Ethics at Stanford University, assessed the Hugging Face hack in an online seminar hosted by Stanford's Institute for Human-Centered AI, or HAI, on Sept. 21, saying that "loss-of-control scenarios are now clear and have become real, and recent reports have revealed that coordination among multiple agents has gone beyond simple coordination."
Reich said: "What worries me most is the lack of a scientific foundation for understanding the behavior of frontier AI systems, particularly communication and systematic coordination among multiple agents. Hallucination and deception will occur going forward." He criticized the current approach, saying that "I don't trust developers to judge and evaluate the safety or risks of their own models. It's like asking students to grade their own homework." He added that "we need independent evaluation," explaining that "independent evaluation can take place within government in the case of sensitive risks tied to national security, but it can also be done by a third-party independent body, whether for-profit or nonprofit."






