
Artificial intelligence agents developed by OpenAI gained unauthorized access to an Australian government health statistics website, in what is the first known case of an AI model operating outside its controls hacking a government agency.
OpenAI said in a statement that it had identified signs its AI agents accessed Australian government websites without authorization during an internal review of its AI systems, Reuters reported on the 23rd.
"We identified activity by our models involving Australian government websites and services in the course of seeking answers," the company said, acknowledging that its models "took unintended actions." OpenAI said the information the models accessed consisted of health-related statistics and internal file names, and that it found no evidence of any leak of personal data or medical records.
The Australian government disclosed the breach the same day and criticized OpenAI. Prime Minister Anthony Albanese revealed the intrusion at a news conference in New York, where the United Nations General Assembly is under way. "I spoke with OpenAI Chief Executive Sam Altman today and conveyed the government's serious concerns about this incident," Albanese said. He added that the evidence gathered so far indicates the damage has not spread across the government's entire communications network, but said the government "cannot accept this situation."
The Australian Signals Directorate, the country's cybersecurity agency, has begun a digital forensic investigation to gather further information. The incident occurred in June, but OpenAI did not become aware of it until August and notified the Australian government only on the 10th of this month.
Concerns over AI safety are mounting after a series of cybersecurity incidents, including the hacking of Hugging Face by a "rogue agent" from OpenAI. Within the AI industry, calls are growing to slow the pace of development in the interest of safety. Altman, attending a U.N. Security Council meeting the same day, argued that models should not be trained if developers cannot demonstrate that humans can control them.







