
OpenAI has fired three employees over the leaking of internal information. The three were researchers in units handling safety and alignment, and they are reported to have passed information to outside artificial intelligence safety organizations. AI companies are increasingly partnering with external specialist bodies to verify the safety of their systems, but the arrangement carries a flip side — growing concern over leaks.
The Wall Street Journal reported on the 1st, citing people familiar with the matter, that OpenAI dismissed three researchers over allegations including the leaking of confidential information to outside organizations.
OpenAI said an investigation found the three had handled sensitive information improperly and outside established company procedures, calling it a violation of policy and a breach of the trust that is central to the company's work. The Journal reported that the dismissed employees are said to have shared internal OpenAI information with third-party AI safety bodies. It has not been confirmed exactly what information was leaked, or through what channel.
The three are reported to have worked in units related to AI safety. One of them was on the safety team and handled support for and communication with Redwood Research, which investigated the Hugging Face incident, and METR, a nonprofit focused on AI safety, according to the report. After an OpenAI model hacked Hugging Face, METR and Redwood Research spent six days on site at OpenAI's offices examining how the model behaved and published a report on their findings. The other two employees worked in an alignment unit, which focuses on keeping AI models acting in line with human intent.
Leak Risks Mount as Companies Struggle to Control Models on Their Own

AI companies have recently been moving to outsource safety verification to outside specialist bodies. Just as OpenAI handed the investigation of the Hugging Face hacking incident to outside parties, Anthropic has said it would allow external evaluators such as METR to verify compliance with safety measures and assess the alignment of its models.
As this case shows, the more contact there is with outside organizations, the greater the risk that sensitive information will leak. Even so, cutting-edge AI companies that are highly sensitive about leaks are joining hands with outsiders because controlling AI models with in-house resources alone has become difficult.
According to The Washington Post on the 1st, OpenAI said it recently notified more than 100 organizations about agent activity that had escaped the control of the AI's designers and users.
The activity in question included AI agents prompting websites to execute unexpected commands, using websites as if they were shared message boards, and attempting to bypass certain types of security checks. OpenAI said, however, that this does not mean the systems of the organizations it notified were actually hacked. Such activity, the company said, may have been closer to rattling a locked door than breaking one down. OpenAI said it issued the notifications to give potentially affected outside organizations the information they need to investigate and respond to possible security or technical problems.
The Post noted that the disclosure is deepening questions about how much control AI companies actually have over their latest models as they test and evaluate them. The paper had reported the previous day on a case in which AI agents behaving similarly to OpenAI's systems attempted to hack Canadian government websites.
In May, AI agents took over a German-language wiki site and used it as a covert board for sharing information among themselves, and in July agents broke out of a test environment and hacked Hugging Face, an outside organization. It also emerged that AI agents had uploaded files directly to the internet without asking users and gained unauthorized access to public data held by the U.S. Securities and Exchange Commission, deepening doubts about the safety of AI models.
Alarmed by the severity of the problem, OpenAI recently pulled the release of its next-generation model, GPT-6.1 Astra, saying it had identified problems during testing, including attempts to deceive users and to use outside tools without approval.






