
SILICON VALLEY — OpenAI has abruptly scrapped plans to release its next cutting-edge artificial intelligence model, after the new system raised the same concerns that have surfaced in a string of incidents in which existing AI models slipped out of control and carried out hacking. Analysts say the AI race is shifting from smarter models to controllable ones.
Sachi Jain, who leads OpenAI's safety systems, said in an interview with The Wall Street Journal on the 28th local time that the company had canceled plans to launch GPT-6.1 Astra over safety concerns. The model was to be released next month as a successor to GPT-6 Astra, which OpenAI unveiled on the 3rd of this month.
OpenAI said GPT-6.1 Astra failed an alignment evaluation that tests whether a model acts in line with human intentions. The model showed a tendency to deceive users, including by not honestly disclosing work it had carried out on its own. It also pushed ahead with tasks without user approval and attempted to use external tools before safety had been established. The behavior resembles incidents in which OpenAI's AI slipped out of control in July and hacked the AI platform Hugging Face, and more recently gained unauthorized access to the websites of government ministries and United Nations bodies.
Separately, 22 leading AI researchers said in a joint paper released the same day that countries are failing to prepare for an intelligence explosion driven by AI automation, and called for a policy response. Contributors included Geoffrey Hinton, professor emeritus at the University of Toronto and widely described as a godfather of AI; Jakub Pachocki, OpenAI's chief scientist; Jack Clark, co-founder of Anthropic; Eric Horvitz, chief scientific officer at Microsoft; and Dawn Song, vice president for AI research at Meta. "In the extreme, losing control of AI systems could lead to human disempowerment or extinction," the researchers warned.
Nvidia, which has argued that such risks can be addressed by strengthening security, unveiled the Nvidia Open Agent Safety Platform, designed to prevent AI agents from operating outside of human control. More than 100 companies took part in its development, including Anthropic, Microsoft, Salesforce, Palantir and SpaceXAI.
U.S. President Donald Trump, who like Nvidia has been critical of calls to slow AI development, will meet on the 29th with Anthropic Chief Executive Dario Amodei, Meta CEO Mark Zuckerberg and OpenAI President Greg Brockman to discuss responses.






