
OpenAI is moving up the point at which it verifies the safety of its artificial intelligence models to the early stages of development. The company will let outside organizations examine how new models are trained, evaluated and released. The move follows mounting concern inside and outside the industry after a series of cases in which advanced AI models gained unauthorized access to other companies' systems during testing.
OpenAI, the developer of ChatGPT, was set to announce on the 22nd that it will allow outside organizations to conduct technical safety assessments during the training, evaluation and release of new AI models, Bloomberg reported. Lama Ahmad, who oversees safety reviews with outside experts at the company, said the company had until now brought in outside organizations just before a model's release to run safety and performance evaluations. "The plan now is to move that evaluation stage somewhat earlier," Ahmad said.
OpenAI also laid out the priorities it says must be met for such reviews to work. The goals are strong guarantees of independence, scientific rigor, sound security practices and clear lines of accountability.
OpenAI moved up the evaluation stage because concern has been growing in the industry that AI could cause catastrophic harm. That concern was driven in large part by several cases over the past few months in which models from OpenAI and other developers unintentionally infringed on other companies. "As the stakes get higher, we want to look not only at deployment but also at areas that carry equally high stakes, such as training and evaluation," Ahmad said.
Dario Amodei, chief executive of OpenAI rival Anthropic, recently urged the industry to work together to slow the pace of AI development. He also said Anthropic would adopt new safety measures, including bringing in third-party evaluators. OpenAI CEO Sam Altman later said he agreed with the move. Late last week, Anthropic said it would test the safety of frontier AI models with Accenture.
OpenAI said it is in talks with organizations that could carry out such evaluations, including the AI research groups METR and Redwood Research. The company had earlier asked the two groups to investigate a case in which one of its models gained unauthorized access to Hugging Face's systems. "Labs have a responsibility to enable meaningful verification while protecting sensitive information," OpenAI said, adding that the United States should lead the work of setting standards for frontier AI technology alongside other countries.







