OpenAI Expands Outside Safety Reviews Into Model Training

OpenAI Expands Outside Safety Reviews Into Model Training

OpenAI plans independent safety checks before AI models launch. Image: ChatGPT

OpenAI plans to expand independent safety assessments into model training and evaluation, giving outside groups earlier access to test high-risk AI systems.

Sep 23, 2026

OpenAI wants outside safety evaluators involved before its models are finished, not just shortly before launch.

The company said Tuesday that it plans to support independent technical safety assessments during training, evaluation, and deployment, expanding beyond the pre-launch reviews it has relied on more heavily in the past.

Lama Ahmad, who leads OpenAI’s work with outside safety experts, told Bloomberg that the company previously focused on bringing third parties in shortly before launch. “As the stakes get higher, we want to make sure we’re also looking at things like training and evaluation, which do have high stakes, in addition to our deployments,” Ahmad said.

OpenAI said the assessments should have strong independence mechanisms, scientific rigor, robust security practices, and clear responsibilities.

The company said it is already talking with multiple potential assessors, including AI research groups METR and Redwood Research. Both groups previously investigated an incident involving OpenAI models gaining unauthorized access to Hugging Face systems.

What outside groups will examine

OpenAI outlined four areas where it wants deeper independent scrutiny.

Assessors could examine whether the evidence behind the company’s safety cases supports its claims, including safeguards used during training, evaluation and deployment. They could also test critical safeguards against jailbreaks and other adversarial attacks and examine defenses against high-risk capabilities in areas such as cybersecurity and biological misuse.

Other work could focus on whether capability evaluations adequately measure risks covered by OpenAI’s Preparedness Framework, as well as whether alignment evaluations can detect serious misalignment.

The company also wants independent investigations into selected incidents involving models acting without authorization or attempting to evade oversight. OpenAI said some assessments could last weeks while others could run for several months, meaning the work is intended to examine safety claims over time rather than serve solely as a final pre-launch check.

Advertisement

The hard part is independence

OpenAI’s proposal comes as AI companies face growing pressure to prove that their safety claims can withstand scrutiny from organizations outside the companies building the models.

The company says assessors should agree on clearly defined claims before testing begins, receive access proportionate to what they are evaluating, and disclose conflicts of interest. It also says sensitive information may require assessments to happen on company-managed devices or at company facilities.

That creates a practical tension. Meaningful safety testing may require access to sensitive model data, internal systems, or security controls, but tight restrictions can also limit how independently outside researchers can validate a company’s claims.

OpenAI says it wants findings to be as transparent as possible while still protecting confidential information, proprietary technology, and security-sensitive details.

What it means for AI development

Conducting assessments during training allows teams to identify vulnerabilities when adjustments to model architecture or safety protocols remain viable, avoiding late-stage discoveries right around release.

However, earlier evaluation windows do not automatically guarantee robust oversight. The true value of this independent scrutiny rests on evaluator qualifications, the depth of system access granted, and the precise scope of claims cleared for testing.

Key operational details also remain unconfirmed. OpenAI has yet to name official assessment partners or outline public access guidelines. In contrast, Anthropic has taken a more concrete step, announcing last week that it would embed evaluators from Accenture to test its frontier models.

Advertisement

What it means for users

While end users of ChatGPT are unlikely to notice immediate changes, the primary value of this initiative unfolds behind the scenes. Independent evaluations could allow OpenAI to detect alignment, security, and misuse risks earlier in the development lifecycle, preventing issues before systems are deployed.

Ultimately, users stand to gain stronger safety protections if these third-party assessments yield actionable insights that OpenAI implements. However, full transparency may be limited, as certain disclosures could be withheld to safeguard proprietary systems or sensitive vulnerabilities.

OpenAI says it plans to expand the independent assessment ecosystem and work toward shared international standards for technical safety reviews as frontier models become more capable. The larger test will be whether outside evaluation becomes a meaningful check on frontier-model development or simply another layer of company-managed assurance.

Let us teach you How to Talk to AI for free! Try our six-minute course at The Neuron Academy and learn a few simple ways to write better prompts and get more useful results from AI, or browse our other AI course for free for seven days. Check out all the lessons here →

Aminu Abdullahi

Aminu Abdullahi is a B2C and B2B technology and finance writer with more than six years of experience covering enterprise IT, cybersecurity, cloud computing, artificial intelligence, fintech, business software, and emerging technologies. He has written for a wide range of technical and business audiences, from IT professionals and cybersecurity leaders to small business owners, executives, and technology buyers. His work has appeared in publications including: TechRepublic eWEEK Channel Insider Geekflare Enterprise Networking Planet eSecurity Planet CIO Insight Webopedia With a background in computer science, Aminu specializes in translating complex technical subjects into clear, practical, and accessible content. His writing helps readers understand emerging technologies, evaluate business software, strengthen cybersecurity strategies, and make more informed decisions about technology investments. Across his work, Aminu focuses on the real-world impact of technology, connecting technical innovation with business value, operational efficiency, security, and long-term digital transformation.