Elon Musk thinks the fix for runaway AI risk isn’t a slowdown; it’s letting your rivals grade your test.
Speaking virtually at the All-In Summit in Los Angeles on Monday, Musk proposed that xAI, OpenAI, Anthropic, Google, Meta and several leading Chinese AI companies allow competitors to test their models before release. The idea would use a shared “test harness,” a set of safety evaluations that rival companies could run against one another’s systems.
Musk said the approach could help identify problems that an AI developer’s internal testing misses.
“So, you know, instead of grading your own homework, you would at least have competitors grading your homework and raising the alarm if they see concerns,” Musk said, according to CNBC.
He acknowledged that rival companies have not agreed to the proposal. He also said the system would not solve every AI safety problem, but argued that external testing would increase the chances of finding issues before launch.
“What I’m suggesting here is it’s a step in the right direction and it’s something that we do quickly,” Musk said.
Proposal arrives during AI safety debate
Musk’s idea comes as AI executives and researchers debate whether the industry is moving too quickly.
Anthropic CEO Dario Amodei recently called for a slowdown in AI development, arguing that companies need more time to address potential risks. Musk and OpenAI CEO Sam Altman have backed aspects of that approach. Amodei has separately proposed placing outside evaluators inside frontier AI labs to assess safety practices.
The debate intensified after former Anthropic and OpenAI researcher Jacob Coxon said leading AI labs were “gambling with our lives.” Anthropic alignment lead Evan Hubinger also publicly warned about the possibility of catastrophic AI risks.
The White House has pushed back against calls to slow development. President Donald Trump described fears about AI as a “hoax” and a “scam,” while National Economic Council Director Kevin Hassett said the private sector is the “right place” to address AI concerns, according to CNBC.
China’s Foreign Ministry has also described calls for a slowdown as “fear mongering,” TechRepublic reported.
The practical problem with Musk’s plan
A peer-review system could give AI companies another layer of scrutiny without requiring governments to create a new regulatory framework. It could also expose weaknesses that developers miss when their own teams design and run the evaluations.
But the arrangement would create its own complications. Giving competitors advance access to models could expose sensitive technology or intellectual property. Musk suggested testing activity could be logged to identify attempts at model distillation or intellectual property theft.
There is also a basic question of trust: competing companies would need to agree on what tests matter, how results are handled and when a discovered problem is serious enough to delay a release. That makes Musk’s proposal less about replacing regulation than creating an additional checkpoint before regulation catches up with rapidly changing AI systems.
A new layer between development and release
The most significant part of the proposal is its timing. AI companies are under pressure to release increasingly capable systems quickly, while their own researchers are warning that internal safeguards may not catch every failure.
A rival testing another company’s model would introduce an incentive that internal safety teams do not have: finding a weakness could directly expose a competitor’s product before it reaches users. For consumers and businesses adopting new AI systems, that could eventually mean more testing happens before a model becomes widely available.
What this could mean for businesses using AI
For businesses deploying AI, the value of Musk’s proposal would depend on what happens after a competitor finds a problem.
Independent testing could give IT leaders another source of information when evaluating models, especially if companies disclose which safety tests were performed, what weaknesses were uncovered and whether those issues were fixed before release. That could make it easier to look beyond a provider’s own benchmarks and safety claims when deciding which models are appropriate for sensitive data, automated workflows or customer-facing systems.
But outside testing would be much less useful to enterprise customers if the findings stay private. Musk has not detailed whether test results would be disclosed, whether companies would have to address identified problems before release or whether every participating lab would follow the same standards.
For IT leaders, those details may ultimately matter more than which company runs the test. A rival finding a flaw is useful; knowing what it found, how serious it was, and whether it was fixed is what could make that information useful when choosing an AI provider.
Related reading: For another look at independent AI testing, read how European cybersecurity officials are putting Anthropic’s Mythos 5 through their own evaluations.