Musk Urges Rival AI Labs to Cross-Test Models for Safety
Rapid gains in generative AI have intensified concerns that frontier models could enable biological weapons, nuclear know-how, cyberattacks or deliberate deception. Safety evaluations are generally designed and run by the developers themselves, creating what critics see as a conflict akin to grading one’s own exam. Elon Musk’s call for cross-company testing seeks to add independent scrutiny and echoes Anthropic Chief Executive Dario Amodei’s warnings that safeguards are failing to keep pace with increasingly capable systems.
Speaking at the All-In Summit in Los Angeles on Sept. 14, 2026, Musk said AI risk was rising “exponentially” and proposed that OpenAI, Anthropic, xAI, Google, Meta and three or four leading Chinese developers run their safety test suites on one another’s models before release. He argued the peer-review approach could win broader agreement than creating a large global regulator and should begin quickly. The tests would probe whether models can help build biological or nuclear weapons, conduct cyberattacks or behave deceptively.
All Coverage
2 original reportsThe Backstory
The history behind this eventMusk Says Anthropic Leads AI, Urges Rival Model Reviews
Rapid advances in artificial intelligence have intensified scrutiny of how powerful models are tested before release. Elon Musk, who leads Tesla and AI developer xAI, argued that rival companies may be well placed to uncover weaknesses because they combine technical expertise with a commercial incentive to challenge competitors. His proposal calls for AI developers to let industry peers test and review new systems before public deployment, adding a market-driven layer of oversight to existing safety evaluations.
In a recent interview, Musk said Anthropic currently leads the AI sector and has a more powerful model that has not yet been released. He did not identify the model, provide benchmark results, specify a launch date or cite any financial figures. Musk also proposed a reciprocal review framework under which competing developers would examine one another’s models before release, contending that industry incentives and technical capabilities could help expose safety risks more effectively.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →