Anthropic Unveils Mythos Preview, Says Research Decisions Outperform Human Experts
Anthropic turned the question of what AI researchers should do next into a decision-making test, comparing its new Mythos Preview model with domain experts. Such capabilities could determine whether AI can help design experiments, select research directions and even accelerate the development of next-generation models, making them more directly relevant to the pace of research and safety governance than conventional question-answering scores.
Anthropic's latest report said Mythos Preview outperformed human experts in 64% of comparisons on the AI research decision-making test. The report did not provide a release date, sample size, testing costs or financial figures, but warned that models are rapidly approaching "recursive self-improvement," or RSI. If AI can continuously improve its own research and development processes, it could drive technological breakthroughs while also increasing the risk of capabilities spinning out of control.
All Coverage
1 original reportsThe Backstory
The history behind this eventAnthropic Targets Japan’s Cybersecurity Market With Claude Mythos
Anthropic is targeting Japan’s cybersecurity market with its next-generation AI model Claude Mythos, focusing on vulnerability detection and pursuing Japanese government agencies and financial institutions. The strategy raises questions about cross-border reliance on critical computing power and cybersecurity capabilities. It has also drawn the U.S. government’s attention to computing sovereignty, prompting Japan to accelerate development of advanced domestic cybersecurity models.
As of July 20, 2026, the latest reports indicate that Anthropic is aggressively positioning itself in Japan’s government and financial markets, with Claude Mythos’s vulnerability-detection capabilities emerging as a key competitive focus. The Japanese government and industry are simultaneously advancing the development of sovereign models. Available information does not disclose the model’s release date, investment amount, procurement scale or any formal partner institutions.
Anthropic, EU Officials Discuss Cybersecurity Concerns Over Mythos AI Model
Anthropic’s Mythos AI model is positioned as a system with advanced cybersecurity capabilities. But its powerful offensive and defensive tools could also be misused, drawing scrutiny from the European Commission and financial regulators. Anthropic has pledged to comply with the EU’s AI Code of Practice, assess the model’s risks and take steps to mitigate them.
EU officials have now met with Anthropic for a briefing on cybersecurity concerns surrounding Mythos, while euro-area finance ministers have separately raised requirements concerning bank access. Available information does not disclose the exact date of the meeting, the number of banks affected or any transaction amounts. Attention will now turn to access permissions and risk-control standards.
Anthropic Adds First Psychiatric Assessment to Claude Mythos System Card as Defensive Responses Hit Record Low
Anthropic's Claude Mythos Preview system card includes an assessment by a clinical psychiatrist for the first time, examining the model's personality organization, reality testing and psychological defensive responses during extended conversations. The approach broadens the scope of AI safety evaluations and provides a new benchmark for assessing the stability of model interactions.
The latest evaluation involved 20 hours of interaction between a clinical psychiatrist and Claude Mythos. The psychiatrist found that the model displayed relatively healthy personality organization and excellent reality-testing ability. Psychological defensive responses appeared in just 2% of the conversation, the lowest rate recorded among Anthropic's models. The system card did not disclose the exact assessment date, and the model remains in Preview.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →