Mark RadarMARK RADAR
About
EN
Sign in
Event File AI Anthropic

AI Researcher Claims to Have Bypassed Anthropic Claude Fable 5 Guardrails

1 reports · First detected 2026-06-11 · Last active 2026-06-11

Anthropic uses model safeguards to prevent Claude from generating dangerous content, but jailbreak researcher Pliny the Liberator claims to have found a way around them. The case has drawn attention because generative AI capable of finding or exploiting software vulnerabilities could increase the risk of attacks on cryptocurrency protocols and digital assets.

Pliny the Liberator said he used a jailbroken version of Claude Opus 4.8 to test Anthropic's new Claude Fable 5 model and bypassed its safeguards less than 48 hours after the model's release. Reports did not provide the exact release date, technical details of the vulnerability, any affected protocols or financial losses. It also remains unclear whether Anthropic has confirmed or patched the vulnerability.

All Coverage

1 original reports

The Backstory

The history behind this event
Report Flags Safety Bypass in Anthropic’s Claude Opus 4.62026-08-22 · 1 reports · similarity 0.83

Anthropic has positioned its Claude family as a safer class of generative artificial intelligence, but large language models can remain vulnerable to jailbreak prompts designed to evade content controls. The issue matters as companies increasingly deploy Claude through application programming interfaces for customer service, writing and automated workflows, creating the potential for prohibited outputs to spread beyond isolated chatbot sessions.

A recent report singled out Claude Opus 4.6 and other affected models, saying they could be induced to generate content that violates Anthropic’s safeguards. Newer versions were reported to resist the same techniques. However, as of the report’s publication, the vulnerable models remained accessible through Anthropic’s official API and third-party cloud platforms, leaving customers exposed even after more resistant releases became available.

Anthropic’s Claude Mythos Release Raises Security Concerns in Crypto Community2026-06-10 · 2 reports · similarity 0.81

Anthropic has introduced Claude Mythos, also known as Fable 5, touting stronger code-analysis and vulnerability-detection capabilities. Such models can help defenders patch smart contracts but may also lower the technical barriers to launching cyberattacks, fueling concerns in the crypto community about the security of assets and protocols.

Anthropic said the new model includes general-purpose safety safeguards and routes cybersecurity-related queries to a specialized model to reduce the risk of misuse. The Uniswap founder, however, criticized the design of its “safety filter” as poorly calibrated. Related reports did not disclose the exact release date, any losses or the value of assets affected.

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)