Researchers Crack Hidden AI Reasoning, Expose Credentials
Reasoning models from Anthropic, OpenAI and Google generate hidden chain-of-thought traces before returning answers. To preserve conversational context without storing full sessions, providers send those traces to clients as encrypted blocks that are replayed in later API calls. The approach is meant to protect proprietary methods and sensitive data, but researchers found the blocks could be reused across sessions, users and models within the same provider, creating avenues for model theft, privacy breaches and covert prompt injection.
Researchers affiliated with ELLIS Institute Tübingen and the Max Planck Institute published their paper on Aug. 10, 2026, after testing the systems in early July. They scanned 315,320 encrypted blocks across 6,708 public trajectories and recovered 367 pieces of personally identifiable information and 182 credentials. Genuine user sessions yielded 62 API keys, 33 passwords and 30 email addresses. Anthropic, OpenAI and Google acknowledged the disclosure; the team said the attacks no longer worked afterward, indicating server-side mitigations had been deployed.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.