A significant architectural flaw in how major AI providers, including OpenAI, Anthropic, and Google, protect the internal “chain-of-thought” reasoning generated by their flagship large language models (LLMs). The research reveals that encrypted reasoning envelopes returned by provider APIs can be replayed into weaker, less-guarded sibling models to extract private reasoning traces in plain text. Detailed by a collaborative research team from the ELLIS Institute Tübingen, the Max Planck Institute, MATS Research, and Snyk, the attack affects the Claude, GPT, and Gemini model ecosystems and requires only standard, unprivileged API access. APIs Flaw Exposes Hidden Reasoning Traces Modern reasoning architectures such as...
Read the full article at the source.
Comments (0)
No comments yet. Be the first to comment!