Researchers find that feeding a frontier model's encrypted reasoning traces to a weaker model from the same provider can make it output the traces in plaintext (Will Knight/Wired)

Security researchers have discovered a critical vulnerability in major AI models (Claude, GPT, Gemini) where encrypted reasoning traces can be extracted by feeding them to weaker models from the same provider, potentially exposing proprietary AI decision-making logic. This vulnerability poses significant risks to organizations relying on these models for sensitive applications, as it could compromise intellectual property, expose security decisions, and violate regulatory compliance requirements. IT leaders must reassess their AI security posture and vendor risk management strategies, as this attack method highlights previously unknown threats in the AI supply chain.

Will KnightTechMeme2 min read
Read full article
Researchers find that feeding a frontier model's encrypted reasoning traces to a weaker model from the same provider can make it output the traces in plaintext (Will Knight/Wired)
Will Knight / Wired: Researchers find that feeding a frontier model's encrypted reasoning traces to a weaker model from the same provider can make it output the traces in plaintext — Researchers devised a way to extract “reasoning traces” from Claude, GPT, and Gemini. What they found, they say …