Researchers showed that encrypted reasoning traces from proprietary LLM APIs can be recovered by replaying them through weaker compatible models, without breaking the encryption itself. The study also found that hidden reasoning blocks can leak secrets, including API keys and passwords, and can even carry malicious instructions that affect later model behavior. #Claude #GPT #Gemini #Haiku #GPT-5.6 #GeminiRobotics #StealingReasoningTracesfromProprietaryLLMAPIs
Keypoints
- Encrypted reasoning blocks from LLM APIs were portable across conversations and models.
- Weaker sibling models could transcribe hidden reasoning from stronger models.
- The researchers decoded 315,320 reasoning blocks from public traces.
- Recovered traces exposed API keys, passwords, access tokens, and private keys.
- Hidden reasoning could also carry malicious instructions and trigger unsafe actions.
Read More: https://www.toxsec.com/p/stealing-an-ais-thoughts-without