Researchers steal reasoning traces from Anthropic, OpenAI and Google APIs by reusing encrypted blocks

The paper "Stealing Reasoning Traces from Proprietary LLM APIs" (arXiv v1 2026-08-10; Willison's write-up 2026-08-11) attacks the hidden-CoT design of B08-01 and B08-06.

Date
10 August 2026
Who
academic researchers (eight authors including Alexander Panfilov and Ilia Shumailov); targets Anthropic, OpenAI and Google
Confidence
High on the paper's abstract; Medium on the vendors' patches (Willison's account
Deep dive
Reasoning II, from o1 and o3 to DeepSeek-R1 and the labs that replicated them

Tier: Supporting · Significance: 3/5 · Org(s): academic researchers (eight authors including Alexander Panfilov and Ilia Shumailov); targets Anthropic, OpenAI and Google · Confidence: High on the paper's abstract; Medium on the vendors' patches (Willison's account) The paper "Stealing Reasoning Traces from Proprietary LLM APIs" (arXiv v1 2026-08-10; Willison's write-up 2026-08-11) attacks the hidden-CoT design of B08-01 and B08-06. Providers keep reasoning private by returning it to the client as encrypted blocks that are passed back on later turns, and the authors found these blocks interchangeable across sessions, users and models within a provider.

Injecting a stronger model's encrypted trace into a weaker, less safeguarded model from the same provider made it decode and print the trace verbatim, without jailbreaking the stronger model; they demonstrate this across Anthropic, OpenAI and Google and say it circumvents anti-distillation mechanisms. Decoding 315,320 reasoning blocks scraped from public repositories, where developers share session logs unaware of what the blocks contain, recovered 367 personal-information artifacts and 182 credentials, and encrypted blocks also offer a route to invisible prompt injection.

After responsible disclosure the three vendors reportedly acknowledged the findings and patched, which per Willison makes reproduction impossible (arXiv 2608.09867, Willison). It matters because encryption of reasoning was a design choice with its own attack surface. It is the technical counterpart to the distillation campaigns of B08-14, and it qualifies Anthropic's documented design in which the signature carries the full thinking and no setting returns the raw chain of thought (B08-06). Sources: arXiv 2608.09867 · Willison, 2026-08-11

Read it in the deep dive