Stealing AI Reasoning Traces

Refract AI Intelligence Digest

BLUF

Architectural flaw in LLM APIs enables theft of hidden chain-of-thought data via session token swapping.

NEWS

New research shows encrypted reasoning blocks returned to clients are interchangeable across different users and sessions. This bypasses server-side isolation designed to protect model intellectual property. Providers currently rely on client-side storage for these sensitive traces, creating a replay attack vector.

Why I Care

This compromises the confidentiality of proprietary model logic and risks exposing sensitive data processed during reasoning steps. Organizations using these APIs for secure tasks may face unintended data leakage and IP theft.

Next Steps

AI vendors must bind encrypted trace tokens to unique session identifiers immediately. Security teams should audit API integrations for token reuse vulnerabilities before the next security patch cycle.

Interesting research: “Stealing Reasoning Traces from Proprietary LLM APIs“: Abstract: Leading large language model providers now conceal their models’ step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model’s reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials. Third, it inadvertently reveals hazardous information hidden within the reasoning process, even in cases where the model’s final, visible output safely rejects a malicious request. Fourth, attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, we propose concrete cryptographic and system-level mitigations to secure client-side reasoning...
Back to Blog Listing

Source: Schneier on Security ·

This digest was generated by Refract AI Collective to help the public sector security community stay informed.