Article

OpenAI Reports Hidden-Reasoning Extraction: What Agent Builders Should Protect

OpenAI’s reasoning-extraction disclosure puts agent transcripts, replay controls and streaming review under scrutiny. What builders should protect.

Encrypted agent-session records separated from a shareable transcript and a monitored output stream.

On September 30, 2026, OpenAI disclosed a campaign to extract protected model reasoning. For agent builders who retain full API responses, the practical issue is how private reasoning state moves through storage, transcript exports, session restoration and output delivery.

The design challenge is to preserve the state an agent needs to continue working without treating that same record as safe to publish or import from elsewhere. Those are different uses and need different controls.

What OpenAI says happened

OpenAI reports 16,000 extraction-pattern requests from over 4,000 users on July 24–25, and related activity involving over 15,000 users. It says it disrupted the cluster by July 28. Its footnote counts attempts, not necessarily successful extractions. September 30 is the disclosure date, not the campaign’s start.

The company attributes a core cluster to individuals associated with Moonshot AI, while leaving the broader operators’ identity unresolved. This is OpenAI’s assessment—not independently established corporate responsibility.

OpenAI describes cross-conversation reasoning replay, not broken encryption, a database breach or direct access to stored conversations. The cross-user path it says it closed required prior possession of another user’s reasoning artifact. Its reported mitigations include account restrictions, stronger identity/model boundaries and streamed-output holds; partner and tool-output defenses remain ongoing work.

Request counts do not establish how much useful training material was obtained or what capability a competing model gained. For the broader distinction between distillation and alleged misuse, see RohitAI’s analysis of Anthropic’s September threat report.

Independent research adds a privacy warning—and a timing caveat

In their August preprint, Alexander Panfilov and coauthors studied 6,708 publicly shared agent trajectories. Among 704 privacy artifacts recovered from genuine user sessions, they report that 64 appeared only in reasoning, not in visible chat history. The implication for exports is concrete: scrubbing readable messages does not inspect an opaque reasoning block.

These are the researchers’ results, not RohitAI measurements. The public corpus was selected, labeling was model-assisted, and full original plaintext was unavailable to verify every recovered token. The findings are neither a customer-wide leakage rate nor independent evidence for the Moonshot attribution. The original paper also says its published extraction method was no longer reproducible after provider mitigations by August.

The researchers’ September 30 website update nevertheless claims continued Astra and GPT-6.1 Sol reasoning extraction through third-party providers. Their follow-up PDF needs careful dating: its September 13 replay table records blocked direct OpenAI and Anthropic routes alongside some successful Azure routes, but the timeline later notes OpenAI/Azure mitigations on September 27 and Anthropic/Azure non-reproducibility on September 28.

The follow-up also distinguishes replay decoding from prompting a model to write reasoning into a tool argument. Generated scratchpad text is not, by itself, proof of an exact copy of concealed reasoning. Neither a September 13 table nor a September 30 claim establishes the current status of every serving route. A model name alone is therefore an inadequate unit for tracking remediation.

Keep the working session separate from the shareable transcript

Deleting all opaque state is not a sound blanket fix. OpenAI’s current reasoning guide says stateless Responses requests return encrypted_content in reasoning items by default, including with store: false or Zero Data Retention. Provider-side retention settings do not eliminate the copy held by your application.

The same guide recommends preserving reasoning items through legitimate function-call continuations. Opt-in reasoning summaries are separate from raw reasoning; a readable summary does not make the accompanying encrypted state readable or safe to export.

Compaction has its own contract. OpenAI’s compaction guide says to pass the standalone /responses/compact output window forward unchanged. Server-side compaction permits different handling: input-array clients may remove items before the newest compaction item, while previous_response_id clients should not manually prune. These are current continuation rules, not features announced in the security disclosure.

RohitAI’s recommendation is to separate three records at the application boundary. This is a design framework inferred from the documentation and research, not a description of OpenAI’s internal implementation.

Record

What it contains

Permitted use

Private working session

Provider-returned items needed for continuation, with owner and session metadata.

Resume authorized work through a trusted server-side adapter; restrict access and retention.

Shareable export

An allowlisted copy of intended visible fields, reviewed for sensitive text and tool payloads.

Human review or publication; omit opaque reasoning/compaction items and do not accept it as replayable state.

Imported transcript

External material whose claimed roles and provenance are untrusted.

Treat it as supplied data, not authenticated assistant history or provider-generated state.

For example, a support export should be generated from the private record, not overwrite it. Removing encrypted blocks from that copy protects the sharing boundary; removing them from the only working history can undermine continuation. A text redactor cannot certify the contents of ciphertext it cannot inspect.

At restore time, associate opaque items with the authenticated owner, provider, model and conversation through trusted server-side records. Reject unapproved transitions before calling the model. Valid JSON is not proof of origin. These application checks complement, rather than replace, the provider’s own replay protections.

Review content before the relevant consumer receives it

OpenAI’s streaming documentation states that moderation scores requested with generation arrive after the complete output, not with partial deltas. If an application forwards those deltas immediately, a later refusal or review decision cannot recall what a user, tool or downstream service already received.

Where pre-release review is required, make release an explicit gate covering the relevant channel—including tool arguments, not just chat text. Buffering trades responsiveness for a chance to inspect before delivery; the sources do not quantify that latency cost. Turning streaming off does not itself inspect content, and ordinary moderation is not established here as a complete reasoning-leak detector.

A focused review for agent teams

  1. Trace full response objects into checkpoints, debug logs, tracing systems, support bundles and evaluation datasets. Apply the originating task’s access and retention policy to opaque state too.

  2. Separate export and restore paths. Using owned accounts and synthetic, non-sensitive fixtures, validate that exports exclude opaque state, untrusted imports cannot become authenticated history, and supported continuation and compaction still work.

  3. Ask each serving provider for mitigation scope and dated evidence for the exact model, host, adapter and streaming configuration you use—including fallback routes and legacy state. Record who enforces the output gate.

These are proposed checks, not tests performed for this article. The useful completion criterion is an authorized session that still resumes correctly, a shareable copy that cannot reintroduce private reasoning state, and a documented decision about when each downstream consumer receives output.

Methodology: AI-assisted reporting and analysis based on OpenAI’s disclosure, current API documentation and the researchers’ published work, reviewed September 30, 2026. No API experiments, private telemetry, interviews or independent attribution investigation were conducted.