// uncategorised

OpenAI Disrupted a Reasoning-Extraction Campaign. The Interesting Part Is the Architectural Flaw, Not the OpenAI vs Moonshot Fight

OpenAI says it identified and disrupted a coordinated campaign that pulled protected reasoning, the model’s internal chain of thought, out of its models. The core of the activity, going back to the first week of July, was attributed to individuals associated with the Chinese company Moonshot AI, though without technical evidence, likely for security reasons. The headlines went to “OpenAI versus Moonshot,” but if you build or deploy LLMs the mechanism matters more: the operators did not break encryption or touch a database, they manipulated model interactions so that protected reasoning was reproduced in visible form, at scale.

That matters because underneath the incident sits an architectural problem that is not unique to OpenAI.

What happened

OpenAI calls it adversarial distillation: the systematic, unauthorized use of one model’s outputs to train, reproduce, or improve another. The timeline:

  • Activity began on July 1, 2026, at low volume.
  • On July 24 and 25 it spiked to 16,000 requests using a specific extraction pattern, from more than 4,000 users.
  • Further investigation found similar prompt-pattern activity across more than 15,000 users.
  • The campaign was fully disrupted on July 28, fraudulent accounts were banned, and additional mitigations were deployed.

Separately, OpenAI closed a pathway that let someone who already held another user’s encrypted reasoning replay it and recover the contents, and added checks that detect and hold streamed output that might expose reasoning.

The architectural flaw that affects everyone

Here is the part worth reading even if you do not care who stole what from whom. A study published in August 2026 (MATS Research, ELLIS Institute Tübingen, and Synk) described an architectural vulnerability affecting Claude, Gemini, and GPT. The encrypted reasoning traces turned out to be fully compatible and interchangeable across different sessions, users, and models within a single provider’s ecosystem.

That enables a concrete technique. Take an encrypted reasoning trace from a strong model, inject it into a weaker and less safeguarded model from the same provider, and the weak model decodes and outputs the trace verbatim, in plaintext. You never jailbreak the capable model directly. The attack goes around the anti-distillation defense sideways, through the weak link in the same ecosystem.

Three consequences follow, and they go well beyond intellectual-property theft:

  • Large-scale private data extraction. If traces are interchangeable across users, there is a path to pull out what was supposed to stay private.
  • Invisible prompt injection. A malicious payload can be hidden entirely inside encrypted blocks that a human never reads but the model processes.
  • Hazardous information leaking from the reasoning process. Even when the model’s final, visible output rejects a harmful request, dangerous detail can remain in the reasoning and be extracted.

Why this is a security problem, not just a business one

Protected reasoning shows how a model works its way to an answer. Extract it and you can both reach sensitive data and help others reproduce the model’s capabilities. OpenAI states plainly that adversarial distillation carries safety and national-security risks: extracted reasoning can be used to train another model without carrying over the safeguards applied to the original’s user-facing output. At scale, distillation accelerates the transfer of advanced capabilities without the same investment in safety, and that gets more serious as models gain ability in dual-use domains.

Put simply: the successor model inherits the capabilities of the original without its brakes. The original’s final answers passed through safety filters. The raw reasoning behind them did not, and that is exactly what gets pulled out.

Not the first time with Moonshot

Distillation accusations are not new to Moonshot AI. Last month Anthropic said Moonshot was quietly relaying its customers’ requests to Claude instead of processing them with its own Kimi model, then showing Claude’s responses back to users. The company was also alleged to have retained a subset of those exchanges to train its chain-of-thought model. That activity is tracked as GTG-16002.

What to take from this

If you deploy LLMs or agents built on them, the takeaway is not who is right in the OpenAI and Moonshot dispute, it is the class of attack itself:

  • An encrypted block is not a trusted block. Since an injection can be hidden inside encrypted reasoning that a human will not see but the model will execute, “encrypted” does not mean “safe.” Everything reaching the model, opaque blocks included, should be treated as potentially hostile input.
  • The weak model in an ecosystem is the way into the strong one. This attack bypassed the strong model’s defenses through its less-safeguarded sibling. If you run several models from one provider, their security is bounded by the weakest link, not the strongest.
  • Reasoning is an attack surface, not just a feature. Hazardous information can live in the chain of thought even when the visible answer is clean. You have to test not only what the model outputs, but what it hides inside.

Testing your own deployments against exactly these techniques, injection inside opaque blocks, leakage through reasoning, bypass via a weaker sibling model, is what LLM jailbreak and guardrail testing, prompt injection testing, and AI agent penetration testing are for.

The takeaway

The story is framed as a corporate fight, but the technical fact under it matters more than the names involved: protected reasoning at major providers turned out to be portable across sessions, users, and models, and that opens distillation of someone else’s capabilities without their safeguards, invisible prompt injection, and leakage of hazardous content from the reasoning process. OpenAI closed a specific pathway and banned accounts, but the class of attack has not gone anywhere. For anyone building on LLMs the lesson is simple: encrypted does not mean trusted, a model stack is only as safe as its weakest member, and a model’s reasoning has to be tested as an attack surface of its own.


// get started

Work with AgentOffense

Tell us about your target and goals. We’ll reply with scope and a fixed-price quote — usually within one business day.

./request_engagement