Signal

OpenAI disrupts alleged Moonshot-linked model-distillation campaign

OpenAI says it blocked a large coordinated effort to extract reasoning behavior from its models, attributing a core cluster to people associated with Moonshot AI while leaving broader attribution uncertain.

1 min read
arXiv paper on extracting protected reasoning traces from proprietary model APIs
Independent researchers document reasoning-trace extraction paths · Credit: Panfilov et al. View source

OpenAI reported on September 30 that it disrupted a coordinated model-distillation campaign after activity that began July 1 surged on July 24 and 25. The company counted 16,000 requests using a relevant extraction pattern from more than 4,000 users and said related prompt patterns appeared across more than 15,000 users before disruption on July 28.

Attribution remains qualified

OpenAI attributes a core cluster to individuals associated with Moonshot AI, the developer of Kimi, but says it is unclear whether all observed operators came from one actor. The report concerns attempted extraction through model interactions; it does not describe a breach of OpenAI's encrypted systems, databases, or stored user conversations.

The attack class is independently documented

Independent researchers have shown ways to recover protected reasoning traces from proprietary model APIs, supporting the broader plausibility of reasoning-trace extraction. That paper does not validate OpenAI's campaign counts or attribution, which remain based on the company's telemetry and investigation.

Sources