For most of the history of machine learning security, stealing a model sounded like a conventional cybersecurity problem. An attacker would compromise an internal system, obtain a checkpoint, copy proprietary training data, or exfiltrate model weights. That threat has not disappeared, but the frontier AI industry is increasingly dealing with a more uncomfortable possibility: an attacker may not need to penetrate the infrastructure at all. If a powerful model is available through an API, every response can potentially become a small piece of training data for another model.
That is the security problem now being described as adversarial distillation.
The term moved sharply into the mainstream of AI security in 2026. On September 30, OpenAI disclosed that it had disrupted a coordinated campaign attempting to extract protected reasoning from its models. OpenAI described the activity as adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to train, reproduce, or improve another model. According to the company, the operators did not compromise an internal database or break the encryption protecting stored conversations. Instead, they manipulated interactions with the models themselves in an effort to reproduce protected reasoning in user-visible form. ओपनएआई
That distinction matters. The attack surface is not simply the infrastructure surrounding the model. The model interface itself becomes part of the extraction surface.
And that changes how AI systems need to be defended.
What Is Adversarial Distillation?
Knowledge distillation is not inherently malicious. It is one of the most established techniques in machine learning. A larger or more capable model acts as a “teacher,” while a smaller “student” model learns from the teacher’s outputs. Instead of reproducing the full training process of the teacher, the student learns a compressed approximation of its behavior.
There are many legitimate reasons to do this. Companies can use distillation to create smaller models with lower inference costs, reduce latency, deploy models on constrained devices, or specialize a general model for a narrow domain. The Frontier Model Forum explicitly distinguishes these authorized uses from adversarial distillation. Its February 2026 issue brief describes adversarial distillation as the unauthorized replication of model capabilities, frequently through covert or indirect access to a teacher model’s outputs. Frontier Model Forum
The security problem begins when an organization uses another provider’s model as an unwilling teacher.
Imagine a frontier model with exceptional performance in software engineering. An attacker does not necessarily need its weights. Instead, the attacker can construct a large set of programming tasks, submit them to the model, store the answers, and use the resulting dataset to fine-tune another model.
Conceptually, the pipeline looks simple:
prompts → frontier model → high-quality outputs → training dataset → student model
The important word here is capability. Adversarial distillation is usually not trying to create a bit-for-bit reconstruction of the victim model. The attacker wants to reproduce enough of a valuable behavior that the student becomes commercially, strategically, or operationally useful.
A model might therefore be partially cloned even when its architecture, original dataset, optimizer state, and parameters remain completely secret.
That is why the boundary between model extraction और model distillation attack is increasingly blurred. Google Threat Intelligence Group describes model extraction attacks as situations in which an adversary systematically probes an accessible model and transfers the information obtained from those interactions into another model using knowledge distillation. Google reported an increase in this kind of activity during 2025 and early 2026. गूगल क्लाउड

Adversarial Distillation Is Not the Same as Stealing Model Weights
The language surrounding model theft can be confusing because several technically different attacks are often described using the same vocabulary.
A traditional model compromise might involve obtaining actual model parameters. That could happen through an infrastructure intrusion, exposed storage, insecure deployment environment, malicious insider, supply-chain compromise, or another conventional attack against the systems holding the model.
Adversarial distillation does not require that.
The student model can use a completely different architecture. It may have fewer parameters. Its weights will not match the teacher’s weights. Its training corpus may be entirely different.
What matters is functional imitation.
This idea predates modern large language models by many years. In their influential 2016 USENIX Security paper, Florian Tramèr and co-authors demonstrated that attackers with black-box query access could reproduce the functionality of machine-learning models exposed through prediction APIs. Their work showed that publicly accessible prediction interfaces could reveal considerably more information about the underlying model than developers might expect. यूज़ेनिक्स
Three years later, Orekondy, Schiele and Fritz demonstrated the same basic principle against more complex image models in Knockoff Nets. Their attacker queried a target model, collected input-output pairs and trained a separate model that reproduced significant portions of the original system’s functionality—even when the attacker had no knowledge of the target’s training data or architecture. Open Access CVF
Natural-language systems proved vulnerable as well. The 2020 research Thieves on Sesame Street! showed that BERT-based NLP services could be approximated using only API access. Strikingly, the researchers found that attackers did not necessarily need the original training dataset or even normal, well-formed language examples to obtain useful training signals from the victim model. arXiv
The foundations of adversarial distillation were therefore visible long before today’s frontier-model race.
What changed was the value of the thing being extracted.
Frontier Models Made Distillation Far More Valuable
An image classifier from 2019 might represent months of engineering work and valuable intellectual property. A frontier reasoning model represents something much larger: enormous computational investment, proprietary training pipelines, reinforcement-learning infrastructure, evaluation systems, human feedback, synthetic data generation, safety research, and accumulated experimentation.
If part of that capability can be transferred by querying the finished model rather than reproducing the entire development process, the economics become extremely attractive.
This is particularly true for capabilities that are difficult to obtain simply by increasing pretraining compute.
Advanced reasoning is one example. Agentic software engineering is another. Long-horizon planning, computer use, scientific reasoning, multimodal interpretation and sophisticated evaluation all potentially represent high-value targets.
The Frontier Model Forum specifically identifies mathematical and scientific reasoning, coding and tool use, multimodal processing and general reasoning as categories that may be targeted through adversarial distillation. Frontier Model Forum
This is one reason modern distillation attacks increasingly look less like indiscriminate scraping and more like structured data collection.
An attacker does not necessarily ask millions of random questions.
The attacker may instead construct a curriculum.
Thousands of debugging tasks can target software-engineering ability. Carefully designed mathematical problems can target reasoning. Tool-use scenarios can teach an agent how to plan actions. Candidate solutions followed by model critiques can generate preference data. Model-generated rubrics can turn the teacher into an evaluator or reward signal.
The victim model becomes part teacher, part dataset generator, and sometimes part automated grader.
That is significantly more powerful than merely copying answers.
Why Chain-of-Thought and Protected Reasoning Became High-Value Targets
The most important evolution in adversarial distillation is the growing interest in reasoning traces.
Earlier model extraction attacks mostly treated a model as a function. Submit an input, collect the prediction, reproduce the mapping.
Reasoning models expose a much richer opportunity.
Suppose a teacher receives a difficult problem and returns only the final answer. That answer contains useful supervision, but relatively little information about how the solution was produced.
If an attacker also obtains detailed intermediate reasoning, the training example becomes far richer.
Instead of:
problem → answer
the attacker may obtain something closer to:
problem → reasoning trajectory → verification steps → answer
A student trained on large numbers of such examples can potentially learn more than the surface-level output distribution.
This is one reason frontier AI providers increasingly avoid exposing unrestricted internal reasoning.
The Frontier Model Forum identifies chain-of-thought exfiltration, chain-of-thought critiquing and model-based grading as particularly significant forms of adversarial distillation. In these scenarios, the attacker is interested not only in what the teacher concludes but in how the model evaluates alternatives, identifies mistakes and moves through difficult reasoning tasks. Frontier Model Forum
The attack disclosed by OpenAI in September 2026 pushed that problem further.
OpenAI said operators attempted to recover protected reasoning using novel interaction patterns. One technique involved taking encrypted reasoning associated with one interaction and presenting it to a model in another interaction with instructions intended to recover or transcribe the concealed content. OpenAI emphasized that this did नहीं mean attackers had broken the underlying encryption or compromised stored conversations. It was an interaction-level extraction problem rather than a conventional cryptographic breach. ओपनएआई
That difference is extremely important for defenders.
Encryption can successfully protect stored information while the application layer still allows that information to reappear through an unexpected path.
AI systems therefore need to treat reasoning artifacts not merely as hidden text, but as security-sensitive objects whose entire lifecycle must be controlled.
The September 2026 OpenAI Incident
OpenAI’s September disclosure provides one of the clearest public examples of what modern adversarial distillation can look like operationally.
According to the company, suspicious activity began on July 1, 2026 at relatively low volume. On July 24 and July 25, OpenAI observed approximately 16,000 requests using a relevant extraction pattern from more than 4,000 users. A broader investigation connected related activity to a cluster exceeding 15,000 users, which the company says it had disrupted by July 28. OpenAI explicitly notes that these numbers describe attempted extractions and do not imply that every attempt succeeded. ओपनएआई
The scale is revealing because it demonstrates why adversarial distillation cannot be detected only by examining individual prompts.
One request may look perfectly reasonable.
Ten requests may still look reasonable.
Even hundreds of prompts distributed among unrelated-looking accounts may not immediately trigger a conventional security rule.
The malicious pattern may only become visible once the provider correlates behavior across thousands of identities, sessions and pieces of infrastructure.
OpenAI said it could not determine that every operator belonged to a single actor. However, the company attributed what it called a core cluster of activity to individuals associated with Moonshot AI, developer of the Kimi model family. That attribution should therefore be understood as OpenAI’s assessment rather than an independently established fact. ओपनएआई
OpenAI responded by restricting associated accounts, strengthening signup and infrastructure controls, expanding monitoring, hardening reasoning protection and coordinating with third-party services. The company also said it shared findings through the Frontier Model Forum and government information-sharing channels. ओपनएआई
The incident illustrates a broader shift in AI security: the defender is no longer protecting a single API key from abuse. It may need to identify a distributed extraction campaign whose individual components resemble legitimate AI usage.
Anthropic Describes Industrial-Scale Distillation Campaigns
OpenAI is not the only frontier model provider raising the alarm.
On February 23, 2026, Anthropic published its own investigation into what it described as industrial-scale distillation campaigns. The company said it had identified activity associated with DeepSeek, Moonshot and MiniMax involving more than 16 million exchanges across approximately 24,000 fraudulent accounts. Anthropic said the campaigns targeted some of Claude’s most valuable capabilities, including agentic reasoning, coding, tool use, data analysis and computer use. मानवजनित
Again, attribution claims made by an affected vendor should be described as that vendor’s findings rather than treated automatically as independently verified evidence. But the operational characteristics Anthropic describes are highly relevant regardless of attribution.
The campaigns allegedly used large numbers of accounts, proxy infrastructure and distributed access paths. Traffic was structured to acquire training examples associated with specific capabilities rather than to use the model as an ordinary end user would.
Anthropic’s report describes activity in which models were used not just to generate answers but also to create reasoning data, judge responses and support reinforcement-learning pipelines. मानवजनित
That is where adversarial distillation begins to resemble an AI development pipeline built on top of someone else’s model.
The frontier model is no longer merely answering questions.
It is being turned into infrastructure for training its potential competitor.
Google Has Observed the Same Attack Class
Google Threat Intelligence Group reached similar conclusions in a February 2026 report.
GTIG said it had observed increasing numbers of model extraction attempts and described distillation attacks as a method through which organizations attempt to clone proprietary logic and reasoning capabilities. Google noted that these attacks can operate through legitimate model access rather than conventional intrusion techniques. गूगल क्लाउड
One disclosed campaign involved attempts to coerce Gemini into exposing richer reasoning information. Google said it identified more than 100,000 prompts associated with the activity and assessed that the campaign appeared designed to reproduce reasoning ability across a broad range of tasks and languages. Google reported that its defensive systems recognized and mitigated the activity. गूगल क्लाउड
This matters because OpenAI, Anthropic and Google are describing essentially the same defensive problem from different vantage points.
The victim is a high-value model.
The attacker obtains access that may initially appear legitimate.
Large numbers of carefully chosen queries are submitted.
Responses are collected.
Those responses become training, evaluation or reinforcement-learning material for another model.
The dangerous operation is distributed across a training pipeline rather than contained inside one obviously malicious request.
Adversarial Distillation Versus Jailbreaking
It is tempting to classify adversarial distillation as another form of jailbreak, but the two attacks have different goals.
A jailbreak primarily tries to change what a model will provide to the attacker.
Adversarial distillation tries to transfer what the model knows or can do into another system.
A jailbreak is often successful if one forbidden answer is obtained. Distillation usually requires accumulation. The attacker’s advantage comes from collecting many high-quality examples over time.
That said, the two attacks can overlap.
A jailbreak may help expose outputs that are especially useful for distillation. An attacker could attempt to bypass restrictions on reasoning, safety boundaries, coding tasks or other protected behavior in order to improve the training dataset.
But defeating the jailbreak does not necessarily defeat the distillation attack.
Perfectly normal answers can still carry substantial information about a model’s capabilities.
This is the fundamental difficulty of defending a model offered as a service: the system must reveal useful behavior in order to be useful.
Every successful response leaks some information about the function being provided.
Adversarial Distillation Versus Prompt Injection
Prompt injection is another adjacent but distinct problem.
In prompt injection, an adversary tries to manipulate the instructions guiding an AI system. This becomes particularly dangerous in agents that can access external tools, files, websites or internal systems.
Adversarial distillation instead targets the model as an information source.
A prompt-injection attack might say, in effect, “ignore the application instructions and perform another action.”
A distillation campaign says, “continue behaving normally, but do it often enough and systematically enough that I can learn from the results.”
This is why ordinary content moderation is not sufficient.
A request does not need to contain malware, hate speech, credential theft or obviously suspicious instructions to participate in an extraction campaign.
Many of the most useful training prompts are completely benign.
The maliciousness exists in the aggregate behavior.
Why Detecting Adversarial Distillation Is Hard
Traditional API abuse controls tend to look for straightforward signals: abnormal request volume, suspicious IP addresses, repeated failures, credential sharing or clearly prohibited content.
Distillation attackers have strong incentives to avoid exactly those patterns.
Requests can be distributed among many accounts. Prompt wording can vary. Access can come through proxies or third-party infrastructure. Different account types can be mixed together. Request volumes can remain below conventional per-account thresholds.
A single account may therefore appear unremarkable.
What defenders need to understand is whether many identities collectively behave like one training operation.
That makes adversarial distillation partly a graph-analysis problem.
Accounts can be related through timing, infrastructure, payment instruments, signup patterns, request structure, overlapping prompt distributions, common target capabilities and similarities in how output is subsequently consumed.
Anthropic says the campaigns it investigated were distinguishable by combinations of repetitive structures, highly concentrated capability targets, synchronized activity and infrastructure relationships. मानवजनित
The defensive unit of analysis therefore cannot always be the request.
Sometimes it must be the campaign.

A Model Can Be Stolen Without Being Reproduced Exactly
One misconception surrounding adversarial distillation is that a student model must achieve near-identical benchmark performance before the attack becomes meaningful.
That standard is too high.
Consider a frontier system that is exceptionally strong in ten different domains. A competitor may care about only one.
If the victim is unusually capable at autonomous code repair, an attacker can focus extraction almost entirely on code repair. If scientific reasoning is valuable, the attacker can build a curriculum around scientific reasoning. If the target is an agent platform, tool-selection behavior may be more valuable than general conversation quality.
The result may be a student that is substantially worse overall while still reproducing enough of the targeted capability to have major economic value.
This also explains why general benchmark comparisons are not necessarily good indicators of whether distillation occurred.
A distilled capability can be narrow.
A student could absorb one strategically important behavior while remaining completely different from the teacher across the rest of its distribution.
Adversarial Distillation Is Becoming an AI Supply-Chain Problem
The longer-term security implications extend beyond individual AI labs.
Modern AI applications rarely call only one model through one interface. Developers increasingly rely on model routers, coding agents, enterprise gateways, inference providers, cloud platforms and orchestration systems.
Each additional layer complicates the question of who is actually interacting with the model.
A provider may think a request comes from a legitimate routing service. The router may believe it is serving an ordinary customer. The customer’s application may itself be automatically generating tasks.
This creates an attribution problem.
It also creates a data-governance problem.
If third-party services forward user conversations to another model for evaluation, transformation or synthetic-data production, users may not understand how far their prompts travel.
The result is that adversarial distillation can intersect with privacy, data residency and enterprise confidentiality issues even when the original goal of the operation is capability extraction rather than user surveillance.
The AI supply chain therefore needs provenance not only for model weights and training datasets, but increasingly for model-generated data.
Organizations may eventually need to answer questions such as: Which model produced this synthetic dataset? Under what license? Through which API? Were the outputs authorized for training? Can their origin be independently demonstrated?
That is a much broader security problem than model theft.
The Most Valuable Asset May No Longer Be the Weights Alone
The AI industry spent years treating model weights as the crown jewels.
They still are.
But adversarial distillation suggests that the effective intellectual property of a frontier model exists at several layers simultaneously.
There are weights, training data, system prompts, reinforcement-learning infrastructure, reward models, evaluators, tool policies, reasoning behavior, safety policies and the behavioral distribution exposed through the API.
A provider may successfully protect the first six while leaking useful information through the last one.
This does not mean API access makes model security impossible. It means the security objective needs to change.
The goal cannot be “prevent the model from revealing any information about itself.” A useful model necessarily reveals information through its behavior.
The realistic objective is to make unauthorized capability transfer expensive, detectable, attributable and incomplete.
That is a much more subtle engineering problem.
Why Rate Limiting Alone Will Not Stop Adversarial Distillation
Rate limiting remains useful, but it is not a complete defense.
A naive attacker sending a million queries through one API key can obviously be throttled.
A sophisticated extraction campaign can distribute the workload.
Ten thousand accounts each submitting a modest number of requests may collectively generate an enormous dataset while remaining below local thresholds.
This is why provider disclosures increasingly emphasize behavioral correlation across accounts and infrastructure rather than simple per-user quotas.
Effective detection needs to consider whether ostensibly separate identities display suspiciously similar training behavior.
For example, defenders may look for unusual concentration around a narrow set of capabilities, synchronized submission patterns, repeated task templates, systematic generation of dataset-like examples and coordinated account creation.
No single feature proves maliciousness.
The confidence comes from correlation.
Output Protection Is Becoming as Important as Input Filtering
Generative AI security has historically focused heavily on the input.
Is the prompt malicious?
Is it attempting a jailbreak?
Does it contain prohibited content?
Is an external document injecting hidden instructions?
Adversarial distillation forces defenders to spend more time thinking about the output side of the system.
What information is the model returning?
Does the output contain protected reasoning?
Can a hidden artifact be replayed elsewhere?
Can tool results accidentally expose internal context?
Can repeated interaction reconstruct information that no single response reveals?
OpenAI’s September 2026 mitigation work reflects this shift. The company said it strengthened checks around streamed output that could reveal reasoning and closed a pathway that could allow someone already possessing another user’s encrypted reasoning artifact to replay it and attempt recovery. ओपनएआई
This resembles a familiar principle from conventional security: sensitive data needs protection throughout its entire lifecycle, not only when stored.
The same principle now applies to reasoning.
Can Watermarking Solve Model Distillation?
Watermarking is frequently proposed as a solution to model extraction, but its role is narrower than it first appears.
A watermark may help establish that a suspicious model’s behavior or training data originated from a particular source. That can improve attribution, forensic analysis or legal enforcement.
But watermarking does not automatically prevent extraction.
A sufficiently sophisticated student may reproduce capabilities without preserving an obvious output signature. Attackers may filter training data, paraphrase responses, combine outputs from multiple models or perform additional training that weakens the watermark.
Earlier model-extraction research has already shown the difficulty of building attribution mechanisms that remain robust against adaptive attackers. The BERT extraction research from 2020, for example, examined defenses including API watermarking and found that stronger adversaries could circumvent some proposed mechanisms. arXiv
Watermarking is therefore more useful as one layer of a defense and provenance system than as a standalone anti-distillation control.
What an Adversarial Distillation Defense Stack Looks Like
No single defensive mechanism is likely to solve the problem.
A practical strategy begins with identity and access controls because large-scale extraction becomes easier when attackers can cheaply create or acquire accounts. Signup abuse, fraudulent payment methods, credential resale and proxy access therefore become part of model security.
The next layer is behavioral detection. Providers need systems capable of connecting apparently unrelated activity and identifying whether requests collectively resemble dataset generation or capability extraction.
Output controls provide another layer. Sensitive reasoning, internal metadata and tool artifacts should not become retrievable simply because an attacker moves them between sessions or changes the way they are presented.
Model-level safeguards can also help make certain extraction strategies less useful without destroying legitimate product quality.
Finally, intelligence sharing matters because attackers can target several frontier providers at once. An infrastructure pattern blocked by one company may appear against another days later.
The fact that the Frontier Model Forum published a dedicated adversarial distillation issue brief in February 2026 is significant. The subject has moved beyond isolated model-security research and into coordinated industry defense. Frontier Model Forum
The Safety Problem Is More Complicated Than Intellectual Property
Most discussion around model extraction begins with intellectual property, and understandably so. Training frontier models requires enormous investment.
But capability transfer creates another problem.
The student does not necessarily inherit the teacher’s safety architecture.
A frontier provider may have spent substantial resources building policy enforcement, abuse monitoring, post-training safety behavior and infrastructure-level restrictions around a model. If another system learns the underlying capability from the teacher’s outputs but does not reproduce those controls, capability and safety can separate.
Both the Frontier Model Forum and major model providers have raised this concern. OpenAI argues that extracted reasoning could be used to train another model without preserving safeguards associated with the original system. ओपनएआई
This becomes especially important as models acquire stronger capabilities in cybersecurity, autonomous software development, scientific research and other dual-use fields.
Distillation can potentially transfer what a model can do without transferring the restrictions governing when it should do it.
That is why adversarial distillation is increasingly treated as an AI safety issue as well as an IP-security issue.
The Legal Boundary Is Less Simple Than the Technical Boundary
It is important not to collapse several different questions into one.
Technically, distillation means transferring behavior or knowledge from one model to another.
Contractually, a model provider may prohibit certain uses of outputs through its terms of service.
From an intellectual-property perspective, the situation can depend on jurisdiction, licensing, the nature of the copied material and the details of the training process.
Security teams should therefore avoid assuming that every instance of one model learning from another is automatically equivalent to criminal theft.
The safer technical distinction is authorization.
Authorized distillation is performed by the model owner, under an appropriate license or with permission.
Adversarial or illicit distillation involves systematic capability extraction without that authorization, often accompanied by attempts to evade controls or conceal the true nature of the activity.
That distinction also prevents defenders from accidentally treating legitimate researchers, enterprise users and developers as attackers merely because their workloads resemble model-training activity.
There Is Another Meaning of “Adversarial Distillation”
One source of confusion around the keyword adversarial distillation is that the phrase existed in machine-learning research before its recent use in frontier-model security.
For example, a 2023 paper in the Journal of Information Security and Applications used “adversarial distillation” in a different context: studying adversarial features and the vulnerability of object-detection models to adversarial examples. ScienceDirect
The contemporary frontier-AI security meaning is different.
Since 2026, organizations including the Frontier Model Forum and OpenAI have increasingly used adversarial distillation to describe unauthorized model-capability extraction through outputs or reasoning. ओपनएआई
For security practitioners, that newer definition is quickly becoming the more operationally important one.
Why AI Security Teams Should Care Even If They Do Not Train Frontier Models
Adversarial distillation is not only a problem for OpenAI, Anthropic or Google.
Any organization exposing a valuable machine-learning capability through an API can potentially face the same fundamental risk.
A financial company might operate a proprietary fraud-scoring model.
A security vendor may provide an AI malware-analysis service.
A healthcare company might expose a specialized medical model.
A startup may have fine-tuned an open model into an unusually capable coding or research assistant.
If competitors can systematically query that system and reproduce the differentiated behavior cheaply enough, the model becomes an extraction target.
Google explicitly warns that organizations operating custom models should monitor API traffic for extraction patterns, not just conventional abuse. गूगल क्लाउड
The more expensive it was to create the capability, the greater the incentive to copy it through the interface rather than reproduce the entire training process.
Adversarial Distillation Changes AI Red Teaming
AI red teaming has traditionally concentrated on questions such as whether the model can be jailbroken, whether prompt injection can redirect an agent, whether sensitive information can be leaked, and whether the model can generate dangerous content.
Those questions remain important.
But frontier systems now need another class of evaluation:
Can normal access to the system be turned into a scalable capability-extraction channel?
That question is different because the red team must examine many interactions together.
Testing should therefore explore whether a model can be systematically coerced into generating unusually rich training material, whether hidden reasoning artifacts can cross trust boundaries, whether coordinated accounts can evade extraction detectors, and whether the system distinguishes legitimate high-volume workflows from dataset harvesting.
The emphasis shifts from isolated prompts to attack economics.
A distillation defense may technically be bypassable, but if it makes extraction ten times more expensive, considerably noisier and easier to attribute, it can still provide meaningful security value.
The Next Model Security Battle Is Over Capability Provenance
There is a broader issue emerging behind adversarial distillation.
The generative AI ecosystem is producing enormous quantities of synthetic training data.
Model A generates examples.
Model B judges them.
Model C improves them.
Model D trains on the resulting dataset.
After enough iterations, it becomes difficult to know where a particular capability originally came from.
This creates a provenance problem.
Future model-security systems may therefore need mechanisms capable of answering not only “was this output generated by AI?” but more sophisticated questions such as “which model family contributed to this training distribution?” and “was this capability obtained through authorized use?”
Those questions are extraordinarily difficult.
But as synthetic data becomes a core part of frontier model training, they may become unavoidable.
Adversarial distillation makes provenance a security primitive.
Adversarial Distillation May Become One of the Defining AI Security Problems
The most important thing about adversarial distillation is not that distillation itself is new.
It is not.
Model extraction has been studied for at least a decade. Researchers demonstrated API-based model stealing long before ChatGPT existed, and black-box functionality cloning was already well established by the end of the 2010s. यूज़ेनिक्स
What is new is the strategic value of the models being queried.
A modern frontier system can encode capabilities generated through enormous amounts of compute, data, post-training and research. When access to that system is sold through an API, an attacker may be able to convert a fraction of those capabilities back into training data.
That creates an unusual security paradox.
AI companies need models to expose intelligence in order to sell intelligence.
But every exposed behavior teaches the outside world something about the system.
The security challenge is therefore not to stop the model from revealing anything. That would make the model useless.
The challenge is to control the rate, richness and context in which capabilities can be transferred, detect when legitimate access becomes systematic extraction, and make large-scale cloning substantially more difficult than ordinary product use.
OpenAI’s September 2026 disclosure demonstrates why this is no longer a theoretical problem. Anthropic and Google have independently described large-scale extraction activity during the same year, while the Frontier Model Forum now treats adversarial distillation as a dedicated frontier-model security issue. ओपनएआई
For years, the most valuable secret in AI security was assumed to be the model file.
The emerging lesson from adversarial distillation is more uncomfortable.
Sometimes you do not need to steal the model.You can teach another model by talking to it.

