A public compendium · updated daily
How frontier intelligence is compressed, priced, and contested.
Distillation moves capability from a large teacher model into a small, cheap student. It is the quiet engine behind most models people actually pay for, the subject of an open dispute between the largest labs, and a live regulatory question in three jurisdictions. This compendium tracks all of it.
Timeline · compiled 3 September 2026 · 111 sources
Cross-cutting timeline of AI distillation, 2006 to 2026
Knowledge distillation began as an academic model-compression trick (Bucila 2006, Hinton 2015) and spent a decade as a research topic before becoming the default way to build small language models (DistilBERT 2019, Gemma, Llama 3.2, Qwen3). In 2024 it turned into a cloud product line (OpenAI Model Distillation, Amazon Bedrock Model Distillation) and a pricing strategy (Gemini Flash, GPT-4o mini, Haiku). DeepSeek-R1's January 2025 release, with six openly distilled students, triggered a market shock and the first public accusation that a rival had distilled a frontier lab's API outputs. By 2026 distillation was a national-security issue: Anthropic, OpenAI and Google all disclosed industrial-scale extraction campaigns within eleven days of each other in February, the White House issued National Security and Technology Memorandum 4 on adversarial distillation, and Congress advanced the Deterring American AI Model Theft Act and the BLADE Act while Beijing publicly rejected the whole framing. This timeline records 130 dated events across research, product, market, policy and legal categories, each with a source URL.
Key figures · 6 figures
Dated events in this timeline
130 events
57 in 2026 so far
Every event carries its own source URL; spans 2006-08-20 to 2026-09-03.
Research milestones
38 events
29% of the record
Papers and public technical reports, from Model Compression (2006) to reasoning-trace extraction (2026).
Product and market events
60 events
52 product, 8 market
Model launches, distillation-as-a-service products, pricing tiers and priced market reactions.
Policy and legal events
32 events
32 of them since January 2025
Bills, memoranda, export-control actions, lab disclosures and disputes. Almost none predate the DeepSeek-R1 release.
Events in the last 12 months
65 events
50% of the whole timeline
From 3 September 2025 to 3 September 2026 — the period covering NSTM-4, the BLADE Act and the three labs’ extraction disclosures.
Span from first to last event
20 years
2006-08-20 → 2026-09-03
First: Bucila, Caruana and Niculescu-Mizil publish 'Model Compression' at KDD 2006. Last: OpenAI launches GPT-6 Astra at $10/$50 per million tokens.
Key findings · 8 findings
Era 1 (2006-2018): distillation as model compression
Bucila, Caruana and Niculescu-Mizil showed in 2006 that a small network trained on the pseudo-labels of a large ensemble could match it; Hinton, Vinyals and Dean formalised soft-target 'dark knowledge' in 2015 and FitNets, sequence-level KD (Kim and Rush) and Born-Again Networks extended it to deeper students, seq2seq models and self-distillation. The technique was framed as deployment engineering, not competition. In parallel, the security literature was already describing the same mechanic as an attack: Tramer et al. stole model functionality through prediction APIs in 2016.
Era 2 (2019-2022): the transformer compression wave
DistilBERT and TinyBERT made distillation the standard recipe for shipping BERT-class models to production, cutting parameters by 40-90% while keeping most accuracy. Hugging Face's DistilBERT became one of the most downloaded models on the Hub and normalised distilled checkpoints as first-class open artifacts. The same window produced the first NLP model-extraction attack papers (Krishna et al., 2019) and self-distillation for vision (DINO, 2021).
Era 3 (2023-2024): black-box distillation from frontier APIs
Stanford Alpaca (52K text-davinci-003 demonstrations for under $600), Vicuna (ShareGPT conversations), Orca (GPT-4 explanation traces) and Distilling Step-by-Step showed that a small open model could inherit behaviour from a closed API using only outputs, not logits. MiniLLM and GKD reframed the loss for generative students (reverse KL, on-policy sampling), while Gudibande et al. warned that imitation closes style gaps faster than capability gaps. Frontier labs responded by hiding chain-of-thought (o1, September 2024) and by selling distillation themselves (OpenAI Model Distillation, October 2024; Bedrock, December 2024).
Era 4 (2025): reasoning distillation and the DeepSeek shock
DeepSeek-R1 (20 January 2025) shipped six distilled dense students from 1.5B to 70B and proved that reasoning traces transfer cheaply; Sky-T1 ($450), s1 (1,000 examples), LIMO (817 examples) and OpenThoughts pushed the data floor lower. Nvidia lost nearly $600B of market value in one session, and within 48 hours OpenAI and the White House alleged DeepSeek had distilled OpenAI outputs. Gemma 3, Llama 4 and Qwen3 all disclosed teacher-student training as core recipe, and CMU's Antidistillation Sampling opened a defensive research line.
Era 5 (2026): adversarial distillation becomes statecraft
Three US labs published extraction evidence within eleven days: OpenAI's memo to the House Select Committee (12 February), Google's Threat Intelligence Group report on a 100,000-prompt Gemini campaign (12 February) and Anthropic's disclosure of 16M+ Claude exchanges from ~24,000 fraudulent accounts (23 February). Anthropic later told the Senate Banking Committee about a 28.8M-exchange campaign it attributes to Alibaba's Qwen lab. OSTP issued NSTM-4 (23 April), the House Foreign Affairs Committee advanced H.R. 8283 43-0, the White House accused Moonshot of distilling Anthropic's Fable for Kimi K3 (22 July), and senators introduced the BLADE Act (5 August). Beijing's Ministry of Commerce rejected the whole framing on 27 July.
By 2026 distillation is inside almost every frontier training run, hostile or not
Hugging Face's mid-2026 survey found teacher-student training in Gemma 4, DeepSeek-V4 (per-domain RL specialists merged by on-policy distillation), Nvidia Nemotron 3 Ultra (10+ domain teachers), Xiaomi MiMo-V2-Flash, GLM-5 (earlier checkpoints as teachers), Qwen3 and Cursor Composer 2.5 (self-distillation). Meta's Muse Glimmer 30B was distilled from the closed Muse Spark, and Apple said its 2026 foundation models are refined with Gemini outputs under a paid partnership. The technique that governments want to sanction is the same one that ships nearly every model on the market.
The legal status of distillation is still unsettled
As of September 2026 no court has ruled on whether API-output distillation is unlawful, and neither Anthropic nor OpenAI has sued any of the accused labs; the disputes rest on terms-of-service breaches, fraud (fake accounts, proxies) and export-control policy rather than copyright or patent. Lawfare's July 2026 analysis argues existing law already covers the fraudulent-access part and warns that broad new IP rights would entrench incumbents. Bills such as H.R. 8283 and S. 5252 explicitly preserve 'legitimate' distillation and rely on public attacker lists, sanctions and export controls instead of new IP rights. The one filed case so far is a securities class action against Alibaba, not an IP suit.
Defences moved from hiding outputs to watermarking and shared threat intelligence
The first defence was concealment: OpenAI hid o1's chain of thought in September 2024. Research defences followed (Antidistillation Sampling, April 2025) and by 2026 the response was operational: Anthropic's classifiers and behavioural fingerprinting, Google's real-time prompt-cluster detection, and an April 2026 arrangement in which OpenAI, Anthropic and Google exchange distillation attack signatures through the Frontier Model Forum. Anthropic's August 2026 text watermark is explicitly a provenance measure for EU transparency compliance rather than an anti-distillation control, and an August 2026 paper showed encrypted reasoning blocks could be replayed across models to recover hidden traces.
Charts · 3 charts
Distillation events per year, stacked by category
events| Year | Research events | Product events | Market events | Policy events | Legal events |
|---|---|---|---|---|---|
| 2006 | 1 | 0 | 0 | 0 | 0 |
| 2014 | 1 | 0 | 0 | 0 | 0 |
| 2015 | 1 | 0 | 0 | 0 | 0 |
| 2016 | 2 | 0 | 0 | 0 | 0 |
| 2018 | 1 | 0 | 0 | 0 | 0 |
| 2019 | 3 | 0 | 0 | 0 | 0 |
| 2020 | 1 | 0 | 0 | 0 | 0 |
| 2021 | 1 | 0 | 0 | 0 | 0 |
| 2022 | 1 | 0 | 0 | 0 | 0 |
| 2023 | 8 | 0 | 0 | 0 | 0 |
| 2024 | 3 | 12 | 0 | 0 | 0 |
| 2025 | 11 | 16 | 2 | 7 | 2 |
| 2026 | 4 | 24 | 6 | 15 | 8 |
Shows the shift from a purely research topic (2006-2022) to a product story (2024) and then a policy and legal story (2025-2026). 2026 is partial: it covers 1 January to 3 September 2026.
Sources: huggingface.co · cnas.org
Cumulative distillation events, 2006 to September 2026
events| Year | Cumulative events events |
|---|---|
| 2006 | 1 |
| 2014 | 2 |
| 2015 | 3 |
| 2016 | 5 |
| 2018 | 6 |
| 2019 | 9 |
| 2020 | 10 |
| 2021 | 11 |
| 2022 | 12 |
| 2023 | 20 |
| 2024 | 35 |
| 2025 | 73 |
| 2026 | 130 |
The curve is close to flat for the first fifteen years and then near-vertical from 2024, when distillation simultaneously became a cloud product, a pricing tier and a national-security file.
Sources: arxiv.org · anthropic.com
Share of timeline events by category
events| Category | Events |
|---|---|
| Research | 38 |
| Product | 52 |
| Market | 8 |
| Policy | 22 |
| Legal | 10 |
Product and research events dominate the record, but policy and legal events are concentrated almost entirely in the 20 months from January 2025.
Sources: iaps.ai
Tables · 1 table
Dated distillation events by year and category, 2006-2026
13 rows| Year | Research events | Product events | Market events | Policy events | Legal events | Total events |
|---|---|---|---|---|---|---|
| 2006 | 1 | 0 | 0 | 0 | 0 | 1 dl.acm.org |
| 2014 | 1 | 0 | 0 | 0 | 0 | 1 arxiv.org |
| 2015 | 1 | 0 | 0 | 0 | 0 | 1 arxiv.org |
| 2016 | 2 | 0 | 0 | 0 | 0 | 2 arxiv.org |
| 2018 | 1 | 0 | 0 | 0 | 0 | 1 arxiv.org |
| 2019 | 3 | 0 | 0 | 0 | 0 | 3 arxiv.org |
| 2020 | 1 | 0 | 0 | 0 | 0 | 1 arxiv.org |
| 2021 | 1 | 0 | 0 | 0 | 0 | 1 arxiv.org |
| 2022 | 1 | 0 | 0 | 0 | 0 | 1 arxiv.org |
| 2023 | 8 | 0 | 0 | 0 | 0 | 8 crfm.stanford.edu |
| 2024 | 3 | 12 | 0 | 0 | 0 | 15 en.wikipedia.org |
| 2025 | 11 | 16 | 2 | 7 | 2 | 38 novasky-ai.github.io |
| 2026 | 4 | 24 | 6 | 15 | 8 | 57 en.wikipedia.org |
Derived from this file's timeline array; every underlying event carries its own source URL. Coverage is denser for 2025-2026 because distillation only became a policy and market story then, so early years understate research activity.
Sources: arxiv.org · huggingface.co · cnas.org
Timeline · 130 events
Bucila, Caruana and Niculescu-Mizil publish 'Model Compression' at KDD 2006
The paper shows that a single small neural network trained on unlabeled data pseudo-labeled by a large ensemble can match the ensemble's accuracy at a fraction of the size. It is the earliest widely cited ancestor of modern knowledge distillation.
Source: dl.acm.orgFitNets introduces hint-based distillation into intermediate layers
Romero et al. train thin, deep students by matching the teacher's intermediate representations ('hints') rather than only its outputs, establishing feature-based distillation.
Source: arxiv.orgHinton, Vinyals and Dean post 'Distilling the Knowledge in a Neural Network'
The paper coins 'distillation', introduces temperature-scaled soft targets and shows ensemble knowledge can be compressed into one model on MNIST and speech. It becomes the canonical citation for the field.
Source: arxiv.orgKim and Rush propose sequence-level knowledge distillation
Sequence-level KD trains a student on the teacher's beam-search outputs instead of per-token distributions, enabling small neural machine translation models. This is the direct ancestor of training on generated reasoning traces.
Source: arxiv.org'Stealing Machine Learning Models via Prediction APIs' presented at USENIX Security
Tramer, Zhang, Juels, Reiter and Ristenpart show that an adversary with only query access to a pay-per-prediction ML API can reconstruct a near-equivalent model, including against commercial services. It is the founding paper of model extraction, the adversarial framing of what labs would later call distillation.
Source: usenix.orgBorn-Again Neural Networks show self-distillation improves students of equal size
Furlanello et al. find that a student distilled from an identically sized teacher can outperform it, complicating the compression-only view of distillation.
Source: arxiv.orgTinyBERT distills BERT with transformer-layer and attention matching
Huawei researchers combine general and task-specific distillation to build a 4-layer BERT that retains most of BERT-base accuracy at 7.5x fewer parameters.
Source: arxiv.orgHugging Face releases DistilBERT
DistilBERT is 40% smaller and 60% faster than BERT-base while retaining about 97% of its language-understanding performance. It makes distilled checkpoints a mainstream open-source artifact.
Source: arxiv.org'Thieves on Sesame Street!' extracts BERT-based APIs for a few hundred dollars
Krishna, Tomar, Parikh, Papernot and Iyyer show that random word sequences plus task heuristics are enough to extract a working copy of a fine-tuned BERT API, and that membership classification and API watermarking both fail against sophisticated adversaries. The paper is the direct NLP ancestor of the 2026 distillation-attack disclosures.
Source: arxiv.orgMobileBERT: task-agnostic distillation for on-device transformers
Sun et al. distil a specially designed inverted-bottleneck BERT-large teacher into a 25M-parameter student that is 4.3x smaller and 5.5x faster than BERT-base, establishing distillation as the standard route to edge NLP.
Source: arxiv.orgDINO frames self-supervised vision learning as self-distillation with no labels
Caron et al. (Meta AI) train a student to match a momentum teacher's softmax outputs on different crops of the same image, producing ViT features that segment objects without labels. Self-distillation becomes a general representation-learning tool, not just a compression trick.
Source: arxiv.orgSelf-Instruct bootstraps instruction data from a model's own generations
Wang et al. generate instructions, inputs and outputs from GPT-3 itself and fine-tune on the filtered result, closing most of the gap to InstructGPT. The pipeline is what Alpaca reuses three months later against a proprietary teacher.
Source: arxiv.orgStanford Alpaca fine-tunes LLaMA 7B on 52K text-davinci-003 demonstrations
Data generation cost under $500 via the OpenAI API and fine-tuning under $100. Alpaca popularises black-box distillation from a proprietary API and prompts OpenAI's terms-of-service clause against training competing models to receive wide attention.
Source: crfm.stanford.eduVicuna-13B fine-tunes LLaMA on ~70K ShareGPT ChatGPT conversations
LMSYS reports about 90% of ChatGPT quality using user-shared conversations scraped from ShareGPT, for roughly $300 of training compute. Vicuna makes conversation-log distillation the default way to build an open chat model in 2023.
Source: lmsys.orgGoogle's 'Distilling Step-by-Step' uses LLM rationales as extra supervision
Hsieh et al. train small task-specific models on chain-of-thought rationales extracted from a large model, outperforming standard fine-tuning with less data.
Source: arxiv.org'The False Promise of Imitating Proprietary LLMs' pushes back on API distillation
Gudibande, Wallace, Snell and co-authors (Berkeley) show that imitation models learn a teacher's style faster than its capability: human raters are fooled, benchmarks are not. The paper frames the capability gap as closable only with vastly more imitation data or a better base model, and is still the standard caution cited against pure black-box distillation.
Source: arxiv.orgMicrosoft's Orca learns from GPT-4 explanation traces
A 13B model trained on step-by-step explanation traces from GPT-4 approaches ChatGPT on several benchmarks, showing that imitation of reasoning, not just answers, transfers capability.
Source: arxiv.orgMiniLLM proposes reverse-KL distillation for generative language models
Gu et al. argue forward KL is ill-suited to open-ended generation and train students with reverse KL on their own samples, improving over standard sequence-level KD.
Source: arxiv.orgGeneralized Knowledge Distillation (GKD) introduces on-policy distillation
Agarwal et al. (Google DeepMind) train the student on its own generated sequences with teacher feedback, fixing train-inference distribution mismatch. On-policy distillation becomes the dominant recipe by 2025-2026.
Source: arxiv.orgAdversarial Diffusion Distillation ships SDXL Turbo: 1-step image generation
Sauer et al. (Stability AI) combine an adversarial loss with score distillation to sample a large diffusion model in 1-4 steps, matching SDXL in four. Distillation moves from language into real-time generative media, a parallel track that later produces LCM, Flux and video-model distillations.
Source: arxiv.orgGoogle releases Gemma; the 2B model is distilled from 7B
Gemma's first generation establishes teacher-student training inside Google's open-model line, a pattern that continues through Gemma 2, 3 and 4.
Source: en.wikipedia.orgAnthropic announces the Claude 3 family including Claude 3 Haiku
Haiku is positioned as the fastest, cheapest tier of a three-tier lineup, cementing the industry pattern of a small model priced far below its flagship sibling.
Source: anthropic.comGoogle says Gemini 1.5 Flash was trained by distillation from 1.5 Pro
At I/O 2024 Google explicitly describes Flash as distilled from Pro, the first time a frontier lab named distillation as the recipe for its flagship low-cost tier. Launch pricing was $0.35 per million input tokens.
Source: blog.googleGemma 2: 9B distilled from 27B, 2B from an unreleased 7B teacher
Google scales distillation into Gemma 2, training smaller models on soft targets from larger ones instead of next-token prediction alone.
Source: en.wikipedia.orgOpenAI launches GPT-4o mini at $0.15/$0.60 per million tokens
GPT-4o mini undercuts Claude 3 Haiku ($0.25/$1.25) and Gemini 1.5 Flash ($0.35/$0.70) and becomes the default student model for OpenAI's later distillation API.
Source: simonwillison.netNvidia's Minitron: compact models via pruning plus distillation
Muralidharan et al. prune Nemotron-4 15B and retrain with distillation using up to 40x fewer tokens per model, reporting 1.8x compute savings for a model family.
Source: arxiv.orgLlama-3.1-Minitron 4B: 'LLM Pruning and Distillation in Practice'
Nvidia applies the Minitron recipe to Llama 3.1 8B and Mistral NeMo 12B, retraining pruned models with 94B tokens of distillation. The models are released on Hugging Face for commercial use.
Source: huggingface.coOpenAI releases o1-preview with a hidden chain of thought
OpenAI forbids attempts to reveal o1's reasoning and cites competitive advantage alongside safety. Hiding traces is widely read as the first commercial anti-distillation defence.
Source: en.wikipedia.orgMeta's Llama 3.2 1B and 3B built by pruning 8B and distilling from 8B/70B logits
Meta states it used single-shot structured pruning from Llama 3.1 8B and incorporated logits from the 8B and 70B models as token-level targets during pre-training.
Source: ai.meta.comOpenAI launches Model Distillation in the API at DevDay
Stored Completions, Evals and fine-tuning are combined into a workflow for distilling GPT-4o or o1-preview outputs into GPT-4o mini. OpenAI offered 2M free daily training tokens on GPT-4o mini through 31 October 2024.
Source: infoq.comAnthropic releases Claude 3.5 Haiku
The small tier is refreshed alongside the upgraded Claude 3.5 Sonnet and the computer-use beta.
Source: en.wikipedia.orgDeepSeek previews R1-Lite, its first reasoning model
DeepSeek-R1-Lite-Preview shows visible chain-of-thought two months before the full R1 and its distilled family.
Source: en.wikipedia.orgAmazon Bedrock Model Distillation enters preview
AWS automates teacher generation, data synthesis and student fine-tuning within the same model family (Anthropic, Meta Llama, Amazon Nova), claiming up to 5x faster and 75% cheaper models with under 2% accuracy loss on RAG tasks.
Source: aws.amazon.comMicrosoft Phi-4 technical report: a 14B model that surpasses its teacher on STEM QA
Phi-4 leans on synthetic data generated with GPT-4-class models, and Microsoft reports it exceeds the teacher on STEM question answering, a milestone for 'weak-to-strong' distillation.
Source: arxiv.orgDeepSeek-V3 released: 671B MoE trained in 2.788M H800 GPU-hours
The technical report (arXiv 27 December) discloses the full training budget, roughly $5.6M at assumed rental prices, which seeds the January 2025 market narrative about cheap frontier models.
Source: arxiv.orgSky-T1-32B-Preview: an o1-class reasoner trained for $450
UC Berkeley's NovaSky team fine-tunes Qwen2.5-32B on 17K traces from QwQ-32B-Preview, reaching 43.3% on AIME 2024 and 82.4% on MATH500.
Source: novasky-ai.github.ioBIS publishes the Framework for Artificial Intelligence Diffusion
The Biden-era interim final rule creates a three-tier country system for advanced chips and, for the first time, controls exports of closed frontier model weights. Compliance was set for 15 May 2025.
Source: wiley.lawDeepSeek-R1 released with six openly distilled students (1.5B to 70B)
Alongside R1 and R1-Zero, DeepSeek publishes dense models distilled from R1 onto Qwen 2.5 and Llama 3 bases under an MIT licence. The paper reports the distilled students beat RL-from-scratch at the same size.
Source: arxiv.orgDeepSeek shock: Nvidia loses nearly $600B in market value in one day
Nvidia falls about 17% and the Nasdaq about 3% as investors reprice AI capex on the belief that cheap distilled and open models reduce demand for frontier compute.
Source: cnbc.comMicrosoft and OpenAI probe DeepSeek-linked API data exfiltration
Bloomberg reports Microsoft security researchers observed individuals believed linked to DeepSeek pulling large volumes of output through OpenAI's API in autumn 2024.
Source: usnews.comOpenThoughts project launches with a 114K-example open reasoning dataset
The Bespoke Labs/Stanford/Berkeley collaboration scales up Bespoke-Stratos and releases OpenThinker-7B, aiming to open-source the reasoning-distillation data recipe.
Source: open-thoughts.aiOpenAI says it has evidence DeepSeek distilled its models; David Sacks cites 'substantial evidence'
OpenAI tells the Financial Times it saw signs of distillation by DeepSeek in violation of its terms; White House AI adviser David Sacks repeats the claim on television. It is the first public accusation of adversarial distillation between frontier labs.
Source: eweek.comGemini 2.0 Flash reaches general availability
Google's second-generation Flash tier becomes the default Gemini model, continuing the distilled-flagship pricing strategy introduced with 1.5 Flash.
Source: en.wikipedia.orgs1: Simple test-time scaling distills 1,000 Gemini Thinking traces
Muennighoff et al. fine-tune Qwen2.5-32B on the s1K set (traces from Gemini 2.0 Flash Thinking) and add 'budget forcing', exceeding o1-preview on competition math.
Source: arxiv.orgLIMO: 817 curated examples yield 63.3% on AIME24
Ye et al. argue reasoning is elicited rather than taught, reaching 95.6% on MATH500 with about 1% of the data used by prior distillation recipes.
Source: arxiv.orgApple publishes 'Distillation Scaling Laws'
Busbridge et al. fit a scaling law for student performance as a function of teacher and student compute, giving compute-optimal distillation recipes for when a teacher exists and when it must be trained.
Source: arxiv.orgNvidia reports record $39.3B quarter one month after the DeepSeek sell-off
Q4 FY2025 revenue rose 78% year on year and data-center revenue hit $35.6B; CEO Jensen Huang argues reasoning models add a new scaling law for compute.
Source: nvidianews.nvidia.comGemma 3: all sizes (1B to 27B) trained with knowledge distillation
Google DeepMind's technical report states every Gemma 3 model is distilled; Gemma3-4B-IT is reported competitive with Gemma2-27B-IT.
Source: arxiv.orgOpenAI's AI Action Plan submission calls DeepSeek 'state-controlled' and urges bans
In its OSTP response OpenAI recommends barring 'PRC-produced' models in Tier 1 countries, citing IP-theft risk after its distillation allegations.
Source: techcrunch.comMeta's Llama 4 Maverick is codistilled from the 2T-parameter Behemoth
Meta describes a novel distillation loss that dynamically weights soft and hard targets, run during Behemoth's own pre-training to amortise cost. Behemoth itself was never released.
Source: ai.meta.comGemini 2.5 Flash released
Google's first Flash model with configurable thinking budgets ships two months after 2.5 Pro.
Source: en.wikipedia.orgAntidistillation Sampling: the first decoding-time defence against reasoning distillation
Savani, Trockman, Feng and co-authors (CMU) perturb the teacher's next-token distribution so that its reasoning traces poison a downstream student while the teacher's own utility is preserved. The paper (NeurIPS 2025) opens a defensive research line that continues through ADS-C and 'The Distillation Game' in 2026.
Source: arxiv.orgQwen3 launches with 'strong-to-weak' distillation for models from 0.6B to 14B
Alibaba distills its 235B-A22B and 32B flagships into smaller students using off-policy then on-policy distillation, reporting better results than RL at roughly one-tenth the GPU hours.
Source: arxiv.orgPhi-4-reasoning distills o3-mini reasoning traces into a 14B model
Microsoft fine-tunes Phi-4 on curated 'teachable' prompts with demonstrations generated by OpenAI's o3-mini.
Source: arxiv.orgAmazon Bedrock Model Distillation reaches general availability
GA adds Nova Premier/Pro, Claude 3.5 Sonnet v2 and Llama 3.3/3.2 pairings plus function-calling data augmentation for agent use cases.
Source: aws-news.comCommerce rescinds the AI Diffusion Rule two days before its compliance date
BIS says the tiered framework would have stifled US innovation and burdened allies; it issues separate guidance warning that using US chips to train Chinese models risks enforcement.
Source: akingump.comAnthropic releases Claude Opus 4 and Sonnet 4
The Claude 4 generation resets the tiering that later distilled-tier releases (Haiku 4.5) are benchmarked against.
Source: en.wikipedia.orgOpenThoughts3: 1.2M-example distillation dataset from QwQ-32B
OpenThinker3-7B scores 53% on AIME 2025 and 54% on GPQA Diamond, beating DeepSeek-R1-Distill-Qwen-7B by 15-20 points via data-recipe ablations alone.
Source: arxiv.orgEU AI Office publishes the final General-Purpose AI Code of Practice
Three chapters (transparency, copyright, safety and security) give GPAI providers a voluntary path to compliance ahead of the 2 August obligations.
Source: lw.comWhite House releases America's AI Action Plan
Ninety-plus actions across innovation, infrastructure and international security; its IP-protection language is later cited as the basis for anti-distillation legislation.
Source: whitehouse.govEU AI Act obligations for general-purpose AI models begin to apply
New GPAI models placed on the EU market must meet transparency and copyright duties; Commission enforcement powers start 2 August 2026 and legacy models have until 2027.
Source: lw.comOpenAI ships gpt-oss-120b and gpt-oss-20b, its first open weights since GPT-2
Both models are Apache 2.0 and run on a single GPU or a 16GB laptop. Releasing open weights removes the API terms-of-service barrier for distillation from an OpenAI-lineage model, and gpt-oss quickly becomes a teacher and a base for third-party students.
Source: openai.comClaude Opus 4.1 released
An incremental flagship update that precedes Anthropic's autumn 4.5 lineup.
Source: en.wikipedia.orgOpenAI launches the GPT-5 family: GPT-5, GPT-5 mini ($0.25/$2) and GPT-5 nano ($0.05/$0.40)
Mini and nano tiers are priced 5x and 25x below GPT-5, extending the three-tier flagship/distilled pattern to OpenAI's fifth generation.
Source: openrouter.aiDeepSeek V3.1 adds hybrid thinking and non-thinking modes
DeepSeek merges its chat and reasoning lines into one model, mirroring Qwen3's design.
Source: en.wikipedia.orgAnthropic bars companies majority-owned by Chinese entities, citing distillation risk
The policy blocks any entity more than 50% owned by companies headquartered in unsupported regions such as China, explicitly naming distillation as a way adversaries could advance their own models. Press coverage followed on 5 September.
Source: anthropic.comClaude Sonnet 4.5 released
Anthropic's mid-tier model becomes the reference point for Haiku 4.5's cost-performance claims two weeks later.
Source: en.wikipedia.orgClaude Haiku 4.5: Sonnet 4-level coding at one-third the cost
Priced at $1/$5 per million tokens, Haiku 4.5 is marketed as matching Sonnet 4 on coding at more than twice the speed, the clearest small-tier value claim of 2025.
Source: anthropic.comThinking Machines: on-policy distillation reaches RL parity at 50-100x less compute
Kevin Lu and colleagues report 7-10x fewer gradient steps than RL and cite Qwen's 1,800 vs 17,920 GPU-hour comparison on AIME'24, attributing the gain to dense per-token reward.
Source: thinkingmachines.aiGoogle launches Gemini 3 Pro and Deep Think
Gemini 3 Pro debuts at 1501 Elo on LMArena and 91.9% GPQA Diamond; it becomes the teacher for the 3 Flash tier released a month later.
Source: blog.googleClaude Opus 4.5 released
Anthropic closes 2025 with its fourth flagship refresh of the year.
Source: en.wikipedia.orgDeepSeek V3.2 released with DeepSeek Sparse Attention
The last V3-line model before V4; it follows the V3.2-Exp preview of 29 September.
Source: en.wikipedia.orgGemini 3 Flash released
Google ships the small tier of its third generation one month after 3 Pro.
Source: en.wikipedia.orgApple and Google announce a multi-year Gemini partnership for Apple Foundation Models
Bloomberg had reported a roughly $1B-per-year deal for a custom 1.2T-parameter Gemini; Apple later says its own models are refined using outputs from Gemini frontier models, a sanctioned, paid form of distillation.
Source: en.wikipedia.orgXiaomi's MiMo-V2-Flash introduces Multi-Teacher On-Policy Distillation (MOPD)
The 309B-total / 15B-active MoE is post-trained by training domain-specialist RL teachers (math, code, agentic, safety) and merging them into one student with dense token-level rewards plus outcome reward models. MOPD is the template that Nvidia's Nemotron 3 and DeepSeek-V4 follow later in 2026.
Source: arxiv.orgClaude Opus 4.6 released
First of Anthropic's 2026 releases; MiniMax markets M2.5 the same month as roughly 1/20th the cost of Opus 4.6.
Source: en.wikipedia.orgZhipu/Z.ai releases GLM-5, a 744B open-weight model with cross-stage self-distillation
GLM-5's final post-training stage is on-policy distillation from its own earlier SFT and RL checkpoints, used to recover capabilities that later RL stages had degraded. It is the clearest production example of a model distilling from past versions of itself.
Source: arxiv.orgOpenAI memo to the House Select Committee on the CCP accuses DeepSeek of ongoing distillation
In 'Updated Stakes for American-Led, Democratic AI', OpenAI says DeepSeek-linked accounts built code to programmatically harvest outputs, used obfuscated third-party routers and unauthorized resellers, and were still distilling ahead of a new model launch.
Source: restofworld.orgGoogle Threat Intelligence Group discloses a 100,000-prompt distillation campaign against Gemini
GTIG says it disrupted commercially motivated model-extraction activity, including one cluster of more than 100,000 prompts engineered to force Gemini to emit reasoning content in the user's language, and attributes activity to actors linked to China, Russia and North Korea. CNAS later valued the extracted material at 1-2.5 billion tokens.
Source: nbcnews.comMiniMax ships M2.5 and M2.5 Lightning at roughly 1/20th of Claude Opus 4.6 pricing
M2.5 is listed at $0.15/$1.20 per million tokens (Lightning $0.30/$2.40) with 80.2% on SWE-bench Verified. It is the sharpest 2026 example of a Chinese open-weight model undercutting a US frontier tier on price while the same lab is named in Anthropic's distillation disclosure eleven days later.
Source: openrouter.aiAlibaba releases Qwen3.5, a 397B-parameter Apache 2.0 MoE with 1M context
The generation immediately preceding the April-June campaign that Anthropic later attributes to Qwen-linked operators.
Source: en.wikipedia.orgClaude Sonnet 4.6 released
Mid-tier refresh that lands six days before Anthropic's distillation-attack disclosure.
Source: en.wikipedia.orgAnthropic discloses industrial-scale distillation attacks by DeepSeek, Moonshot and MiniMax
About 24,000 fraudulent accounts generated over 16 million exchanges: MiniMax 13M+ (agentic coding), Moonshot 3.4M+ (agentic reasoning, computer use), DeepSeek 150K+ (reasoning and reward-model data), organised into 'hydra cluster' proxy networks of 20,000+ accounts. Anthropic announces classifiers, behavioral fingerprinting, tighter research-account controls and indicator-sharing with other labs.
Source: anthropic.comIAPS publishes 'AI Distillation Attacks: The Case for Targeted Government Intervention'
The policy memo argues for narrowly scoped intervention that punishes fraudulent-access distillation without chilling legitimate teacher-student training, and becomes the analytic reference for the April congressional and executive actions.
Source: iaps.aiOpenAI releases GPT-5.4; mini and nano follow on 17 March
GPT-5.4 mini goes to free-tier users while nano stays API-only. Both small tiers are priced about four times above their GPT-5 equivalents, the first time the distilled tier moved up rather than down in price.
Source: en.wikipedia.orgAI Foundation Model Transparency Act of 2026 (H.R. 8094) introduced
Rep. Don Beyer's bill would require disclosure of training data and methods for foundation models, indirectly touching teacher-student provenance.
Source: congress.govGoogle releases Gemma 4 (E2B, E4B, 26B A4B, 31B), again trained with distillation
Google's release log dates the first Gemma 4 checkpoints to 31 March 2026, with the public announcement on 2 April; multi-token-prediction variants follow on 16 April and a 12B unified model on 3 June. Hugging Face's mid-2026 survey describes post-training as distillation from a large instruction-tuned teacher, extending the Gemma 2/3 recipe.
Source: ai.google.devOpenAI, Anthropic and Google begin sharing distillation attack signatures via the Frontier Model Forum
Bloomberg reports the three labs exchange query-distribution patterns, prompt structures, IP fingerprints and account-creation behaviour the way cyber-threat intelligence is shared, so a signature detected at one lab can be flagged at the others within hours. Antitrust uncertainty limits how much can be exchanged.
Source: techbrew.comDeterring American AI Model Theft Act (H.R. 8283) introduced by Rep. Huizenga
The bill creates a public 'AI Model Extraction Attackers List', discretionary sanctions and a Commerce/State information-sharing channel, while stating that legitimate distillation remains a valuable research tool.
Source: huizenga.house.govHouse Select Committee hearing: 'China's Illicit Campaign to Steal and Subvert American AI'
Witness testimony describes coordinated distillation campaigns using proxy accounts and jailbreaks against US frontier models.
Source: docs.house.govHouse Foreign Affairs Committee marks up H.R. 8283 and reports it favorably, 43-0
A technical amendment swaps State Department references for Commerce. The unanimous vote one week after introduction signals bipartisan consensus on treating adversarial distillation as a sanctionable act. IAPS records the favorable report on 23 April.
Source: iaps.aiOSTP issues National Security and Technology Memorandum 4 on adversarial distillation
Director Michael Kratsios says foreign entities, principally in China, run 'deliberate, industrial-scale campaigns' to distill US frontier AI systems using tens of thousands of proxies; agencies are told to share intelligence, co-develop defences and explore accountability measures.
Source: defenseone.comState Department cables posts and issues a demarche to Beijing over distillation
US diplomats are instructed to warn host governments about model extraction by DeepSeek, Moonshot AI and MiniMax, and a formal protest is delivered to Beijing.
Source: justsecurity.orgDeepSeek previews V4-Flash (284B) and V4-Pro (1.6T), trained with on-policy distillation from domain experts
Per Hugging Face's survey, each domain (math, code, agents) gets an RL-trained specialist teacher and the unified model is trained with reverse-KL against them.
Source: huggingface.coHouse Homeland Security and Select Committee on China open a joint investigation into PRC AI models
Chairmen Garbarino and Moolenaar send letters to Anysphere (Cursor) and Airbnb, citing 'unauthorized model distillation and other illicit techniques' by DeepSeek, Alibaba, Moonshot AI and MiniMax and questioning Cursor Composer 2's use of a Moonshot open-weight base. The probe later expands to DoorDash.
Source: homeland.house.govCursor ships Composer 2.5, trained by self-distillation with textual feedback
The same model acts as teacher and student: hint-conditioned outputs supervise unhinted ones through per-token KL, and on-policy distillation repairs localized failures (bad tool calls, premature stops) without rewriting whole rollouts. Cursor reports 25x more synthetic tasks than Composer 2 and near-frontier coding scores at lower token price.
Source: cursor.comClaude Opus 4.8 released
Last Opus 4.x release before the Claude 5 generation; falls inside the window of the alleged Qwen-linked extraction campaign (22 April to 5 June 2026).
Source: en.wikipedia.orgApple says WWDC 2026 foundation models are refined using outputs from Gemini frontier models
Craig Federighi describes Apple's third-generation models as trained on proprietary data and refined with Gemini outputs under the January partnership, an openly licensed instance of the same technique at the centre of the enforcement fights.
Source: en.wikipedia.orgMiniMax launches M3 'frontier coding' model
Four months after Anthropic's report, MiniMax ships a 1M-context multimodal coding model; MiniMax did not publicly rebut the distillation findings.
Source: en.wikipedia.orgCNAS publishes 'Adversarial Distillation' report
Remler and Hayum estimate the 16M Claude exchanges could represent 150-400 billion extracted tokens (against DeepSeek-R1's entire 6.4B-token SFT set) and map a six-actor supply chain including token mixers and transfer stations such as CloseAI, BianXie AI, One-API, New-API, OpenRouter and Eden AI.
Source: cnas.orgNvidia releases Nemotron 3 Ultra, a 550B open model built by multi-teacher on-policy distillation
More than ten domain-specialised teachers were trained separately and consolidated into the 550B-total / 55B-active hybrid Mamba-MoE student through dense token-level guidance on student-generated rollouts. The technical report is dated 9 June 2026.
Source: research.nvidia.comAnthropic releases Claude Fable 5 and Claude Mythos 5
Fable 5 is the model the White House later says Moonshot distilled for Kimi K3; Mythos 5 is limited-access.
Source: en.wikipedia.orgEuropean Commission publishes the Code of Practice on Transparency of AI-Generated Content
The voluntary code operationalises AI Act Article 50 marking and labelling duties; around 190 organisations, including Anthropic, had signed by 31 July 2026. It is the direct cause of Claude's text watermark two months later.
Source: digital-strategy.ec.europa.euAnthropic tells the Senate Banking Committee of a 28.8M-exchange campaign it links to Alibaba's Qwen lab
About 25,000 fraudulent accounts queried Claude between 22 April and 5 June 2026, targeting software engineering, agentic reasoning and long-horizon tasks; Anthropic policy head Sarah Heck calls it the largest known distillation attack on a commercial model. The letter became public on 24 June. Alibaba denies wrongdoing.
Source: cnbc.comAnthropic suspends Fable 5 and Mythos 5 worldwide under a Commerce export-control directive
The directive required restricting access by foreign nationals inside and outside the US; unable to verify nationality in real time, Anthropic pulled both models for all users, beginning a 19-day outage that made frontier model weights an explicit export-control object.
Source: anthropic.comAlibaba shares fall about 7.3% over two sessions after the distillation letter becomes public
Bloomberg's 24 June report on Anthropic's Senate letter is followed by a 2.7% drop that day and 4.7% the next, the first clearly priced market reaction to a distillation accusation since the January 2025 Nvidia sell-off.
Source: dandodiary.comCommerce lifts the export controls on Fable 5 and Mythos 5
Secretary Howard Lutnick says Anthropic no longer needs an export licence after agreeing to proactively detect and address model security risks, work with the government on standards for future models, and report malicious activity. Fable 5 returns globally on 1-2 July; Mythos 5 had been re-approved for US organisations on 26 June.
Source: cnbc.comChinese open-weight models briefly reach 63% of US enterprise tokens on OpenRouter
In the first week of July 2026 Chinese-origin models accounted for about 63% of tokens routed by US firms on OpenRouter, up from under 10% a year earlier, with DeepSeek the single largest vendor at 17.6% and Qwen next at 13.9%. The share is the commercial reason distillation became a policy fight.
Source: finance.yahoo.comHugging Face survey: distillation is now inside nearly every 2026 frontier release
The post catalogues Gemma 4, DeepSeek-V4, MiMo-V2-Flash (multi-teacher on-policy distillation), GLM-5 (checkpoint-as-teacher), Nvidia Nemotron 3 Ultra (10+ domain teachers), Qwen3 and Cursor Composer 2.5 (self-distillation).
Source: huggingface.coLawfare argues the US should respond to distillation without new IP law
Bahrad Sokhansanj writes that distillation uses only public-facing outputs, that copyright, patent and trade-secret law do not clearly reach it, and that the right levers are access security, fraud detection, lab-government information sharing and existing Computer Fraud and Abuse Act provisions. The piece becomes the standard counterweight to the sanctions bills.
Source: lawfaremedia.orgMoonshot releases Kimi K3, a 2.8T-parameter open model
Positioned as a frontier rival to US models and launched via hosted products and API with weights promised for 27 July; within a week the White House alleges it was built by distilling Anthropic's Fable. A Moonshot executive denied on 21 July that K3 was 'a distilled replica of an existing model'.
Source: en.wikipedia.orgWhite House accuses Moonshot of distilling Anthropic's Fable for Kimi K3; Treasury says sanctions are 'on the table'
OSTP director Kratsios says Moonshot built a platform that switched access methods to evade detection and also alleges use of prohibited Nvidia GB300 servers via Thailand; Treasury Secretary Bessent warns of sanctions and Entity List designations. Independent researchers note only ~16 days separate Fable 5's 1 July return and K3's launch.
Source: techcrunch.comMATS researchers report Kimi K3 and GLM 5.2 adopting Claude's persona
Benji Berczi and Kyuhee Kim find Kimi K3 self-identifying as Claude in several of ten trials until a 20 July server-side change, and that telling GLM 5.2 'you are Claude' raises its uncensored answer rate on sensitive PRC questions from 17% to 85%. The authors stress this is evidence that Claude's self-concept sits in the weights, not proof of distillation.
Source: theregister.comChina's Ministry of Commerce publicly rejects the US distillation allegations
MOFCOM calls the accusations evidence-free, 'a typical example of AI hegemony' and a double standard, notes that many US AI companies have distilled Chinese models, cites nearly 200 US startups asking Washington not to cut off access to Chinese open-source models, and warns of 'all necessary measures'. It is Beijing's first comprehensive response.
Source: cset.georgetown.eduNPR: allegations of distillation spark debate over whether it is even illegal
Legal analysts note the disputes rest on terms of service and fraud rather than copyright; US officials estimate unauthorized distillation costs US labs up to $6B a year.
Source: npr.orgDeepSeek V4-Flash official release
The 284B model exits preview as the 0731 build; V4-Pro (1.6T) follows on 13 August.
Source: en.wikipedia.org'Stealing Reasoning Traces from Proprietary LLM APIs' shows encrypted reasoning blocks are portable
Researchers find that encrypted reasoning objects from OpenAI, Anthropic and Google are interchangeable across sessions, users and models, so injecting a strong model's encrypted trace into a weaker sibling makes it emit the reasoning in plaintext. They decrypt 315,320 publicly shared blocks, recovering 367 PII artifacts and 182 credentials; providers mitigated the main attack during August 2026.
Source: arxiv.orgEuropean Commission gains enforcement powers over general-purpose AI models
One year after GPAI obligations took effect, the AI Office can now demand documentation, obtain model access for evaluation, order corrective measures and fine providers up to the greater of EUR 15M or 3% of worldwide turnover. Transparency duties for AI-generated content apply from the same date.
Source: digital-strategy.ec.europa.euAlibaba releases Qwen3.8-Max, a 2.4T-parameter model
Previewed at WAIC on 19 July and shipped with published API pricing on 3 August, roughly two months after Anthropic's Senate letter attributing a 28.8M-exchange campaign to Qwen-linked operators. Qwen3.8-27B open weights follow on 14 August.
Source: en.wikipedia.orgSecurities class action filed against Alibaba in SDNY over undisclosed distillation conduct
The complaint alleges Alibaba described AI risks as future possibilities while already running thousands of accounts to distill Anthropic's Claude. It is the first court filing anywhere arising from the 2026 distillation disputes; no lab has yet sued a rival directly.
Source: dandodiary.comSenators introduce the BLADE Act (S. 5252) to block large-scale adversarial distillation
Hagerty, Scott, Kim and Cortez Masto's bipartisan bill directs the executive branch to publicly list foreign distillation actors, work with industry on detection, and authorises Commerce export controls and Treasury sanctions against them.
Source: hagerty.senate.govByteDance founder Zhang Yiming bans distillation inside the Seed AI team
Zhang tells an internal all-hands that ByteDance will not distil closed or open-weight models, including Kimi K3, even at the cost of falling behind domestic rivals, and Seed issues a policy with API-level detection. It is the first Chinese lab to renounce the practice outright.
Source: technode.comMeta releases Muse Glimmer 30B, distilled from the closed Muse Spark
A 29.6B dense vision-language agentic model with a ViT-G/14 encoder and 128K context, Apache 2.0, quantised to run under 20GB of RAM. It is the clearest 2026 case of a US lab shipping open weights that are an admitted distillation of its own unreleased flagship.
Source: neowin.netAnthropic announces it will watermark text generated by Claude
The commitment follows Anthropic's July 2026 signature on the EU Code of Practice on Transparency of AI-Generated Content and applies at the model level to every Claude surface for models released after 2 August.
Source: techcrunch.comDeepSeek V4-Pro (1.6T parameters) reaches general availability
The 0813 build ends a preview that began on 24 April and completes the V4 rollout across app, web and API. V4's post-training merges RL-trained domain specialists into the unified model by on-policy distillation.
Source: api-docs.deepseek.comAnthropic publishes how Claude's text watermark works
Anthropic uses a SynthID-Text-style scheme that biases the randomness used to choose among equally good word options with a cryptographic key, leaving no hidden characters and no token cost. Anthropic frames it as provenance for EU transparency compliance rather than an anti-distillation control, since it only marks words Claude actually chooses.
Source: anthropic.comAnthropic releases Claude Fable 5.1 and Claude Mythos 5.1
Same weights with different safeguard levels: Fable 5.1 is generally available with cheaper cache reads and fewer false-positive refusals, while Mythos 5.1 stays limited to vetted US organisations under trusted-access programs.
Source: techcrunch.comOpenAI launches GPT-6 Astra at $10/$50 per million tokens
OpenAI calls Astra a generational leap and the first model to trigger its highest internal cyber safeguards; president Greg Brockman frames it as the start of AGI. At 40x the input price of GPT-5.4 mini it resets the gap that the next distilled tier will have to close.
Source: cnbc.com
Glossary · 15 terms
- Knowledge distillation (KD)
- Training a smaller 'student' model to reproduce the behaviour of a larger 'teacher', classically by matching the teacher's softened output probabilities (Hinton et al., 2015).
- Model compression
- The 2006 precursor to KD in which a compact model is trained on unlabeled data pseudo-labeled by a large ensemble (Bucila, Caruana and Niculescu-Mizil).
- Soft targets / dark knowledge
- The full probability distribution a teacher assigns over outputs; its relative probabilities on wrong classes carry information a hard label does not.
- Sequence-level distillation
- Training a student on whole sequences generated by the teacher rather than per-token distributions (Kim and Rush, 2016); the basis of fine-tuning on reasoning traces.
- Black-box distillation
- Distillation using only a teacher's text outputs obtained via an API, with no access to weights or logits (Alpaca, Vicuna, Orca, the DeepSeek accusations).
- Model extraction attack
- The security framing of the same mechanic: querying a deployed model to reconstruct its functionality (Tramer et al., 2016; Krishna et al., 2019). US policy documents in 2026 use 'model extraction' and 'adversarial distillation' interchangeably.
- On-policy distillation
- The student generates its own samples and the teacher scores them token by token (GKD, 2023); reported to match RL at 50-100x lower compute in 2025 and the default post-training stage in 2026.
- Multi-teacher on-policy distillation (MOPD)
- Training several domain-specialist teachers by RL (math, code, agents, safety) and merging them into one student with dense token-level rewards on student rollouts; used by MiMo-V2-Flash, Nemotron 3 Ultra and DeepSeek-V4.
- Reverse KL
- A distillation loss that penalises the student for placing mass where the teacher does not; preferred for generative models (MiniLLM, DeepSeek-V4).
- Self-distillation
- A model teaching a same-size or earlier version of itself: Born-Again Networks (2018), DINO (2021), GLM-5's cross-stage checkpoint distillation and Cursor Composer 2.5's hint-conditioned KL.
- Pruning plus distillation
- Removing layers or widths from a large model and recovering accuracy with distillation (Nvidia Minitron, Llama 3.2 1B/3B).
- Antidistillation sampling
- A decoding-time defence that perturbs the teacher's next-token distribution so its reasoning traces are poor training data for a student while its own answers stay useful (Savani et al., 2025).
- Hydra cluster
- Anthropic's term for networks of thousands of fraudulent accounts, often behind commercial proxies, that spread distillation traffic thinly enough to look like ordinary usage.
- Adversarial distillation / distillation attack
- Unauthorised, large-scale extraction of a proprietary model's outputs through fraudulent accounts, proxies or jailbreaks for the purpose of training a competing model; defined in the US as a national-security concern by NSTM-4 (April 2026).
- Model Extraction Attackers List
- The public list of foreign entities engaged in distillation attacks proposed by H.R. 8283 and the BLADE Act, to be maintained with Commerce and Treasury for sanctions and export controls.
Sources · 111 sources
Every figure on this page comes from one of these primary sources. Compiled 3 September 2026.
- Model Compression (KDD 2006)
- Distilling the Knowledge in a Neural Network
- Sequence-Level Knowledge Distillation
- Stealing Machine Learning Models via Prediction APIs
- Thieves on Sesame Street! Model Extraction of BERT-based APIs
- DistilBERT
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
- Emerging Properties in Self-Supervised Vision Transformers (DINO)
- Self-Instruct: Aligning Language Model with Self Generated Instructions
- Alpaca: A Strong, Replicable Instruction-Following Model
- Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT Quality
- The False Promise of Imitating Proprietary LLMs
- Orca: Progressive Learning from Complex Explanation Traces of GPT-4
- On-Policy Distillation of Language Models (GKD)
- Adversarial Diffusion Distillation (SDXL Turbo)
- Gemini 1.5 Flash announcement (I/O 2024)
- OpenAI DevDay 2024 announcements
- Llama 3.2 announcement
- Amazon Bedrock Model Distillation (preview)
- DeepSeek-V3 Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
- Sky-T1: Train your own O1 preview model within $450
- Nvidia sheds almost $600 billion in market cap
- OpenAI accuses DeepSeek of knowledge distillation
- Distillation Scaling Laws
- NVIDIA Q4 and fiscal 2025 results
- Gemma 3 Technical Report
- OpenAI calls DeepSeek 'state-controlled'
- Antidistillation Sampling
- The Llama 4 herd
- Qwen3 Technical Report
- BIS rescinds AI Diffusion Rule and issues new guidance
- White House unveils America's AI Action Plan
- EU AI Act GPAI obligations in force; final Code of Practice
- Introducing gpt-oss
- Updating restrictions of sales to unsupported regions
- Introducing Claude Haiku 4.5
- On-Policy Distillation
- A new era of intelligence with Gemini 3
- MiMo-V2-Flash Technical Report
- GLM-5: from Vibe Coding to Agentic Engineering
- OpenAI accuses DeepSeek of malpractice ahead of AI launch
- Google says attackers used 100,000+ prompts to try to clone its AI
- MiniMax M2.5 API pricing
- Detecting and preventing distillation attacks
- AI Distillation Attacks: The Case for Targeted Government Intervention
- AI Foundation Model Transparency Act of 2026 (H.R. 8094)
- Gemma release log
- OpenAI, Anthropic, Google team up against Chinese distillation
- Huizenga introduces legislation to stop China and Russia from stealing American AI
- China's Illicit Campaign to Steal and Subvert American AI (witness statement)
- China has 'deliberate, industrial-scale campaigns' to steal AI models
- From Diagnosis to Deterrence: The Emerging U.S. Response to Distillation
- Chairmen Garbarino, Moolenaar announce joint investigation into PRC AI models
- AI Distillation Attacks: Executive and Congressional Action Can Go Further
- Introducing Composer 2.5
- Adversarial Distillation
- NVIDIA Nemotron 3 Ultra Technical Report
- Strong backing for the Code of Practice on Transparency of AI-generated Content
- Anthropic accuses Alibaba of campaign to 'brazenly' and 'illicitly' distill Claude
- Redeploying Claude Fable 5
- Anthropic says Trump admin has lifted export controls on Claude Fable 5 and Mythos 5
- China AI models capture 63% of U.S. OpenRouter usage
- Distillation in 2026 (so far): which frontier models use it and how
- Responding to AI Distillation Without Panic
- Treasury threatens sanctions after White House claims Moonshot distilled Anthropic's Fable
- Impostor Chinese models pretend they're Claude
- MOFCOM press spokesperson answers questions on model distillation (translation)
- How to Fix the AI Model Theft Bill Before It Becomes Law
- Allegations of AI distillation spark debate about IP theft. But is it illegal?
- Commission starts enforcing AI Act rules and new transparency requirements
- Securities suit against Alibaba combines two key litigation trends
- Hagerty, colleagues introduce the BLADE Act
- Zhang Yiming says ByteDance's Seed team won't rely on AI distillation
- Meta releases Muse Glimmer, a 30B open agentic AI model
- Anthropic says it will watermark text generated by its AI models
- DeepSeek-V4-Pro GA release
- How Claude's text watermarking works
- Stealing Reasoning Traces from Proprietary LLM APIs
- Anthropic's new Fable release is cheaper, less restrictive
- OpenAI announces rollout of GPT-6 Astra model
- FitNets: Hints for Thin Deep Nets
- Born-Again Neural Networks
- TinyBERT: Distilling BERT for Natural Language Understanding
- Distilling Step-by-Step!
- MiniLLM: Knowledge Distillation of Large Language Models
- Gemma (language model)
- Introducing the next generation of Claude
- GPT-4o mini pricing and analysis
- Compact Language Models via Pruning and Knowledge Distillation (Minitron)
- Llama-3.1-Minitron-4B-Width-Base model card
- OpenAI o1
- Claude (language model)
- DeepSeek
- Phi-4 Technical Report
- BIS rescinds AI Diffusion Rule
- Microsoft probing if DeepSeek-linked group improperly obtained OpenAI data
- Open Thoughts launch
- Gemini (language model)
- s1: Simple test-time scaling
- LIMO: Less is More for Reasoning
- Phi-4-reasoning Technical Report
- Amazon Bedrock Model Distillation is now generally available
- OpenThoughts: Data Recipes for Reasoning Models
- GPT-5 nano pricing
- Apple Intelligence
- Qwen
- GPT-5.4
- MiniMax (company)
- Moonshot AI
- Anthropic newsroom