Timeline · compiled 3 September 2026 · 111 sources

Cross-cutting timeline of AI distillation, 2006 to 2026

Knowledge distillation began as an academic model-compression trick (Bucila 2006, Hinton 2015) and spent a decade as a research topic before becoming the default way to build small language models (DistilBERT 2019, Gemma, Llama 3.2, Qwen3). In 2024 it turned into a cloud product line (OpenAI Model Distillation, Amazon Bedrock Model Distillation) and a pricing strategy (Gemini Flash, GPT-4o mini, Haiku). DeepSeek-R1's January 2025 release, with six openly distilled students, triggered a market shock and the first public accusation that a rival had distilled a frontier lab's API outputs. By 2026 distillation was a national-security issue: Anthropic, OpenAI and Google all disclosed industrial-scale extraction campaigns within eleven days of each other in February, the White House issued National Security and Technology Memorandum 4 on adversarial distillation, and Congress advanced the Deterring American AI Model Theft Act and the BLADE Act while Beijing publicly rejected the whole framing. This timeline records 130 dated events across research, product, market, policy and legal categories, each with a source URL.

Key figures · 6 figures

Dated events in this timeline

130 events

57 in 2026 so far

Every event carries its own source URL; spans 2006-08-20 to 2026-09-03.

dl.acm.org

Research milestones

38 events

29% of the record

Papers and public technical reports, from Model Compression (2006) to reasoning-trace extraction (2026).

arxiv.org

Product and market events

60 events

52 product, 8 market

Model launches, distillation-as-a-service products, pricing tiers and priced market reactions.

huggingface.co

Policy and legal events

32 events

32 of them since January 2025

Bills, memoranda, export-control actions, lab disclosures and disputes. Almost none predate the DeepSeek-R1 release.

iaps.ai

Events in the last 12 months

65 events

50% of the whole timeline

From 3 September 2025 to 3 September 2026 — the period covering NSTM-4, the BLADE Act and the three labs’ extraction disclosures.

anthropic.com

Span from first to last event

20 years

2006-08-20 → 2026-09-03

First: Bucila, Caruana and Niculescu-Mizil publish 'Model Compression' at KDD 2006. Last: OpenAI launches GPT-6 Astra at $10/$50 per million tokens.

cnbc.com

Key findings · 8 findings

  1. Era 1 (2006-2018): distillation as model compression

    Bucila, Caruana and Niculescu-Mizil showed in 2006 that a small network trained on the pseudo-labels of a large ensemble could match it; Hinton, Vinyals and Dean formalised soft-target 'dark knowledge' in 2015 and FitNets, sequence-level KD (Kim and Rush) and Born-Again Networks extended it to deeper students, seq2seq models and self-distillation. The technique was framed as deployment engineering, not competition. In parallel, the security literature was already describing the same mechanic as an attack: Tramer et al. stole model functionality through prediction APIs in 2016.

    Sources dl.acm.org · arxiv.org · arxiv.org · usenix.org

  2. Era 2 (2019-2022): the transformer compression wave

    DistilBERT and TinyBERT made distillation the standard recipe for shipping BERT-class models to production, cutting parameters by 40-90% while keeping most accuracy. Hugging Face's DistilBERT became one of the most downloaded models on the Hub and normalised distilled checkpoints as first-class open artifacts. The same window produced the first NLP model-extraction attack papers (Krishna et al., 2019) and self-distillation for vision (DINO, 2021).

    Sources arxiv.org · arxiv.org · arxiv.org · arxiv.org

  3. Era 3 (2023-2024): black-box distillation from frontier APIs

    Stanford Alpaca (52K text-davinci-003 demonstrations for under $600), Vicuna (ShareGPT conversations), Orca (GPT-4 explanation traces) and Distilling Step-by-Step showed that a small open model could inherit behaviour from a closed API using only outputs, not logits. MiniLLM and GKD reframed the loss for generative students (reverse KL, on-policy sampling), while Gudibande et al. warned that imitation closes style gaps faster than capability gaps. Frontier labs responded by hiding chain-of-thought (o1, September 2024) and by selling distillation themselves (OpenAI Model Distillation, October 2024; Bedrock, December 2024).

    Sources crfm.stanford.edu · arxiv.org · arxiv.org · infoq.com

  4. Era 4 (2025): reasoning distillation and the DeepSeek shock

    DeepSeek-R1 (20 January 2025) shipped six distilled dense students from 1.5B to 70B and proved that reasoning traces transfer cheaply; Sky-T1 ($450), s1 (1,000 examples), LIMO (817 examples) and OpenThoughts pushed the data floor lower. Nvidia lost nearly $600B of market value in one session, and within 48 hours OpenAI and the White House alleged DeepSeek had distilled OpenAI outputs. Gemma 3, Llama 4 and Qwen3 all disclosed teacher-student training as core recipe, and CMU's Antidistillation Sampling opened a defensive research line.

    Sources arxiv.org · cnbc.com · eweek.com · arxiv.org

  5. Era 5 (2026): adversarial distillation becomes statecraft

    Three US labs published extraction evidence within eleven days: OpenAI's memo to the House Select Committee (12 February), Google's Threat Intelligence Group report on a 100,000-prompt Gemini campaign (12 February) and Anthropic's disclosure of 16M+ Claude exchanges from ~24,000 fraudulent accounts (23 February). Anthropic later told the Senate Banking Committee about a 28.8M-exchange campaign it attributes to Alibaba's Qwen lab. OSTP issued NSTM-4 (23 April), the House Foreign Affairs Committee advanced H.R. 8283 43-0, the White House accused Moonshot of distilling Anthropic's Fable for Kimi K3 (22 July), and senators introduced the BLADE Act (5 August). Beijing's Ministry of Commerce rejected the whole framing on 27 July.

    Sources anthropic.com · nbcnews.com · iaps.ai · cset.georgetown.edu

  6. By 2026 distillation is inside almost every frontier training run, hostile or not

    Hugging Face's mid-2026 survey found teacher-student training in Gemma 4, DeepSeek-V4 (per-domain RL specialists merged by on-policy distillation), Nvidia Nemotron 3 Ultra (10+ domain teachers), Xiaomi MiMo-V2-Flash, GLM-5 (earlier checkpoints as teachers), Qwen3 and Cursor Composer 2.5 (self-distillation). Meta's Muse Glimmer 30B was distilled from the closed Muse Spark, and Apple said its 2026 foundation models are refined with Gemini outputs under a paid partnership. The technique that governments want to sanction is the same one that ships nearly every model on the market.

    Sources huggingface.co · research.nvidia.com · cursor.com · en.wikipedia.org

  7. The legal status of distillation is still unsettled

    As of September 2026 no court has ruled on whether API-output distillation is unlawful, and neither Anthropic nor OpenAI has sued any of the accused labs; the disputes rest on terms-of-service breaches, fraud (fake accounts, proxies) and export-control policy rather than copyright or patent. Lawfare's July 2026 analysis argues existing law already covers the fraudulent-access part and warns that broad new IP rights would entrench incumbents. Bills such as H.R. 8283 and S. 5252 explicitly preserve 'legitimate' distillation and rely on public attacker lists, sanctions and export controls instead of new IP rights. The one filed case so far is a securities class action against Alibaba, not an IP suit.

    Sources itif.org · npr.org · lawfaremedia.org · dandodiary.com

  8. Defences moved from hiding outputs to watermarking and shared threat intelligence

    The first defence was concealment: OpenAI hid o1's chain of thought in September 2024. Research defences followed (Antidistillation Sampling, April 2025) and by 2026 the response was operational: Anthropic's classifiers and behavioural fingerprinting, Google's real-time prompt-cluster detection, and an April 2026 arrangement in which OpenAI, Anthropic and Google exchange distillation attack signatures through the Frontier Model Forum. Anthropic's August 2026 text watermark is explicitly a provenance measure for EU transparency compliance rather than an anti-distillation control, and an August 2026 paper showed encrypted reasoning blocks could be replayed across models to recover hidden traces.

    Sources arxiv.org · anthropic.com · techbrew.com · arxiv.org

Charts · 3 charts

Distillation events per year, stacked by category

events
The values plotted in “Distillation events per year, stacked by category”, in events.
YearResearch eventsProduct eventsMarket eventsPolicy eventsLegal events
200610000
201410000
201510000
201620000
201810000
201930000
202010000
202110000
202210000
202380000
2024312000
20251116272
20264246158

Shows the shift from a purely research topic (2006-2022) to a product story (2024) and then a policy and legal story (2025-2026). 2026 is partial: it covers 1 January to 3 September 2026.

Sources: huggingface.co · cnas.org

Cumulative distillation events, 2006 to September 2026

events
The values plotted in “Cumulative distillation events, 2006 to September 2026”, in events.
YearCumulative events events
20061
20142
20153
20165
20186
20199
202010
202111
202212
202320
202435
202573
2026130

The curve is close to flat for the first fifteen years and then near-vertical from 2024, when distillation simultaneously became a cloud product, a pricing tier and a national-security file.

Sources: arxiv.org · anthropic.com

Share of timeline events by category

events
The values plotted in “Share of timeline events by category”, in events.
CategoryEvents
Research38
Product52
Market8
Policy22
Legal10

Product and research events dominate the record, but policy and legal events are concentrated almost entirely in the 20 months from January 2025.

Sources: iaps.ai

Tables · 1 table

Dated distillation events by year and category, 2006-2026

13 rows
Dated distillation events by year and category, 2006-2026 — Every event in this timeline counted by calendar year and category. Years with no recorded event are omitted. Counts are derived from the timeline array below, so each cell traces to individual sourced entries. — Units: Research in events; Product in events; Market in events; Policy in events; Legal in events; Total in events.
YearResearch eventsProduct eventsMarket eventsPolicy eventsLegal eventsTotal events
2006100001 dl.acm.org
2014100001 arxiv.org
2015100001 arxiv.org
2016200002 arxiv.org
2018100001 arxiv.org
2019300003 arxiv.org
2020100001 arxiv.org
2021100001 arxiv.org
2022100001 arxiv.org
2023800008 crfm.stanford.edu
202431200015 en.wikipedia.org
2025111627238 novasky-ai.github.io
2026424615857 en.wikipedia.org

Derived from this file's timeline array; every underlying event carries its own source URL. Coverage is denser for 2025-2026 because distillation only became a policy and market story then, so early years understate research activity.

Sources: arxiv.org · huggingface.co · cnas.org

Timeline · 130 events

  1. research

    Bucila, Caruana and Niculescu-Mizil publish 'Model Compression' at KDD 2006

    The paper shows that a single small neural network trained on unlabeled data pseudo-labeled by a large ensemble can match the ensemble's accuracy at a fraction of the size. It is the earliest widely cited ancestor of modern knowledge distillation.

    Source: dl.acm.org
  2. research

    FitNets introduces hint-based distillation into intermediate layers

    Romero et al. train thin, deep students by matching the teacher's intermediate representations ('hints') rather than only its outputs, establishing feature-based distillation.

    Source: arxiv.org
  3. research

    Hinton, Vinyals and Dean post 'Distilling the Knowledge in a Neural Network'

    The paper coins 'distillation', introduces temperature-scaled soft targets and shows ensemble knowledge can be compressed into one model on MNIST and speech. It becomes the canonical citation for the field.

    Source: arxiv.org
  4. research

    Kim and Rush propose sequence-level knowledge distillation

    Sequence-level KD trains a student on the teacher's beam-search outputs instead of per-token distributions, enabling small neural machine translation models. This is the direct ancestor of training on generated reasoning traces.

    Source: arxiv.org
  5. research

    'Stealing Machine Learning Models via Prediction APIs' presented at USENIX Security

    Tramer, Zhang, Juels, Reiter and Ristenpart show that an adversary with only query access to a pay-per-prediction ML API can reconstruct a near-equivalent model, including against commercial services. It is the founding paper of model extraction, the adversarial framing of what labs would later call distillation.

    Source: usenix.org
  6. research

    Born-Again Neural Networks show self-distillation improves students of equal size

    Furlanello et al. find that a student distilled from an identically sized teacher can outperform it, complicating the compression-only view of distillation.

    Source: arxiv.org
  7. research

    TinyBERT distills BERT with transformer-layer and attention matching

    Huawei researchers combine general and task-specific distillation to build a 4-layer BERT that retains most of BERT-base accuracy at 7.5x fewer parameters.

    Source: arxiv.org
  8. research

    Hugging Face releases DistilBERT

    DistilBERT is 40% smaller and 60% faster than BERT-base while retaining about 97% of its language-understanding performance. It makes distilled checkpoints a mainstream open-source artifact.

    Source: arxiv.org
  9. research

    'Thieves on Sesame Street!' extracts BERT-based APIs for a few hundred dollars

    Krishna, Tomar, Parikh, Papernot and Iyyer show that random word sequences plus task heuristics are enough to extract a working copy of a fine-tuned BERT API, and that membership classification and API watermarking both fail against sophisticated adversaries. The paper is the direct NLP ancestor of the 2026 distillation-attack disclosures.

    Source: arxiv.org
  10. research

    MobileBERT: task-agnostic distillation for on-device transformers

    Sun et al. distil a specially designed inverted-bottleneck BERT-large teacher into a 25M-parameter student that is 4.3x smaller and 5.5x faster than BERT-base, establishing distillation as the standard route to edge NLP.

    Source: arxiv.org
  11. research

    DINO frames self-supervised vision learning as self-distillation with no labels

    Caron et al. (Meta AI) train a student to match a momentum teacher's softmax outputs on different crops of the same image, producing ViT features that segment objects without labels. Self-distillation becomes a general representation-learning tool, not just a compression trick.

    Source: arxiv.org
  12. research

    Self-Instruct bootstraps instruction data from a model's own generations

    Wang et al. generate instructions, inputs and outputs from GPT-3 itself and fine-tune on the filtered result, closing most of the gap to InstructGPT. The pipeline is what Alpaca reuses three months later against a proprietary teacher.

    Source: arxiv.org
  13. research

    Stanford Alpaca fine-tunes LLaMA 7B on 52K text-davinci-003 demonstrations

    Data generation cost under $500 via the OpenAI API and fine-tuning under $100. Alpaca popularises black-box distillation from a proprietary API and prompts OpenAI's terms-of-service clause against training competing models to receive wide attention.

    Source: crfm.stanford.edu
  14. research

    Vicuna-13B fine-tunes LLaMA on ~70K ShareGPT ChatGPT conversations

    LMSYS reports about 90% of ChatGPT quality using user-shared conversations scraped from ShareGPT, for roughly $300 of training compute. Vicuna makes conversation-log distillation the default way to build an open chat model in 2023.

    Source: lmsys.org
  15. research

    Google's 'Distilling Step-by-Step' uses LLM rationales as extra supervision

    Hsieh et al. train small task-specific models on chain-of-thought rationales extracted from a large model, outperforming standard fine-tuning with less data.

    Source: arxiv.org
  16. research

    'The False Promise of Imitating Proprietary LLMs' pushes back on API distillation

    Gudibande, Wallace, Snell and co-authors (Berkeley) show that imitation models learn a teacher's style faster than its capability: human raters are fooled, benchmarks are not. The paper frames the capability gap as closable only with vastly more imitation data or a better base model, and is still the standard caution cited against pure black-box distillation.

    Source: arxiv.org
  17. research

    Microsoft's Orca learns from GPT-4 explanation traces

    A 13B model trained on step-by-step explanation traces from GPT-4 approaches ChatGPT on several benchmarks, showing that imitation of reasoning, not just answers, transfers capability.

    Source: arxiv.org
  18. research

    MiniLLM proposes reverse-KL distillation for generative language models

    Gu et al. argue forward KL is ill-suited to open-ended generation and train students with reverse KL on their own samples, improving over standard sequence-level KD.

    Source: arxiv.org
  19. research

    Generalized Knowledge Distillation (GKD) introduces on-policy distillation

    Agarwal et al. (Google DeepMind) train the student on its own generated sequences with teacher feedback, fixing train-inference distribution mismatch. On-policy distillation becomes the dominant recipe by 2025-2026.

    Source: arxiv.org
  20. research

    Adversarial Diffusion Distillation ships SDXL Turbo: 1-step image generation

    Sauer et al. (Stability AI) combine an adversarial loss with score distillation to sample a large diffusion model in 1-4 steps, matching SDXL in four. Distillation moves from language into real-time generative media, a parallel track that later produces LCM, Flux and video-model distillations.

    Source: arxiv.org
  21. product

    Google releases Gemma; the 2B model is distilled from 7B

    Gemma's first generation establishes teacher-student training inside Google's open-model line, a pattern that continues through Gemma 2, 3 and 4.

    Source: en.wikipedia.org
  22. product

    Anthropic announces the Claude 3 family including Claude 3 Haiku

    Haiku is positioned as the fastest, cheapest tier of a three-tier lineup, cementing the industry pattern of a small model priced far below its flagship sibling.

    Source: anthropic.com
  23. product

    Google says Gemini 1.5 Flash was trained by distillation from 1.5 Pro

    At I/O 2024 Google explicitly describes Flash as distilled from Pro, the first time a frontier lab named distillation as the recipe for its flagship low-cost tier. Launch pricing was $0.35 per million input tokens.

    Source: blog.google
  24. product

    Gemma 2: 9B distilled from 27B, 2B from an unreleased 7B teacher

    Google scales distillation into Gemma 2, training smaller models on soft targets from larger ones instead of next-token prediction alone.

    Source: en.wikipedia.org
  25. product

    OpenAI launches GPT-4o mini at $0.15/$0.60 per million tokens

    GPT-4o mini undercuts Claude 3 Haiku ($0.25/$1.25) and Gemini 1.5 Flash ($0.35/$0.70) and becomes the default student model for OpenAI's later distillation API.

    Source: simonwillison.net
  26. research

    Nvidia's Minitron: compact models via pruning plus distillation

    Muralidharan et al. prune Nemotron-4 15B and retrain with distillation using up to 40x fewer tokens per model, reporting 1.8x compute savings for a model family.

    Source: arxiv.org
  27. research

    Llama-3.1-Minitron 4B: 'LLM Pruning and Distillation in Practice'

    Nvidia applies the Minitron recipe to Llama 3.1 8B and Mistral NeMo 12B, retraining pruned models with 94B tokens of distillation. The models are released on Hugging Face for commercial use.

    Source: huggingface.co
  28. product

    OpenAI releases o1-preview with a hidden chain of thought

    OpenAI forbids attempts to reveal o1's reasoning and cites competitive advantage alongside safety. Hiding traces is widely read as the first commercial anti-distillation defence.

    Source: en.wikipedia.org
  29. product

    Meta's Llama 3.2 1B and 3B built by pruning 8B and distilling from 8B/70B logits

    Meta states it used single-shot structured pruning from Llama 3.1 8B and incorporated logits from the 8B and 70B models as token-level targets during pre-training.

    Source: ai.meta.com
  30. product

    OpenAI launches Model Distillation in the API at DevDay

    Stored Completions, Evals and fine-tuning are combined into a workflow for distilling GPT-4o or o1-preview outputs into GPT-4o mini. OpenAI offered 2M free daily training tokens on GPT-4o mini through 31 October 2024.

    Source: infoq.com
  31. product

    Anthropic releases Claude 3.5 Haiku

    The small tier is refreshed alongside the upgraded Claude 3.5 Sonnet and the computer-use beta.

    Source: en.wikipedia.org
  32. product

    DeepSeek previews R1-Lite, its first reasoning model

    DeepSeek-R1-Lite-Preview shows visible chain-of-thought two months before the full R1 and its distilled family.

    Source: en.wikipedia.org
  33. product

    Amazon Bedrock Model Distillation enters preview

    AWS automates teacher generation, data synthesis and student fine-tuning within the same model family (Anthropic, Meta Llama, Amazon Nova), claiming up to 5x faster and 75% cheaper models with under 2% accuracy loss on RAG tasks.

    Source: aws.amazon.com
  34. research

    Microsoft Phi-4 technical report: a 14B model that surpasses its teacher on STEM QA

    Phi-4 leans on synthetic data generated with GPT-4-class models, and Microsoft reports it exceeds the teacher on STEM question answering, a milestone for 'weak-to-strong' distillation.

    Source: arxiv.org
  35. product

    DeepSeek-V3 released: 671B MoE trained in 2.788M H800 GPU-hours

    The technical report (arXiv 27 December) discloses the full training budget, roughly $5.6M at assumed rental prices, which seeds the January 2025 market narrative about cheap frontier models.

    Source: arxiv.org
  36. research

    Sky-T1-32B-Preview: an o1-class reasoner trained for $450

    UC Berkeley's NovaSky team fine-tunes Qwen2.5-32B on 17K traces from QwQ-32B-Preview, reaching 43.3% on AIME 2024 and 82.4% on MATH500.

    Source: novasky-ai.github.io
  37. policy

    BIS publishes the Framework for Artificial Intelligence Diffusion

    The Biden-era interim final rule creates a three-tier country system for advanced chips and, for the first time, controls exports of closed frontier model weights. Compliance was set for 15 May 2025.

    Source: wiley.law
  38. research

    DeepSeek-R1 released with six openly distilled students (1.5B to 70B)

    Alongside R1 and R1-Zero, DeepSeek publishes dense models distilled from R1 onto Qwen 2.5 and Llama 3 bases under an MIT licence. The paper reports the distilled students beat RL-from-scratch at the same size.

    Source: arxiv.org
  39. market

    DeepSeek shock: Nvidia loses nearly $600B in market value in one day

    Nvidia falls about 17% and the Nasdaq about 3% as investors reprice AI capex on the belief that cheap distilled and open models reduce demand for frontier compute.

    Source: cnbc.com
  40. research

    OpenThoughts project launches with a 114K-example open reasoning dataset

    The Bespoke Labs/Stanford/Berkeley collaboration scales up Bespoke-Stratos and releases OpenThinker-7B, aiming to open-source the reasoning-distillation data recipe.

    Source: open-thoughts.ai
  41. product

    Gemini 2.0 Flash reaches general availability

    Google's second-generation Flash tier becomes the default Gemini model, continuing the distilled-flagship pricing strategy introduced with 1.5 Flash.

    Source: en.wikipedia.org
  42. research

    s1: Simple test-time scaling distills 1,000 Gemini Thinking traces

    Muennighoff et al. fine-tune Qwen2.5-32B on the s1K set (traces from Gemini 2.0 Flash Thinking) and add 'budget forcing', exceeding o1-preview on competition math.

    Source: arxiv.org
  43. research

    LIMO: 817 curated examples yield 63.3% on AIME24

    Ye et al. argue reasoning is elicited rather than taught, reaching 95.6% on MATH500 with about 1% of the data used by prior distillation recipes.

    Source: arxiv.org
  44. research

    Apple publishes 'Distillation Scaling Laws'

    Busbridge et al. fit a scaling law for student performance as a function of teacher and student compute, giving compute-optimal distillation recipes for when a teacher exists and when it must be trained.

    Source: arxiv.org
  45. market

    Nvidia reports record $39.3B quarter one month after the DeepSeek sell-off

    Q4 FY2025 revenue rose 78% year on year and data-center revenue hit $35.6B; CEO Jensen Huang argues reasoning models add a new scaling law for compute.

    Source: nvidianews.nvidia.com
  46. research

    Gemma 3: all sizes (1B to 27B) trained with knowledge distillation

    Google DeepMind's technical report states every Gemma 3 model is distilled; Gemma3-4B-IT is reported competitive with Gemma2-27B-IT.

    Source: arxiv.org
  47. policy

    OpenAI's AI Action Plan submission calls DeepSeek 'state-controlled' and urges bans

    In its OSTP response OpenAI recommends barring 'PRC-produced' models in Tier 1 countries, citing IP-theft risk after its distillation allegations.

    Source: techcrunch.com
  48. product

    Meta's Llama 4 Maverick is codistilled from the 2T-parameter Behemoth

    Meta describes a novel distillation loss that dynamically weights soft and hard targets, run during Behemoth's own pre-training to amortise cost. Behemoth itself was never released.

    Source: ai.meta.com
  49. product

    Gemini 2.5 Flash released

    Google's first Flash model with configurable thinking budgets ships two months after 2.5 Pro.

    Source: en.wikipedia.org
  50. research

    Antidistillation Sampling: the first decoding-time defence against reasoning distillation

    Savani, Trockman, Feng and co-authors (CMU) perturb the teacher's next-token distribution so that its reasoning traces poison a downstream student while the teacher's own utility is preserved. The paper (NeurIPS 2025) opens a defensive research line that continues through ADS-C and 'The Distillation Game' in 2026.

    Source: arxiv.org
  51. product

    Qwen3 launches with 'strong-to-weak' distillation for models from 0.6B to 14B

    Alibaba distills its 235B-A22B and 32B flagships into smaller students using off-policy then on-policy distillation, reporting better results than RL at roughly one-tenth the GPU hours.

    Source: arxiv.org
  52. research

    Phi-4-reasoning distills o3-mini reasoning traces into a 14B model

    Microsoft fine-tunes Phi-4 on curated 'teachable' prompts with demonstrations generated by OpenAI's o3-mini.

    Source: arxiv.org
  53. product

    Amazon Bedrock Model Distillation reaches general availability

    GA adds Nova Premier/Pro, Claude 3.5 Sonnet v2 and Llama 3.3/3.2 pairings plus function-calling data augmentation for agent use cases.

    Source: aws-news.com
  54. policy

    Commerce rescinds the AI Diffusion Rule two days before its compliance date

    BIS says the tiered framework would have stifled US innovation and burdened allies; it issues separate guidance warning that using US chips to train Chinese models risks enforcement.

    Source: akingump.com
  55. product

    Anthropic releases Claude Opus 4 and Sonnet 4

    The Claude 4 generation resets the tiering that later distilled-tier releases (Haiku 4.5) are benchmarked against.

    Source: en.wikipedia.org
  56. research

    OpenThoughts3: 1.2M-example distillation dataset from QwQ-32B

    OpenThinker3-7B scores 53% on AIME 2025 and 54% on GPQA Diamond, beating DeepSeek-R1-Distill-Qwen-7B by 15-20 points via data-recipe ablations alone.

    Source: arxiv.org
  57. policy

    EU AI Office publishes the final General-Purpose AI Code of Practice

    Three chapters (transparency, copyright, safety and security) give GPAI providers a voluntary path to compliance ahead of the 2 August obligations.

    Source: lw.com
  58. policy

    White House releases America's AI Action Plan

    Ninety-plus actions across innovation, infrastructure and international security; its IP-protection language is later cited as the basis for anti-distillation legislation.

    Source: whitehouse.gov
  59. policy

    EU AI Act obligations for general-purpose AI models begin to apply

    New GPAI models placed on the EU market must meet transparency and copyright duties; Commission enforcement powers start 2 August 2026 and legacy models have until 2027.

    Source: lw.com
  60. product

    OpenAI ships gpt-oss-120b and gpt-oss-20b, its first open weights since GPT-2

    Both models are Apache 2.0 and run on a single GPU or a 16GB laptop. Releasing open weights removes the API terms-of-service barrier for distillation from an OpenAI-lineage model, and gpt-oss quickly becomes a teacher and a base for third-party students.

    Source: openai.com
  61. product

    Claude Opus 4.1 released

    An incremental flagship update that precedes Anthropic's autumn 4.5 lineup.

    Source: en.wikipedia.org
  62. product

    OpenAI launches the GPT-5 family: GPT-5, GPT-5 mini ($0.25/$2) and GPT-5 nano ($0.05/$0.40)

    Mini and nano tiers are priced 5x and 25x below GPT-5, extending the three-tier flagship/distilled pattern to OpenAI's fifth generation.

    Source: openrouter.ai
  63. product

    DeepSeek V3.1 adds hybrid thinking and non-thinking modes

    DeepSeek merges its chat and reasoning lines into one model, mirroring Qwen3's design.

    Source: en.wikipedia.org
  64. policy

    Anthropic bars companies majority-owned by Chinese entities, citing distillation risk

    The policy blocks any entity more than 50% owned by companies headquartered in unsupported regions such as China, explicitly naming distillation as a way adversaries could advance their own models. Press coverage followed on 5 September.

    Source: anthropic.com
  65. product

    Claude Sonnet 4.5 released

    Anthropic's mid-tier model becomes the reference point for Haiku 4.5's cost-performance claims two weeks later.

    Source: en.wikipedia.org
  66. product

    Claude Haiku 4.5: Sonnet 4-level coding at one-third the cost

    Priced at $1/$5 per million tokens, Haiku 4.5 is marketed as matching Sonnet 4 on coding at more than twice the speed, the clearest small-tier value claim of 2025.

    Source: anthropic.com
  67. research

    Thinking Machines: on-policy distillation reaches RL parity at 50-100x less compute

    Kevin Lu and colleagues report 7-10x fewer gradient steps than RL and cite Qwen's 1,800 vs 17,920 GPU-hour comparison on AIME'24, attributing the gain to dense per-token reward.

    Source: thinkingmachines.ai
  68. product

    Google launches Gemini 3 Pro and Deep Think

    Gemini 3 Pro debuts at 1501 Elo on LMArena and 91.9% GPQA Diamond; it becomes the teacher for the 3 Flash tier released a month later.

    Source: blog.google
  69. product

    Claude Opus 4.5 released

    Anthropic closes 2025 with its fourth flagship refresh of the year.

    Source: en.wikipedia.org
  70. product

    DeepSeek V3.2 released with DeepSeek Sparse Attention

    The last V3-line model before V4; it follows the V3.2-Exp preview of 29 September.

    Source: en.wikipedia.org
  71. product

    Gemini 3 Flash released

    Google ships the small tier of its third generation one month after 3 Pro.

    Source: en.wikipedia.org
  72. market

    Apple and Google announce a multi-year Gemini partnership for Apple Foundation Models

    Bloomberg had reported a roughly $1B-per-year deal for a custom 1.2T-parameter Gemini; Apple later says its own models are refined using outputs from Gemini frontier models, a sanctioned, paid form of distillation.

    Source: en.wikipedia.org
  73. research

    Xiaomi's MiMo-V2-Flash introduces Multi-Teacher On-Policy Distillation (MOPD)

    The 309B-total / 15B-active MoE is post-trained by training domain-specialist RL teachers (math, code, agentic, safety) and merging them into one student with dense token-level rewards plus outcome reward models. MOPD is the template that Nvidia's Nemotron 3 and DeepSeek-V4 follow later in 2026.

    Source: arxiv.org
  74. product

    Claude Opus 4.6 released

    First of Anthropic's 2026 releases; MiniMax markets M2.5 the same month as roughly 1/20th the cost of Opus 4.6.

    Source: en.wikipedia.org
  75. product

    Zhipu/Z.ai releases GLM-5, a 744B open-weight model with cross-stage self-distillation

    GLM-5's final post-training stage is on-policy distillation from its own earlier SFT and RL checkpoints, used to recover capabilities that later RL stages had degraded. It is the clearest production example of a model distilling from past versions of itself.

    Source: arxiv.org
  76. market

    MiniMax ships M2.5 and M2.5 Lightning at roughly 1/20th of Claude Opus 4.6 pricing

    M2.5 is listed at $0.15/$1.20 per million tokens (Lightning $0.30/$2.40) with 80.2% on SWE-bench Verified. It is the sharpest 2026 example of a Chinese open-weight model undercutting a US frontier tier on price while the same lab is named in Anthropic's distillation disclosure eleven days later.

    Source: openrouter.ai
  77. product

    Alibaba releases Qwen3.5, a 397B-parameter Apache 2.0 MoE with 1M context

    The generation immediately preceding the April-June campaign that Anthropic later attributes to Qwen-linked operators.

    Source: en.wikipedia.org
  78. product

    Claude Sonnet 4.6 released

    Mid-tier refresh that lands six days before Anthropic's distillation-attack disclosure.

    Source: en.wikipedia.org
  79. policy

    IAPS publishes 'AI Distillation Attacks: The Case for Targeted Government Intervention'

    The policy memo argues for narrowly scoped intervention that punishes fraudulent-access distillation without chilling legitimate teacher-student training, and becomes the analytic reference for the April congressional and executive actions.

    Source: iaps.ai
  80. product

    OpenAI releases GPT-5.4; mini and nano follow on 17 March

    GPT-5.4 mini goes to free-tier users while nano stays API-only. Both small tiers are priced about four times above their GPT-5 equivalents, the first time the distilled tier moved up rather than down in price.

    Source: en.wikipedia.org
  81. policy

    AI Foundation Model Transparency Act of 2026 (H.R. 8094) introduced

    Rep. Don Beyer's bill would require disclosure of training data and methods for foundation models, indirectly touching teacher-student provenance.

    Source: congress.gov
  82. product

    Google releases Gemma 4 (E2B, E4B, 26B A4B, 31B), again trained with distillation

    Google's release log dates the first Gemma 4 checkpoints to 31 March 2026, with the public announcement on 2 April; multi-token-prediction variants follow on 16 April and a 12B unified model on 3 June. Hugging Face's mid-2026 survey describes post-training as distillation from a large instruction-tuned teacher, extending the Gemma 2/3 recipe.

    Source: ai.google.dev
  83. market

    OpenAI, Anthropic and Google begin sharing distillation attack signatures via the Frontier Model Forum

    Bloomberg reports the three labs exchange query-distribution patterns, prompt structures, IP fingerprints and account-creation behaviour the way cyber-threat intelligence is shared, so a signature detected at one lab can be flagged at the others within hours. Antitrust uncertainty limits how much can be exchanged.

    Source: techbrew.com
  84. policy

    Deterring American AI Model Theft Act (H.R. 8283) introduced by Rep. Huizenga

    The bill creates a public 'AI Model Extraction Attackers List', discretionary sanctions and a Commerce/State information-sharing channel, while stating that legitimate distillation remains a valuable research tool.

    Source: huizenga.house.gov
  85. policy

    House Select Committee hearing: 'China's Illicit Campaign to Steal and Subvert American AI'

    Witness testimony describes coordinated distillation campaigns using proxy accounts and jailbreaks against US frontier models.

    Source: docs.house.gov
  86. policy

    House Foreign Affairs Committee marks up H.R. 8283 and reports it favorably, 43-0

    A technical amendment swaps State Department references for Commerce. The unanimous vote one week after introduction signals bipartisan consensus on treating adversarial distillation as a sanctionable act. IAPS records the favorable report on 23 April.

    Source: iaps.ai
  87. policy

    OSTP issues National Security and Technology Memorandum 4 on adversarial distillation

    Director Michael Kratsios says foreign entities, principally in China, run 'deliberate, industrial-scale campaigns' to distill US frontier AI systems using tens of thousands of proxies; agencies are told to share intelligence, co-develop defences and explore accountability measures.

    Source: defenseone.com
  88. policy

    State Department cables posts and issues a demarche to Beijing over distillation

    US diplomats are instructed to warn host governments about model extraction by DeepSeek, Moonshot AI and MiniMax, and a formal protest is delivered to Beijing.

    Source: justsecurity.org
  89. product

    DeepSeek previews V4-Flash (284B) and V4-Pro (1.6T), trained with on-policy distillation from domain experts

    Per Hugging Face's survey, each domain (math, code, agents) gets an RL-trained specialist teacher and the unified model is trained with reverse-KL against them.

    Source: huggingface.co
  90. policy

    House Homeland Security and Select Committee on China open a joint investigation into PRC AI models

    Chairmen Garbarino and Moolenaar send letters to Anysphere (Cursor) and Airbnb, citing 'unauthorized model distillation and other illicit techniques' by DeepSeek, Alibaba, Moonshot AI and MiniMax and questioning Cursor Composer 2's use of a Moonshot open-weight base. The probe later expands to DoorDash.

    Source: homeland.house.gov
  91. product

    Cursor ships Composer 2.5, trained by self-distillation with textual feedback

    The same model acts as teacher and student: hint-conditioned outputs supervise unhinted ones through per-token KL, and on-policy distillation repairs localized failures (bad tool calls, premature stops) without rewriting whole rollouts. Cursor reports 25x more synthetic tasks than Composer 2 and near-frontier coding scores at lower token price.

    Source: cursor.com
  92. product

    Claude Opus 4.8 released

    Last Opus 4.x release before the Claude 5 generation; falls inside the window of the alleged Qwen-linked extraction campaign (22 April to 5 June 2026).

    Source: en.wikipedia.org
  93. product

    Apple says WWDC 2026 foundation models are refined using outputs from Gemini frontier models

    Craig Federighi describes Apple's third-generation models as trained on proprietary data and refined with Gemini outputs under the January partnership, an openly licensed instance of the same technique at the centre of the enforcement fights.

    Source: en.wikipedia.org
  94. product

    MiniMax launches M3 'frontier coding' model

    Four months after Anthropic's report, MiniMax ships a 1M-context multimodal coding model; MiniMax did not publicly rebut the distillation findings.

    Source: en.wikipedia.org
  95. policy

    CNAS publishes 'Adversarial Distillation' report

    Remler and Hayum estimate the 16M Claude exchanges could represent 150-400 billion extracted tokens (against DeepSeek-R1's entire 6.4B-token SFT set) and map a six-actor supply chain including token mixers and transfer stations such as CloseAI, BianXie AI, One-API, New-API, OpenRouter and Eden AI.

    Source: cnas.org
  96. product

    Nvidia releases Nemotron 3 Ultra, a 550B open model built by multi-teacher on-policy distillation

    More than ten domain-specialised teachers were trained separately and consolidated into the 550B-total / 55B-active hybrid Mamba-MoE student through dense token-level guidance on student-generated rollouts. The technical report is dated 9 June 2026.

    Source: research.nvidia.com
  97. product

    Anthropic releases Claude Fable 5 and Claude Mythos 5

    Fable 5 is the model the White House later says Moonshot distilled for Kimi K3; Mythos 5 is limited-access.

    Source: en.wikipedia.org
  98. policy

    European Commission publishes the Code of Practice on Transparency of AI-Generated Content

    The voluntary code operationalises AI Act Article 50 marking and labelling duties; around 190 organisations, including Anthropic, had signed by 31 July 2026. It is the direct cause of Claude's text watermark two months later.

    Source: digital-strategy.ec.europa.eu
  99. policy

    Anthropic suspends Fable 5 and Mythos 5 worldwide under a Commerce export-control directive

    The directive required restricting access by foreign nationals inside and outside the US; unable to verify nationality in real time, Anthropic pulled both models for all users, beginning a 19-day outage that made frontier model weights an explicit export-control object.

    Source: anthropic.com
  100. market

    Alibaba shares fall about 7.3% over two sessions after the distillation letter becomes public

    Bloomberg's 24 June report on Anthropic's Senate letter is followed by a 2.7% drop that day and 4.7% the next, the first clearly priced market reaction to a distillation accusation since the January 2025 Nvidia sell-off.

    Source: dandodiary.com
  101. policy

    Commerce lifts the export controls on Fable 5 and Mythos 5

    Secretary Howard Lutnick says Anthropic no longer needs an export licence after agreeing to proactively detect and address model security risks, work with the government on standards for future models, and report malicious activity. Fable 5 returns globally on 1-2 July; Mythos 5 had been re-approved for US organisations on 26 June.

    Source: cnbc.com
  102. product

    Claude Sonnet 5 released

    Anthropic's mid-tier joins the Claude 5 generation.

    Source: en.wikipedia.org
  103. market

    Chinese open-weight models briefly reach 63% of US enterprise tokens on OpenRouter

    In the first week of July 2026 Chinese-origin models accounted for about 63% of tokens routed by US firms on OpenRouter, up from under 10% a year earlier, with DeepSeek the single largest vendor at 17.6% and Qwen next at 13.9%. The share is the commercial reason distillation became a policy fight.

    Source: finance.yahoo.com
  104. research

    Hugging Face survey: distillation is now inside nearly every 2026 frontier release

    The post catalogues Gemma 4, DeepSeek-V4, MiMo-V2-Flash (multi-teacher on-policy distillation), GLM-5 (checkpoint-as-teacher), Nvidia Nemotron 3 Ultra (10+ domain teachers), Qwen3 and Cursor Composer 2.5 (self-distillation).

    Source: huggingface.co
  105. product

    Moonshot releases Kimi K3, a 2.8T-parameter open model

    Positioned as a frontier rival to US models and launched via hosted products and API with weights promised for 27 July; within a week the White House alleges it was built by distilling Anthropic's Fable. A Moonshot executive denied on 21 July that K3 was 'a distilled replica of an existing model'.

    Source: en.wikipedia.org
  106. product

    Claude Opus 5 released

    Anthropic's flagship of the Claude 5 generation.

    Source: anthropic.com
  107. research

    MATS researchers report Kimi K3 and GLM 5.2 adopting Claude's persona

    Benji Berczi and Kyuhee Kim find Kimi K3 self-identifying as Claude in several of ten trials until a 20 July server-side change, and that telling GLM 5.2 'you are Claude' raises its uncensored answer rate on sensitive PRC questions from 17% to 85%. The authors stress this is evidence that Claude's self-concept sits in the weights, not proof of distillation.

    Source: theregister.com
  108. policy

    China's Ministry of Commerce publicly rejects the US distillation allegations

    MOFCOM calls the accusations evidence-free, 'a typical example of AI hegemony' and a double standard, notes that many US AI companies have distilled Chinese models, cites nearly 200 US startups asking Washington not to cut off access to Chinese open-source models, and warns of 'all necessary measures'. It is Beijing's first comprehensive response.

    Source: cset.georgetown.edu
  109. product

    DeepSeek V4-Flash official release

    The 284B model exits preview as the 0731 build; V4-Pro (1.6T) follows on 13 August.

    Source: en.wikipedia.org
  110. research

    'Stealing Reasoning Traces from Proprietary LLM APIs' shows encrypted reasoning blocks are portable

    Researchers find that encrypted reasoning objects from OpenAI, Anthropic and Google are interchangeable across sessions, users and models, so injecting a strong model's encrypted trace into a weaker sibling makes it emit the reasoning in plaintext. They decrypt 315,320 publicly shared blocks, recovering 367 PII artifacts and 182 credentials; providers mitigated the main attack during August 2026.

    Source: arxiv.org
  111. policy

    European Commission gains enforcement powers over general-purpose AI models

    One year after GPAI obligations took effect, the AI Office can now demand documentation, obtain model access for evaluation, order corrective measures and fine providers up to the greater of EUR 15M or 3% of worldwide turnover. Transparency duties for AI-generated content apply from the same date.

    Source: digital-strategy.ec.europa.eu
  112. product

    Alibaba releases Qwen3.8-Max, a 2.4T-parameter model

    Previewed at WAIC on 19 July and shipped with published API pricing on 3 August, roughly two months after Anthropic's Senate letter attributing a 28.8M-exchange campaign to Qwen-linked operators. Qwen3.8-27B open weights follow on 14 August.

    Source: en.wikipedia.org
  113. policy

    Senators introduce the BLADE Act (S. 5252) to block large-scale adversarial distillation

    Hagerty, Scott, Kim and Cortez Masto's bipartisan bill directs the executive branch to publicly list foreign distillation actors, work with industry on detection, and authorises Commerce export controls and Treasury sanctions against them.

    Source: hagerty.senate.gov
  114. market

    ByteDance founder Zhang Yiming bans distillation inside the Seed AI team

    Zhang tells an internal all-hands that ByteDance will not distil closed or open-weight models, including Kimi K3, even at the cost of falling behind domestic rivals, and Seed issues a policy with API-level detection. It is the first Chinese lab to renounce the practice outright.

    Source: technode.com
  115. product

    Meta releases Muse Glimmer 30B, distilled from the closed Muse Spark

    A 29.6B dense vision-language agentic model with a ViT-G/14 encoder and 128K context, Apache 2.0, quantised to run under 20GB of RAM. It is the clearest 2026 case of a US lab shipping open weights that are an admitted distillation of its own unreleased flagship.

    Source: neowin.net
  116. product

    Anthropic announces it will watermark text generated by Claude

    The commitment follows Anthropic's July 2026 signature on the EU Code of Practice on Transparency of AI-Generated Content and applies at the model level to every Claude surface for models released after 2 August.

    Source: techcrunch.com
  117. product

    DeepSeek V4-Pro (1.6T parameters) reaches general availability

    The 0813 build ends a preview that began on 24 April and completes the V4 rollout across app, web and API. V4's post-training merges RL-trained domain specialists into the unified model by on-policy distillation.

    Source: api-docs.deepseek.com
  118. product

    Anthropic publishes how Claude's text watermark works

    Anthropic uses a SynthID-Text-style scheme that biases the randomness used to choose among equally good word options with a cryptographic key, leaving no hidden characters and no token cost. Anthropic frames it as provenance for EU transparency compliance rather than an anti-distillation control, since it only marks words Claude actually chooses.

    Source: anthropic.com
  119. product

    Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1

    Same weights with different safeguard levels: Fable 5.1 is generally available with cheaper cache reads and fewer false-positive refusals, while Mythos 5.1 stays limited to vetted US organisations under trusted-access programs.

    Source: techcrunch.com
  120. product

    OpenAI launches GPT-6 Astra at $10/$50 per million tokens

    OpenAI calls Astra a generational leap and the first model to trigger its highest internal cyber safeguards; president Greg Brockman frames it as the start of AGI. At 40x the input price of GPT-5.4 mini it resets the gap that the next distilled tier will have to close.

    Source: cnbc.com

Glossary · 15 terms

Knowledge distillation (KD)
Training a smaller 'student' model to reproduce the behaviour of a larger 'teacher', classically by matching the teacher's softened output probabilities (Hinton et al., 2015).
Model compression
The 2006 precursor to KD in which a compact model is trained on unlabeled data pseudo-labeled by a large ensemble (Bucila, Caruana and Niculescu-Mizil).
Soft targets / dark knowledge
The full probability distribution a teacher assigns over outputs; its relative probabilities on wrong classes carry information a hard label does not.
Sequence-level distillation
Training a student on whole sequences generated by the teacher rather than per-token distributions (Kim and Rush, 2016); the basis of fine-tuning on reasoning traces.
Black-box distillation
Distillation using only a teacher's text outputs obtained via an API, with no access to weights or logits (Alpaca, Vicuna, Orca, the DeepSeek accusations).
Model extraction attack
The security framing of the same mechanic: querying a deployed model to reconstruct its functionality (Tramer et al., 2016; Krishna et al., 2019). US policy documents in 2026 use 'model extraction' and 'adversarial distillation' interchangeably.
On-policy distillation
The student generates its own samples and the teacher scores them token by token (GKD, 2023); reported to match RL at 50-100x lower compute in 2025 and the default post-training stage in 2026.
Multi-teacher on-policy distillation (MOPD)
Training several domain-specialist teachers by RL (math, code, agents, safety) and merging them into one student with dense token-level rewards on student rollouts; used by MiMo-V2-Flash, Nemotron 3 Ultra and DeepSeek-V4.
Reverse KL
A distillation loss that penalises the student for placing mass where the teacher does not; preferred for generative models (MiniLLM, DeepSeek-V4).
Self-distillation
A model teaching a same-size or earlier version of itself: Born-Again Networks (2018), DINO (2021), GLM-5's cross-stage checkpoint distillation and Cursor Composer 2.5's hint-conditioned KL.
Pruning plus distillation
Removing layers or widths from a large model and recovering accuracy with distillation (Nvidia Minitron, Llama 3.2 1B/3B).
Antidistillation sampling
A decoding-time defence that perturbs the teacher's next-token distribution so its reasoning traces are poor training data for a student while its own answers stay useful (Savani et al., 2025).
Hydra cluster
Anthropic's term for networks of thousands of fraudulent accounts, often behind commercial proxies, that spread distillation traffic thinly enough to look like ordinary usage.
Adversarial distillation / distillation attack
Unauthorised, large-scale extraction of a proprietary model's outputs through fraudulent accounts, proxies or jailbreaks for the purpose of training a competing model; defined in the US as a national-security concern by NSTM-4 (April 2026).
Model Extraction Attackers List
The public list of foreign entities engaged in distillation attacks proposed by H.R. 8283 and the BLADE Act, to be maintained with Commerce and Treasury for sanctions and export controls.

Sources · 111 sources

Every figure on this page comes from one of these primary sources. Compiled 3 September 2026.

  1. Model Compression (KDD 2006)ACM · 20 August 2006 · paper
  2. Distilling the Knowledge in a Neural NetworkarXiv · 9 March 2015 · paper
  3. Sequence-Level Knowledge DistillationarXiv · 25 June 2016 · paper
  4. Stealing Machine Learning Models via Prediction APIsUSENIX Security · 11 August 2016 · paper
  5. Thieves on Sesame Street! Model Extraction of BERT-based APIsarXiv · 27 October 2019 · paper
  6. DistilBERTarXiv · 2 October 2019 · paper
  7. MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesarXiv · 6 April 2020 · paper
  8. Emerging Properties in Self-Supervised Vision Transformers (DINO)arXiv (Meta AI) · 29 April 2021 · paper
  9. Self-Instruct: Aligning Language Model with Self Generated InstructionsarXiv · 20 December 2022 · paper
  10. Alpaca: A Strong, Replicable Instruction-Following ModelStanford CRFM · 13 March 2023 · blog
  11. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT QualityLMSYS · 30 March 2023 · blog
  12. The False Promise of Imitating Proprietary LLMsarXiv (UC Berkeley) · 25 May 2023 · paper
  13. Orca: Progressive Learning from Complex Explanation Traces of GPT-4arXiv · 5 June 2023 · paper
  14. On-Policy Distillation of Language Models (GKD)arXiv · 23 June 2023 · paper
  15. Adversarial Diffusion Distillation (SDXL Turbo)arXiv (Stability AI) · 28 November 2023 · paper
  16. Gemini 1.5 Flash announcement (I/O 2024)Google · 14 May 2024 · blog
  17. OpenAI DevDay 2024 announcementsInfoQ · 1 October 2024 · news
  18. Llama 3.2 announcementMeta · 25 September 2024 · blog
  19. Amazon Bedrock Model Distillation (preview)AWS · 3 December 2024 · blog
  20. DeepSeek-V3 Technical ReportarXiv · 27 December 2024 · paper
  21. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RLarXiv · 22 January 2025 · paper
  22. Sky-T1: Train your own O1 preview model within $450NovaSky (UC Berkeley) · 10 January 2025 · blog
  23. Nvidia sheds almost $600 billion in market capCNBC · 27 January 2025 · news
  24. OpenAI accuses DeepSeek of knowledge distillationeWeek · 30 January 2025 · news
  25. Distillation Scaling LawsarXiv (Apple) · 12 February 2025 · paper
  26. NVIDIA Q4 and fiscal 2025 resultsNVIDIA · 26 February 2025 · filing
  27. Gemma 3 Technical ReportarXiv (Google DeepMind) · 25 March 2025 · paper
  28. OpenAI calls DeepSeek 'state-controlled'TechCrunch · 13 March 2025 · news
  29. Antidistillation SamplingarXiv (CMU) · 18 April 2025 · paper
  30. The Llama 4 herdMeta · 5 April 2025 · blog
  31. Qwen3 Technical ReportarXiv (Alibaba) · 14 May 2025 · paper
  32. BIS rescinds AI Diffusion Rule and issues new guidanceAkin Gump · 13 May 2025 · law
  33. White House unveils America's AI Action PlanWhite House · 23 July 2025 · law
  34. EU AI Act GPAI obligations in force; final Code of PracticeLatham & Watkins · 2 August 2025 · law
  35. Introducing gpt-ossOpenAI · 5 August 2025 · blog
  36. Updating restrictions of sales to unsupported regionsAnthropic · 4 September 2025 · blog
  37. Introducing Claude Haiku 4.5Anthropic · 15 October 2025 · blog
  38. On-Policy DistillationThinking Machines Lab · 27 October 2025 · blog
  39. A new era of intelligence with Gemini 3Google · 18 November 2025 · blog
  40. MiMo-V2-Flash Technical ReportarXiv (Xiaomi LLM-Core) · 8 January 2026 · paper
  41. GLM-5: from Vibe Coding to Agentic EngineeringarXiv (Z.ai / Zhipu) · 11 February 2026 · paper
  42. OpenAI accuses DeepSeek of malpractice ahead of AI launchRest of World · 12 February 2026 · news
  43. Google says attackers used 100,000+ prompts to try to clone its AINBC News · 12 February 2026 · news
  44. MiniMax M2.5 API pricingOpenRouter · 12 February 2026 · pricing
  45. Detecting and preventing distillation attacksAnthropic · 23 February 2026 · blog
  46. AI Distillation Attacks: The Case for Targeted Government InterventionIAPS · March 2026 · paper
  47. AI Foundation Model Transparency Act of 2026 (H.R. 8094)US Congress · 26 March 2026 · law
  48. Gemma release logGoogle · 3 June 2026 · docs
  49. OpenAI, Anthropic, Google team up against Chinese distillationTech Brew · 7 April 2026 · news
  50. Huizenga introduces legislation to stop China and Russia from stealing American AIUS House of Representatives · 20 April 2026 · law
  51. China's Illicit Campaign to Steal and Subvert American AI (witness statement)US House Select Committee on the CCP · 16 April 2026 · filing
  52. China has 'deliberate, industrial-scale campaigns' to steal AI modelsDefense One · 23 April 2026 · news
  53. From Diagnosis to Deterrence: The Emerging U.S. Response to DistillationJust Security · 5 May 2026 · blog
  54. Chairmen Garbarino, Moolenaar announce joint investigation into PRC AI modelsUS House Committee on Homeland Security · 29 April 2026 · law
  55. AI Distillation Attacks: Executive and Congressional Action Can Go FurtherIAPS · 12 May 2026 · blog
  56. Introducing Composer 2.5Cursor · 18 May 2026 · blog
  57. Adversarial DistillationCNAS · 2 June 2026 · paper
  58. NVIDIA Nemotron 3 Ultra Technical ReportNVIDIA · 9 June 2026 · paper
  59. Strong backing for the Code of Practice on Transparency of AI-generated ContentEuropean Commission · 10 June 2026 · law
  60. Anthropic accuses Alibaba of campaign to 'brazenly' and 'illicitly' distill ClaudeCNBC · 24 June 2026 · news
  61. Redeploying Claude Fable 5Anthropic · 30 June 2026 · blog
  62. Anthropic says Trump admin has lifted export controls on Claude Fable 5 and Mythos 5CNBC · 30 June 2026 · news
  63. China AI models capture 63% of U.S. OpenRouter usageGuruFocus via Yahoo Finance · July 2026 · news
  64. Distillation in 2026 (so far): which frontier models use it and howHugging Face · 8 July 2026 · blog
  65. Responding to AI Distillation Without PanicLawfare · 16 July 2026 · blog
  66. Treasury threatens sanctions after White House claims Moonshot distilled Anthropic's FableTechCrunch · 22 July 2026 · news
  67. Impostor Chinese models pretend they're ClaudeThe Register · 27 July 2026 · news
  68. MOFCOM press spokesperson answers questions on model distillation (translation)CSET Georgetown · 27 July 2026 · law
  69. How to Fix the AI Model Theft Bill Before It Becomes LawITIF · 28 July 2026 · blog
  70. Allegations of AI distillation spark debate about IP theft. But is it illegal?NPR · 28 July 2026 · news
  71. Commission starts enforcing AI Act rules and new transparency requirementsEuropean Commission · 2 August 2026 · law
  72. Securities suit against Alibaba combines two key litigation trendsThe D&O Diary · 4 August 2026 · filing
  73. Hagerty, colleagues introduce the BLADE ActUS Senate · 5 August 2026 · law
  74. Zhang Yiming says ByteDance's Seed team won't rely on AI distillationTechNode · 6 August 2026 · news
  75. Meta releases Muse Glimmer, a 30B open agentic AI modelNeowin · 10 August 2026 · news
  76. Anthropic says it will watermark text generated by its AI modelsTechCrunch · 11 August 2026 · news
  77. DeepSeek-V4-Pro GA releaseDeepSeek · 13 August 2026 · docs
  78. How Claude's text watermarking worksAnthropic · 14 August 2026 · blog
  79. Stealing Reasoning Traces from Proprietary LLM APIsarXiv · August 2026 · paper
  80. Anthropic's new Fable release is cheaper, less restrictiveTechCrunch · 1 September 2026 · news
  81. OpenAI announces rollout of GPT-6 Astra modelCNBC · 3 September 2026 · news
  82. FitNets: Hints for Thin Deep NetsarXiv · 19 December 2014 · paper
  83. Born-Again Neural NetworksarXiv · 12 May 2018 · paper
  84. TinyBERT: Distilling BERT for Natural Language UnderstandingarXiv (Huawei Noah’s Ark Lab) · 23 September 2019 · paper
  85. Distilling Step-by-Step!arXiv (Google) · 3 May 2023 · paper
  86. MiniLLM: Knowledge Distillation of Large Language ModelsarXiv · 14 June 2023 · paper
  87. Gemma (language model)Wikipedia · August 2026 · news
  88. Introducing the next generation of ClaudeAnthropic · 4 March 2024 · blog
  89. GPT-4o mini pricing and analysisSimon Willison · 18 July 2024 · blog
  90. Compact Language Models via Pruning and Knowledge Distillation (Minitron)arXiv (NVIDIA) · 19 July 2024 · paper
  91. Llama-3.1-Minitron-4B-Width-Base model cardNVIDIA via Hugging Face · 21 August 2024 · docs
  92. OpenAI o1Wikipedia · August 2026 · news
  93. Claude (language model)Wikipedia · 1 September 2026 · news
  94. DeepSeekWikipedia · August 2026 · news
  95. Phi-4 Technical ReportarXiv (Microsoft) · 12 December 2024 · paper
  96. BIS rescinds AI Diffusion RuleWiley Rein · 13 May 2025 · law
  97. Microsoft probing if DeepSeek-linked group improperly obtained OpenAI dataUS News / Reuters · 28 January 2025 · news
  98. Open Thoughts launchOpen Thoughts · 28 January 2025 · blog
  99. Gemini (language model)Wikipedia · August 2026 · news
  100. s1: Simple test-time scalingarXiv (Stanford) · 31 January 2025 · paper
  101. LIMO: Less is More for ReasoningarXiv · 5 February 2025 · paper
  102. Phi-4-reasoning Technical ReportarXiv (Microsoft) · 30 April 2025 · paper
  103. Amazon Bedrock Model Distillation is now generally availableAWS News · 1 May 2025 · blog
  104. OpenThoughts: Data Recipes for Reasoning ModelsarXiv · 4 June 2025 · paper
  105. GPT-5 nano pricingOpenRouter · 7 August 2025 · pricing
  106. Apple IntelligenceWikipedia · August 2026 · news
  107. QwenWikipedia · August 2026 · news
  108. GPT-5.4Wikipedia · August 2026 · news
  109. MiniMax (company)Wikipedia · August 2026 · news
  110. Moonshot AIWikipedia · August 2026 · news
  111. Anthropic newsroomAnthropic · 1 September 2026 · blog