================================================================================================ GLOBAL DISTILLATION — FULL TEXT DIGEST ================================================================================================ Site: https://global-distillation.com/ Licence: MIT (see https://global-distillation.com/LICENSE). Free to quote, excerpt, cite and redistribute with attribution. Cite as: Global Distillation, "How frontier intelligence is compressed, priced and contested", https://global-distillation.com/, accessed . Dataset updated: 2026-09-04 (live signals refresh daily at 06:17 UTC) Digest built: from data/*.json as of 2026-09-04, by scripts/gen-llms.mjs Machine data: https://global-distillation.com/data/ — one JSON file per section, schema at https://global-distillation.com/data/SCHEMA.md Q&A pairs: https://global-distillation.com/data/answers.json This is the whole compendium in one plain-text file, for language models and answer engines. It contains every perspective summary, every key finding, every headline figure with its unit and primary-source URL, every comparison table rendered as pipe-separated text with units in the header, every chart series as data points, the complete dated timeline, one paragraph on each distillation method, the merged glossary and the full source register. Nothing here is generated prose about the subject: it is a rendering of the published dataset. Where a primary source never published a figure, the cell reads "undisclosed" or "n/a" rather than an estimate. CONTENTS 1. What distillation is, and how this compendium is sourced 2. Live signals (refreshed daily) 3. Academic perspective — https://global-distillation.com/academic 4. Financial perspective — https://global-distillation.com/financial 5. Political perspective — https://global-distillation.com/political 6. Company perspective — https://global-distillation.com/company 7. Developer perspective — https://global-distillation.com/developer 8. Customer perspective — https://global-distillation.com/customer 9. Method library — https://global-distillation.com/library 10. Timeline 2006-2026 — https://global-distillation.com/timeline 11. Glossary (merged across all files) Sources are listed at the end of each section, in the order the section cites them, and per-row and per-chart source URLs appear inline in the tables and chart blocks themselves. ================================================================================================ 1. WHAT DISTILLATION IS, AND HOW THIS COMPENDIUM IS SOURCED ================================================================================================ Knowledge distillation (KD): Training a student model to reproduce the behaviour of a teacher model, using the teacher's outputs, internal states or generated data as the supervision signal instead of (or alongside) ground-truth labels. Teacher: The model whose behaviour is being copied. Usually larger, slower or more expensive than the student; occasionally the same size (self-distillation) or even a peer (mutual learning). Student: The model being trained. Its capacity ceiling, not the loss function, is usually what limits how much of the teacher survives the transfer. Sourcing: every numeric claim in this dataset carries a source URL, and the register at the end of each section lists the principal works. Sources are primary wherever one exists — papers, model cards, official pricing pages, regulatory filings, statutes and named public statements. The six perspective files and the two reference files are edited by hand and each carries its own "updated" date; data/live.json is machine-generated daily. Accusations between companies are reported as accusations, with the accuser, the claim, the evidence offered and the outcome each recorded separately. Page URLs: Overview: https://global-distillation.com/ Academic: https://global-distillation.com/academic Financial: https://global-distillation.com/financial Political: https://global-distillation.com/political Company: https://global-distillation.com/company Developer: https://global-distillation.com/developer Customer: https://global-distillation.com/customer Method library: https://global-distillation.com/library Timeline: https://global-distillation.com/timeline Compare: https://global-distillation.com/compare Methodology: https://global-distillation.com/methodology ================================================================================================ 2. LIVE SIGNALS (REFRESHED DAILY) ================================================================================================ Automatically refreshed daily at 06:17 UTC by scripts/update-data.mjs. Last refresh: 2026-09-04T22:27:24.062Z. Endpoint: https://global-distillation.com/data/live.json. - arXiv papers matching "knowledge distillation" (all time): 5,372 papers - arXiv papers matching "knowledge distillation" (last 30 days): 78 papers Year | arXiv papers (papers) ---- | --------------------- 2015 | 2 2016 | 6 2017 | 20 2018 | 54 2019 | 163 2020 | 342 2021 | 467 2022 | 641 2023 | 807 2024 | 1016 2025 | 1122 2026 | 731 - Hugging Face models matching "distill": 16,329 models Model | Downloads, 30 days (downloads) | Likes (likes) ----- | ------------------------ | ------------- deepseek-ai/DeepSeek-R1-Distill-Qwen-32B | 562818 | 1608 deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | 367771 | 881 deepseek-ai/DeepSeek-R1-Distill-Llama-8B | 379197 | 875 deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B | 461220 | 1571 distilbert/distilbert-base-uncased | 7067963 | 1127 Qwen/Qwen3-8B | 13232997 | 1343 google/gemma-3-4b-it | 1626709 | 1470 meta-llama/Llama-3.2-3B-Instruct | 1419885 | 2512 microsoft/Phi-4-mini-instruct | 456948 | 829 HuggingFaceTB/SmolLM3-3B | 581340 | 1022 Repository | Stars (stars) | Description ---------- | ------------- | ----------- huggingface/trl | 19223 | n/a arcee-ai/DistillKit | 1052 | n/a pytorch/torchtune | 5804 | n/a NVIDIA/NeMo | 18388 | n/a axolotl-ai-cloud/axolotl | 12440 | n/a unslothai/unsloth | 75624 | n/a hiyouga/LLaMA-Factory | 74577 | n/a vllm-project/vllm | 90979 | n/a sgl-project/sglang | 35476 | n/a huggingface/open-r1 | 26452 | n/a open-thoughts/open-thoughts | 2331 | n/a EleutherAI/lm-evaluation-harness | 13890 | n/a SafeAILab/EAGLE | 2524 | n/a NovaSky-AI/SkyThought | 3399 | n/a simplescaling/s1 | 6667 | n/a Recent news items tracked by the daily refresh: - 2026-07-30 — Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it — https://www.ctgt.ai/research/distillation-censorship-transfer - 2026-07-25 — The tech world is suddenly obsessed with one concept in AI: Distillation — https://www.cnbc.com/2026/07/25/hat-is-distillation-and-why-is-everyone-so-obsessed-with-it-this-week.html - 2026-07-22 — Top White House official escalating the fight over Moonshot AI's Kimi K3 model — https://www.businessinsider.com/white-house-kimi-k3-moonshot-ai-distillation-2026-7 - 2026-07-19 — Updating IP Regulations for AI Distillation — https://www.marble.onl/posts/copyright_vs_ai.html - 2026-07-16 — Responding to AI Distillation Without Panic — https://www.lawfaremedia.org/article/responding-to-ai-distillation-without-panic - 2026-07-13 — A brief history of distillation in AI — https://twitter.com/SergioPaniego/status/2073066275819991472 - 2026-07-06 — American A.I. Companies Say Chinese Copycats Are Quickly Catching Up — https://www.nytimes.com/2026/07/06/technology/ai-distillation-china.html - 2026-06-26 — Anthropic Accuses Alibaba of Largest AI Distillation Attack: 28.8M Fraudulent — https://yipzap.com/anthropic-accuses-alibaba-of-largest-ai-distillation-attack-28-8m-fraudulent-exchanges/ - 2026-05-25 — 96% Correct Next Token Prediction, with No DNN, No Training, Autodistilled Model — https://mltechniques.com/2026/05/25/96-correct-next-token-prediction-with-no-dnn-no-training-auto-distilled-model/ - 2026-04-30 — I over-engineered my simple AI backend: distillation, router, embedding etc. — https://sisyphusconsulting.org/case-studies/2026/04/01/scaling-llms-at-the-edge/ - 2026-04-29 — US accuses China of industrial-scale AI model distillation, will share Intel — https://thenextweb.com/news/us-white-house-ai-model-distillation-china-theft - 2026-04-28 — When model distillation becomes a diplomatic incident — https://underlines.news/2026/04/26/us-orders-global-diplomatic-warning-on-chinese-distillation-of-ai-models - 2026-04-25 — White House Memo on Adversarial Distillation of American AI Models [pdf] — https://whitehouse.gov/wp-content/uploads/2026/04/NSTM-4.pdf - 2026-03-08 — When Distillation Strips the Soul: Safety Comparison of a Claude-Distilled Model — https://netrork.com/blog/when-distillation-strips-the-soul/ - 2026-03-03 — Show HN: Aside – Local meeting capture with vault-native AI distillation — https://github.com/jshph/aside/ - 2026-02-14 — AI could eat itself: Competitors (..) steal their secrets and clone them — https://www.theregister.com/2026/02/14/ai_risk_distillation_attacks/ - 2026-02-12 — Google identifies over 100k prompts used in distillation attacks — https://cloud.google.com/blog/topics/threat-intelligence/distillation-experimentation-integration-ai-adversarial-use - 2026-01-07 — Show HN: Feedstack – SerpAPI/OpenAI extract feature requests from Reddit/forums — https://feedstack.app - 2025-11-22 — Show HN: Reverse Jailbreaking a Psychopathic AI via Identity Injection — https://github.com/DRawson5570/AI-Wisdom-Distillation - 2025-07-20 — Distillation makes AI models smaller and cheaper — https://www.quantamagazine.org/how-distillation-makes-ai-models-smaller-and-cheaper-20250718/ ================================================================================================ 3. ACADEMIC PERSPECTIVE — ACADEMIC VIEW OF AI DISTILLATION: FROM DARK KNOWLEDGE TO ON-POLICY REASONING TRANSFER ================================================================================================ Page: https://global-distillation.com/academic Data: https://global-distillation.com/data/academic.json Updated: 2026-09-04 Content: 8 key figures, 6 tables, 6 charts, 37 dated events, 22 glossary terms, 48 sources SUMMARY -------- Knowledge distillation began as a model-compression trick — Buciluă, Caruana and Niculescu-Mizil compressed an ensemble into a single net in 2006, and Hinton, Vinyals and Dean gave it its modern soft-target formulation in 2015 (now ~25.9k citations). For a decade the literature organised itself along three axes — what is transferred (response/logit, feature, relation), how the teacher is available (white-box logits vs black-box text), and when the student samples (offline, online, self) — with BERT-era students such as DistilBERT, TinyBERT and MobileBERT retaining 96.8-99.2% of teacher quality at 1.7-7.5x smaller size (DistilBERT is 40% smaller; TinyBERT 7.5x, MobileBERT 4.3x). The January 2025 release of the DeepSeek-R1-Distill series turned distillation from a compression tool into the default way to transfer reasoning: an 800k-trace SFT run on Qwen2.5-32B reached AIME 2024 pass@1 of 72.6 versus 47.0 for large-scale RL applied to the same base model, and triggered a wave of ultra-cheap replications (Sky-T1 at under $450, s1 at 1,000 samples and ~7 H100-hours, LIMO at 817 samples, Bespoke-Stratos at 17k). Since late 2025 the frontier has shifted from off-policy trace imitation to on-policy distillation, where the student samples and the teacher grades every token: arXiv submissions mentioning on-policy distillation jumped from 38 in 2025 to 359 in the first eight months of 2026, and Qwen3, DeepSeek-V4 and Nemotron 3 Ultra all use it as a primary post-training stage — Qwen3 reports matching-or-better results at roughly one tenth of RL GPU-hours. The open research questions are now theoretical rather than engineering: distillation scaling laws, the capacity gap between teacher and student, whether students can ever exceed their teachers, and the homogenisation/model-collapse risk of a literature increasingly trained on its own outputs. KEY FIGURES ----------- - Citations: Hinton, Vinyals & Dean (2015): 25,899 citations (the field's founding text) Semantic Scholar citation count for arXiv:1503.02531, retrieved 2026-09-04 Source: https://api.semanticscholar.org/graph/v1/paper/arXiv:1503.02531?fields=title,year,citationCount - arXiv papers with 'knowledge distillation' in abstract (2025): 1,084 papers (+11.6% vs 2024 (971)) arXiv API count, submittedDate 2025-01-01 to 2025-12-31 Source: https://export.arxiv.org/api/query?search_query=abs:%22knowledge%20distillation%22%20AND%20submittedDate:%5B202501010000%20TO%20202512312359%5D&max_results=1 - Same query, 2026 year-to-date (through 2026-09-04): 709 papers (on pace for ~1,050 full-year) arXiv API count, submittedDate 2026-01-01 to 2026-09-04 Source: https://export.arxiv.org/api/query?search_query=abs:%22knowledge%20distillation%22%20AND%20submittedDate:%5B202601010000%20TO%20202609042359%5D&max_results=1 - arXiv papers mentioning 'on-policy distillation' (2026 YTD): 359 papers (9.4x the 38 seen in all of 2025) arXiv API full-text field count; the single fastest-growing sub-topic in the distillation literature Source: https://export.arxiv.org/api/query?search_query=all:%22on-policy%20distillation%22%20AND%20submittedDate:%5B202601010000%20TO%20202609042359%5D&max_results=1 - Citations: DeepSeek-R1 (Nature, 2025): 5,597 citations (first peer-reviewed open-weight frontier LLM) Semantic Scholar count for the Nature version of arXiv:2501.12948, retrieved 2026-09-04 Source: https://api.semanticscholar.org/graph/v1/paper/arXiv:2501.12948?fields=title,year,citationCount - Cheapest published reasoning distillation run (s1-32B): 7 H100 GPU-hours (1,000 training samples) 26 minutes on 16 NVIDIA H100s, fine-tuning Qwen2.5-32B-Instruct on the s1K trace set Source: https://arxiv.org/abs/2501.19393 - Qwen3-8B: on-policy distillation vs RL GPU-hours: 10 x cheaper (1,800 vs 17,920 GPU-hours) Qwen3 Technical Report Table 21; distillation also scored higher (AIME'24 74.4 vs 67.6) Source: https://arxiv.org/html/2505.09388v1 - Compute efficiency of a distilled 8B vs training the same model from scratch: 2,000 x (2026 controlled benchmark study) "creating a distilled 8B model is over 2,000 times more compute-efficient than training its vanilla counterpart" Source: https://arxiv.org/abs/2602.20164 KEY FINDINGS ------------ 1. Distillation is now measurably better than RL at instilling reasoning in mid-size models DeepSeek's own ablation is the cleanest evidence: applying large-scale RL directly to Qwen2.5-32B (DeepSeek-R1-Zero-Qwen-32B) reached AIME 2024 pass@1 of 47.0 and MATH-500 of 91.6, while plain supervised fine-tuning on 800k traces sampled from the 671B DeepSeek-R1 teacher reached 72.6 and 94.3 on the same base model. Qwen3 reproduced the pattern a few months later with on-policy logit distillation: 74.4 vs 67.6 AIME'24 at one tenth the GPU-hours. The 2025 follow-up by Kim et al. refines the claim — RL with verifiable rewards raises pass@1 but often not pass@k, whereas distillation can raise both when it injects genuinely new knowledge. Sources: https://arxiv.org/html/2501.12948v1 https://arxiv.org/html/2505.09388v1 https://arxiv.org/abs/2505.14216 2. The field's centre of gravity moved from off-policy traces to on-policy token-level grading Off-policy distillation trains the student on the teacher's own perfect outputs, so errors compound at inference — exposure bias that the 2026 survey by Song and Zheng argues scales roughly with the square of sequence length. On-policy distillation instead samples trajectories from the student and has the teacher score each token, combining RL's distribution match with distillation's dense signal. arXiv mentions went from 38 in 2025 to 359 in the first eight months of 2026, and by mid-2026 Qwen3, DeepSeek-V4 and NVIDIA's Nemotron 3 Ultra (and GLM-5, per a Hugging Face community write-up) had all made it a primary post-training stage. Sources: https://arxiv.org/abs/2604.00626 https://thinkingmachines.ai/blog/on-policy-distillation/ https://huggingface.co/blog/sergiopaniego/distillation-2026 3. Sample efficiency collapsed by three orders of magnitude in a single year DeepSeek-R1-Distill used 800,000 curated reasoning traces. Within weeks, Berkeley's Sky-T1 matched o1-preview-class math with 17k traces for under $450, Bespoke-Stratos hit AIME 2024 63.3 with the same 17k budget, Stanford's s1 reached AIME 56.7 and MATH-500 93.0 on 1,000 examples in 26 minutes on 16 H100s, and LIMO reported 63.3/95.6 from 817 samples — roughly 1% of the data used by prior approaches. The academic lesson is that a strong base model already contains most of the reasoning capability; the traces mainly teach a format and a search policy. Sources: https://novasky-ai.github.io/posts/sky-t1/ https://huggingface.co/bespokelabs/Bespoke-Stratos-32B https://arxiv.org/abs/2501.19393 https://arxiv.org/abs/2502.03387 4. Retention degrades sharply below ~14B parameters on hard reasoning, but barely at all on easier benchmarks Across the DeepSeek-R1-Distill family the same teacher yields very different retention depending on task difficulty and student size. On MATH-500 the 1.5B student already retains 86.2% of the 671B teacher and the 32B student retains 96.9%. On AIME 2024 the 1.5B student retains only 36.2% while the 32B retains 91.0%. Retention is therefore not a property of the method but of the interaction between benchmark difficulty and student capacity — the 'capacity gap' that Cho and Hariharan identified in vision in 2019 and that Kajitsuka et al. revisited for chain-of-thought distillation in April 2026. Sources: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B https://arxiv.org/abs/1910.01348 https://arxiv.org/abs/2604.08880 5. Distillation now has a scaling law, and it says teachers can be too strong Apple's Distillation Scaling Laws (ICML 2025) is a large-scale controlled study of distillation — students from 143M to 12.6B parameters, teachers spanning a similar range, up to 512B training tokens, figures taken from the paper body and the Apple ML Research write-up rather than the abstract — and produced a law predicting student cross-entropy from the compute split between teacher and student. The headline practical result: distillation beats supervised learning only up to a compute level that scales predictably with student size, and a teacher that is too capable for the student's budget makes things worse rather than better. Sources: https://arxiv.org/abs/2502.08606 https://machinelearning.apple.com/research/distillation-scaling-laws 6. Pretraining-time distillation is now standard at the frontier, not just a fine-tuning trick Gemma 2's 2B and 9B models replaced next-token prediction with distillation from a larger teacher and were trained on more than 50x the compute-optimal token count; the paper's ablation shows a 2B model trained on 500B tokens scoring 60.3 average from scratch versus 67.7 distilled. Gemma 3 refined this by sampling 256 teacher logits per token, renormalising, and using cross-entropy over that sample. Meta pruned Llama 3.1 8B in one shot and used logits from Llama 3.1 8B and 70B as token-level targets during pretraining of Llama 3.2 1B and 3B. NVIDIA's Minitron showed pruning-plus-distillation needs up to 40x fewer tokens per model than training from scratch. Sources: https://arxiv.org/html/2408.00118v1 https://arxiv.org/html/2503.19786v1 https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ https://arxiv.org/abs/2407.14679 7. The divergence you minimise determines what kind of student you get Classical KD minimises forward KL, which is mode-covering: the student spreads mass over regions the teacher barely visits, which is fine for classification but produces hallucinated low-probability text in generation. MiniLLM (ICLR 2024) swapped in reverse KL, which is mode-seeking, and reported lower exposure bias, better calibration and stronger long-text generation across 120M-13B students. GKD generalised this to a JSD family evaluated on student-sampled sequences, and by 2026 reverse KL over on-policy rollouts had become the default objective in industrial post-training pipelines. Sources: https://arxiv.org/abs/2306.08543 https://arxiv.org/abs/2306.13649 https://thinkingmachines.ai/blog/on-policy-distillation/ 8. A small student can beat a much larger teacher when the transferred signal is rationales, not labels Distilling Step-by-Step extracted natural-language rationales alongside labels and trained a 770M T5 that outperformed few-shot-prompted 540B PaLM while using only 80% of the available data — a ~700x parameter reduction. Orca (13B) learned from GPT-4 explanation traces and beat Vicuna-13B by over 100% on Big-Bench Hard and 42% on AGIEval. MobileBERT even exceeds the same-size BERT-base baseline on SQuAD (its actual teacher is a custom IB-BERT-LARGE, which it does not beat) (F1 90.0 vs 88.5 on v1.1, 79.2 vs 77.1 on v2.0) at 4.3x smaller and 5.5x faster. The pattern: intermediate supervision, not just the final answer, is what closes the gap. Sources: https://arxiv.org/abs/2305.02301 https://arxiv.org/abs/2306.02707 https://arxiv.org/abs/2004.02984 9. Distillation has become measurable — and the measurements suggest widespread homogenisation The Quantification of Large Language Model Distillation framework (ACL 2025) proposes Response Similarity Evaluation and Identity Consistency Evaluation to estimate how heavily a model was distilled from another, and reports that most well-known closed- and open-source LLMs exhibit high distillation degrees, with base models more distilled than aligned ones. Read alongside Shumailov et al.'s Nature result that recursive training on generated data destroys distribution tails, this defines a genuine research risk: a literature that trains overwhelmingly on frontier-model outputs may be narrowing the diversity it depends on. Sources: https://arxiv.org/abs/2501.12619 https://www.nature.com/articles/s41586-024-07566-y 10. Reproducibility improved in 2025-2026 — distillation is one of the few frontier techniques the open literature can actually replicate Hugging Face's Open-R1 reproduced DeepSeek's reported MATH-500 result for R1-Distill-Qwen-32B (95.6 vs 94.3 reported, per the Open-R1 repository evaluation table) and released OpenR1-Math-220k and Mixture-of-Thoughts as fully open training data. OpenThoughts ran 1,000+ controlled ablations to build OpenThoughts3-1.2M, whose 7B student beat DeepSeek-R1-Distill-Qwen-7B by 15.3 points on AIME 2025, 17.2 on LiveCodeBench and 20.5 on GPQA Diamond. DeepSeek-R1 itself became the first major open-weight LLM published after independent peer review, in Nature in September 2025. Sources: https://github.com/huggingface/open-r1 https://huggingface.co/blog/open-r1 https://arxiv.org/abs/2506.04178 https://www.nature.com/articles/s41586-025-09422-z TABLES -------- TABLE: Taxonomy of distillation methods [id: kd-method-taxonomy, 19 rows] The three canonical axes of the KD literature — what knowledge is transferred, how the teacher is accessed, and whether the student's own samples are used — mapped onto the seminal work for each family. Method family | Knowledge transferred | Teacher access | Sampling regime | Seminal work | Year | Citations (citations) | Source URL ------------- | --------------------- | -------------- | --------------- | ------------ | ---- | --------------------- | ---------- Ensemble compression | Labels on unlabelled transfer set | black-box | offline | Buciluă, Caruana & Niculescu-Mizil, Model Compression | 2006 | 2907 | https://dl.acm.org/doi/10.1145/1150402.1150464 Response / logit-based | Temperature-softened output distribution ('dark knowledge') | white-box | offline | Hinton, Vinyals & Dean | 2015 | 25899 | https://arxiv.org/abs/1503.02531 Feature-based (hints) | Intermediate activations via a regressor | white-box | offline | FitNets (Romero et al.) | 2015 | 4873 | https://arxiv.org/abs/1412.6550 Attention transfer | Spatial attention maps | white-box | offline | Zagoruyko & Komodakis | 2017 | 3209 | https://arxiv.org/abs/1612.03928 Relation-based | Pairwise/triplet structure of the embedding space | white-box | offline | Relational KD (Park et al.) | 2019 | 2038 | https://arxiv.org/abs/1904.05068 Contrastive representation | Mutual information between teacher and student features | white-box | offline | CRD (Tian, Krishnan & Isola) | 2020 | 1406 | https://arxiv.org/abs/1910.10699 Sequence-level KD | Teacher-generated output sequences as hard targets | black-box | offline | Kim & Rush | 2016 | 1478 | https://arxiv.org/abs/1606.07947 Self-attention relation | Query-key and value-value relation matrices | white-box | offline | MiniLM (Wang et al.) | 2020 | 2550 | https://arxiv.org/abs/2002.10957 Online / mutual | Peer predictions, no fixed teacher | white-box | online | Deep Mutual Learning (Zhang et al.) | 2018 | 2033 | https://arxiv.org/abs/1706.00384 Self-distillation | The model's own earlier generation | white-box | self | Born-Again Neural Networks (Furlanello et al.) | 2018 | 1279 | https://arxiv.org/abs/1805.04770 Reverse-KL policy distillation | Mode-seeking match to teacher distribution | white-box | on-policy | MiniLLM (Gu et al.) | 2024 | 116 | https://arxiv.org/abs/2306.08543 Generalized on-policy KD | Generalized JSD on student-generated sequences | white-box | on-policy | GKD (Agarwal et al.) | 2024 | 732 | https://arxiv.org/abs/2306.13649 Rationale / CoT distillation | Natural-language rationales as a second training signal | black-box | offline | Distilling Step-by-Step (Hsieh et al.) | 2023 | 1074 | https://arxiv.org/abs/2305.02301 Explanation-trace imitation | Full step-by-step GPT-4 explanation traces | black-box | offline | Orca (Mukherjee et al.) | 2023 | 433 | https://arxiv.org/abs/2306.02707 Prune + distil | Logits used to recover accuracy after structured pruning | white-box | offline | Minitron (Muralidharan et al.) | 2024 | 189 | https://arxiv.org/abs/2407.14679 Draft-model distillation | Alignment of a small drafter to the target for speculative decoding | white-box | on/off-policy | DistillSpec (Zhou et al.) | 2024 | 171 | https://arxiv.org/abs/2310.08461 Dataset distillation | A synthetic dataset rather than a model | n/a | offline | Dataset Distillation (Wang et al.) | 2018 | 386 | https://arxiv.org/abs/1811.10959 Cross-tokenizer distillation | Logits mapped across mismatched vocabularies | white-box | off/on-policy | Universal Logit Distillation (Boizard et al.) | 2024 | 60 | https://arxiv.org/abs/2402.12030 Multi-teacher on-policy distillation | Weighted reverse KL against >10 domain specialists on student rollouts | white-box | on-policy | Nemotron 3 Ultra (NVIDIA) | 2026 | 14 | https://arxiv.org/abs/2606.15007 Notes: Citation counts retrieved from the Semantic Scholar Graph API on 2026-09-04. 'Year' is the year of the archival venue where one exists, otherwise the arXiv year. MiniLLM's count (116) could not be confirmed: the Semantic Scholar record appears to be split between the arXiv preprint (arXiv:2306.08543) and the ICLR 2024 proceedings entry, so the figure shown is likely an undercount and is not comparable to neighbouring rows such as GKD. Sources: https://api.semanticscholar.org/graph/v1/paper/batch https://arxiv.org/abs/2006.05525 https://arxiv.org/abs/2402.13116 TABLE: Student-vs-teacher retention on hard reasoning benchmarks [id: student-teacher-retention, 13 rows] Reported pass@1 scores for distilled students against their teachers, with retention computed as student/teacher. Shows how retention collapses on the hardest benchmark (AIME) for small students while staying near-parity on MATH-500. Student | Student params (B) | Teacher | Benchmark | Teacher (%) | Student (%) | Retention (%) | Source URL ------- | ------------------ | ------- | --------- | ----------- | ----------- | ------------- | ---------- DeepSeek-R1-Distill-Qwen-1.5B | 1.5 | DeepSeek-R1 (671B) | AIME 2024 | 79.8 | 28.9 | 36.2 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Llama-8B | 8 | DeepSeek-R1 (671B) | AIME 2024 | 79.8 | 50.4 | 63.2 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-7B | 7 | DeepSeek-R1 (671B) | AIME 2024 | 79.8 | 55.5 | 69.5 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-14B | 14 | DeepSeek-R1 (671B) | AIME 2024 | 79.8 | 69.7 | 87.3 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Llama-70B | 70 | DeepSeek-R1 (671B) | AIME 2024 | 79.8 | 70 | 87.7 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-32B | 32 | DeepSeek-R1 (671B) | AIME 2024 | 79.8 | 72.6 | 91 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-1.5B | 1.5 | DeepSeek-R1 (671B) | MATH-500 | 97.3 | 83.9 | 86.2 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-32B | 32 | DeepSeek-R1 (671B) | MATH-500 | 97.3 | 94.3 | 96.9 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-32B | 32 | DeepSeek-R1 (671B) | GPQA Diamond | 71.5 | 62.1 | 86.9 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B Sky-T1-32B-Preview | 32 | QwQ-32B-Preview | AIME 2024 | 50 | 43.3 | 86.6 | https://novasky-ai.github.io/posts/sky-t1/ Bespoke-Stratos-32B | 32 | DeepSeek-R1 (671B) | AIME 2024 | 79.8 | 63.3 | 79.3 | https://huggingface.co/bespokelabs/Bespoke-Stratos-32B s1.1-32B | 32 | DeepSeek-R1 (671B) | MATH-500 | 97.3 | 95.4 | 98 | https://arxiv.org/html/2501.19393v3 Qwen3-8B (on-policy distilled) | 8 | Qwen3-32B | AIME 2024 | 81.4 | 74.4 | 91.4 | https://arxiv.org/html/2505.09388v1 Notes: Retention = student score / teacher score, computed from the sources' reported pass@1 figures. DeepSeek-R1 GPQA Diamond teacher score is 71.5 as reported in the R1 paper. Cross-paper comparisons carry evaluation-harness differences: Open-R1 independently measured DeepSeek-R1-Distill-Qwen-32B at 95.6 on MATH-500 versus 94.3 reported — roughly a 1.3-point harness spread, large enough to swamp several of the cross-paper retention deltas in this table. Sources: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B https://arxiv.org/html/2501.12948v1 https://github.com/huggingface/open-r1 TABLE: The BERT era: how much quality survived compression [id: bert-era-compression, 6 rows] Encoder-model distillation results that set the field's expectations before the LLM wave. These are the numbers every later paper benchmarks against. Student | Teacher | Technique | Size reduction | Speedup | Quality retained | Year | Source URL ------- | ------- | --------- | -------------- | ------- | ---------------- | ---- | ---------- DistilBERT | BERT-base | Pretraining-time logit KD + cosine-distance loss | 40% smaller | 60% faster | 97% of language-understanding capability | 2019 | https://arxiv.org/abs/1910.01108 Patient-KD BERT (6L) | BERT-base | Learn from multiple intermediate layers | ~2x fewer layers | undisclosed | task-dependent; established multi-layer KD for BERT | 2019 | https://arxiv.org/abs/1908.09355 TinyBERT (4 layers) | BERT-base | Two-stage transformer distillation (pretraining + task) | 7.5x smaller | 9.4x faster | >96.8% of teacher on GLUE | 2020 | https://arxiv.org/abs/1909.10351 TinyBERT (6 layers) | BERT-base | Two-stage transformer distillation | 2x smaller | undisclosed | on par with teacher | 2020 | https://arxiv.org/abs/1909.10351 MobileBERT | IB-BERT-large (custom teacher) | Bottleneck architecture + progressive knowledge transfer | 4.3x smaller | 5.5x faster (62 ms on Pixel 4) | GLUE 77.7 vs BERT-base 78.3; SQuAD v1.1 F1 90.0 vs 88.5 (exceeds teacher reference) | 2020 | https://arxiv.org/abs/2004.02984 MiniLM | BERT-base / UniLM | Deep self-attention relation distillation (Q-K and V-V) | task-agnostic 6-layer students | undisclosed | state of the art for task-agnostic compression at the time | 2020 | https://arxiv.org/abs/2002.10957 Notes: MobileBERT's BERT-base GLUE reference of 78.3 is derived from the paper's statement that MobileBERT's 77.7 is '0.6 lower than BERT_BASE'. Speed and size figures are as reported by the authors on their own hardware and are not directly comparable across papers. Sources: https://arxiv.org/abs/1910.01108 https://arxiv.org/abs/1909.10351 https://arxiv.org/abs/2004.02984 https://arxiv.org/abs/2002.10957 https://arxiv.org/abs/1908.09355 TABLE: Compute and data cost per distillation recipe [id: compute-cost-per-method, 14 rows] What each published recipe actually cost, in the units its authors reported. The spread — from 7 GPU-hours to a full pretraining run — is the single most striking fact about the 2025-2026 literature. Run / method | Approach | Teacher | Training data | Compute | Reported cost | Source URL ------------ | -------- | ------- | ------------- | ------- | ------------- | ---------- s1-32B (Stanford, 2025) | Offline SFT on curated traces + budget forcing | Gemini 2.0 Flash Thinking | 1,000 samples (s1K) | 26 min on 16x H100 (~7 GPU-hours) | undisclosed in the paper (~$50 widely reported in press coverage) | https://arxiv.org/abs/2501.19393 LIMO (2025) | Offline SFT on hand-curated reasoning chains | curated, multi-source | 817 samples | undisclosed | undisclosed | https://arxiv.org/abs/2502.03387 Sky-T1-32B-Preview (Berkeley NovaSky, 2025) | Offline SFT on rejection-sampled teacher traces | QwQ-32B-Preview | 17k (10k math, 5k code, 1k science/puzzle) | 19 hours on 8x H100 (152 GPU-hours) | under $450 | https://novasky-ai.github.io/posts/sky-t1/ Bespoke-Stratos-32B (2025) | Offline SFT, Sky-T1 pipeline with modified filtering | DeepSeek-R1 | 17k (47x fewer than R1-Distill) | trace generation ~1.5 hours with DeepSeek-R1 | undisclosed | https://www.bespokelabs.ai/blog/bespoke-stratos-the-unreasonable-effectiveness-of-reasoning-distillation Stanford Alpaca-7B (2023) | Black-box output imitation (self-instruct) | text-davinci-003 | 52k instruction-following demonstrations | undisclosed | under $500 API + under $600 total | https://crfm.stanford.edu/2023/03/13/alpaca.html DeepSeek-R1-Distill series (2025) | Offline SFT only, no RL on the student | DeepSeek-R1 (671B MoE, 37B active) | 800k rejection-sampled traces, 2-3 epochs | undisclosed | undisclosed | https://arxiv.org/html/2501.12948v1 OpenThoughts3 / OpenThinker3-7B (2025) | Offline SFT on an ablation-optimised recipe | QwQ-32B | 1.2M examples | 1,000+ controlled pipeline experiments | undisclosed | https://arxiv.org/abs/2506.04178 Qwen3-8B strong-to-weak distillation (2025) | Off-policy then on-policy logit KL | Qwen3-32B / Qwen3-235B-A22B | on-policy rollouts | 1,800 GPU-hours (vs 17,920 for the RL alternative) | ~1/10 the GPU-hours of RL | https://arxiv.org/html/2505.09388v1 On-policy distillation of Qwen3-8B (Thinking Machines, 2025) | Reverse-KL token grading on student rollouts | Qwen3-32B | 77k prompts x 4 samples, ~150 steps | 8.4e19 teacher FLOPs + 8.2e19 student FLOPs | 9-30x cheaper than the SFT extrapolation, depending on teacher-FLOP amortisation | https://thinkingmachines.ai/blog/on-policy-distillation/ Minitron / Nemotron-4 15B to 8B and 4B (2024) | Structured pruning + KD retraining | Nemotron-4 15B | <3% of the original pretraining data | up to 40x fewer training tokens per model | 1.8x compute saving for the full model family | https://arxiv.org/abs/2407.14679 Llama 3.2 1B / 3B (Meta, 2024) | One-shot structured pruning + logit distillation in pretraining | Llama 3.1 8B and 70B | pretraining corpus with teacher logits as token-level targets | undisclosed | undisclosed | https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ Gemma 2 2B (Google, 2024) | Distillation replaces next-token prediction in pretraining | larger Gemma teacher (7B in ablation, 27B for release) | >50x compute-optimal tokens; 500B-token ablation | full pretraining run | ablation: 60.3 avg from scratch vs 67.7 distilled | https://arxiv.org/html/2408.00118v1 Gemma 3 (Google, 2025) | Pretraining KD with 256 sampled teacher logits per token | larger instruction-tuned Gemma teacher | full pretraining corpus | full pretraining run | undisclosed | https://arxiv.org/html/2503.19786v1 Distilled 8B vs from-scratch 8B (benchmark study, 2026) | Controlled comparison of distillation vs vanilla pretraining | various | n/a | >2,000x more compute-efficient than the vanilla counterpart | undisclosed | https://arxiv.org/abs/2602.20164 Notes: Costs are as reported by the original authors and use different accounting (rented GPU-hours, API spend, FLOPs). They are not directly comparable; treat them as order-of-magnitude signals. 'undisclosed' means no figure was published, not that the run was free. Sources: https://novasky-ai.github.io/posts/sky-t1/ https://arxiv.org/abs/2501.19393 https://arxiv.org/html/2505.09388v1 https://thinkingmachines.ai/blog/on-policy-distillation/ https://arxiv.org/abs/2407.14679 TABLE: Distillation vs reinforcement learning, head to head [id: distillation-vs-rl, 4 rows] Two independent controlled comparisons on the same base model, plus the theoretical result that explains the difference. Study | Base model | RL result | Distillation result | Verdict | Source URL ----- | ---------- | --------- | ------------------- | ------- | ---------- DeepSeek-R1 (2025), Table 6 | Qwen2.5-32B | AIME 2024 47.0, MATH-500 91.6 (DeepSeek-R1-Zero-Qwen-32B, large-scale RL) | AIME 2024 72.6, MATH-500 94.3 (SFT on 800k R1 traces, no RL) | Distillation wins by 25.6 AIME points on identical base | https://arxiv.org/html/2501.12948v1 Qwen3 Technical Report (2025), Table 21 | Qwen3-8B | AIME'24 67.6, AIME'25 55.5; 17,920 GPU-hours | AIME'24 74.4, AIME'25 65.5; 1,800 GPU-hours | Distillation wins by 6.8 / 10.0 points at 1/10 the compute | https://arxiv.org/html/2505.09388v1 Kim et al. (2025), RL vs Distillation | various reasoning students | RLVR raises pass@1 but often not pass@k; gains concentrate on easy questions | Distillation can raise both accuracy and capability when it injects new knowledge | RL sharpens; distillation can genuinely extend | https://arxiv.org/abs/2505.14216 Thinking Machines Lab (2025) | Qwen3-8B from a 400k SFT checkpoint | reference RL trajectory to the same target | reaches teacher performance ~7-10x faster; 50-100x total compute reduction accounting for context and batch differences | On-policy distillation dominates RL on this task | https://thinkingmachines.ai/blog/on-policy-distillation/ Notes: The DeepSeek and Qwen comparisons are the field's two cleanest controlled ablations because both hold the base model fixed. Both come from the labs that shipped the models, so independent replication (e.g. Open-R1) matters. Sources: https://arxiv.org/html/2501.12948v1 https://arxiv.org/html/2505.09388v1 https://arxiv.org/abs/2505.14216 TABLE: Open research problems as of September 2026 [id: open-problems, 8 rows] Where the theory is still behind the practice. Problem | State of the art | Why it is unresolved | Source URL ------- | ---------------- | -------------------- | ---------- Capacity gap | Cho & Hariharan (2019) showed bigger teachers are not better teachers and proposed teacher early-stopping; Kajitsuka et al. (2026) show the effect varies widely by task and teacher-student pairing in CoT distillation | No predictive rule for choosing the right teacher for a given student budget | https://arxiv.org/abs/2604.08880 Can a student exceed its teacher? | MobileBERT exceeds the same-size BERT-base baseline on SQuAD, but not its own teacher (IB-BERT-LARGE); no clean published case of a student exceeding its own teacher is offered here. 2026 work explores objectives that deliberately push past the teacher distribution | Standard KD objectives have an imitation ceiling by construction; exceeding it requires an extra signal (verifier, search, or new data) | https://arxiv.org/abs/2004.02984 Scaling laws | Apple's Distillation Scaling Laws (ICML 2025) fit student loss to the teacher/student compute split across 143M-12.6B students | The law is fit on pretraining cross-entropy, not on downstream reasoning, and does not yet cover on-policy regimes | https://machinelearning.apple.com/research/distillation-scaling-laws Does KD actually match the teacher function? | Stanton et al. (2021) found students often fail to match the teacher's predictive distribution even when generalisation improves | Optimisation, not capacity, appears to be the bottleneck — and it remains poorly characterised | https://arxiv.org/abs/2106.05945 Tokenizer mismatch | ULD (2024), approximate likelihood matching (NeurIPS 2025), byte-level interfaces and projection-guided methods (2026) | Logit-level transfer across different vocabularies is still lossy; most cross-family distillation falls back to black-box text | https://arxiv.org/abs/2503.20083 Homogenisation and model collapse | Distillation-degree metrics (RSE/ICE, ACL 2025) report high distillation degrees across well-known LLMs; Shumailov et al. (Nature 2024) show recursive synthetic training destroys distribution tails | No agreed measurement of how much real-data grounding a training corpus needs to stay safe | https://arxiv.org/abs/2501.12619 Evaluation contamination in the distillation loop | Independent replications (Open-R1) differ from reported numbers by ~1.3 points on MATH-500 | Teacher traces are generated on the same benchmark families used for evaluation; harness differences compound the ambiguity | https://github.com/huggingface/open-r1 Attribution and provenance | Response Similarity Evaluation and Identity Consistency Evaluation give a first quantitative handle | No method reliably proves which teacher a given open-weight model was distilled from | https://arxiv.org/abs/2501.12619 Notes: Each row states the strongest published position as of 2026-09-04, not a consensus. Sources: https://arxiv.org/abs/2502.08606 https://arxiv.org/abs/2106.05945 https://arxiv.org/abs/2604.00626 CHART DATA ---------- CHART: arXiv papers with 'knowledge distillation' in the abstract, 2015-2026 [id: papers-per-year, type: bar, unit: papers] Year | arXiv submissions (papers) 2015 | 2 2016 | 6 2017 | 18 2018 | 51 2019 | 159 2020 | 327 2021 | 451 2022 | 624 2023 | 777 2024 | 971 2025 | 1084 2026 | 709 Notes: Counts retrieved from the arXiv API on 2026-09-04 using search_query=abs:"knowledge distillation" restricted per submission year. 2026 covers 1 January to 4 September only, so the full year is on pace for roughly 1,050. Growth is 542x from 2015 to 2025. Sources: https://export.arxiv.org/api/query?search_query=abs:%22knowledge%20distillation%22%20AND%20submittedDate:%5B202501010000%20TO%20202512312359%5D&max_results=1 https://info.arxiv.org/help/api/user-manual.html CHART: Retention of teacher score vs student size (DeepSeek-R1-Distill family) [id: retention-vs-student-size, type: scatter, unit: %] Student parameters (B) | AIME 2024 (hard) (%) 1.5 | 36.2 7 | 69.5 8 | 63.2 14 | 87.3 32 | 91 70 | 87.7 Student parameters (B) | MATH-500 (moderate) (%) 1.5 | 86.2 7 | 95.4 8 | 91.6 14 | 96.5 32 | 96.9 70 | 97.1 Student parameters (B) | GPQA Diamond (knowledge) (%) 1.5 | 47.3 7 | 68.7 8 | 68.5 14 | 82.7 32 | 86.9 70 | 91.2 Notes: Retention = student pass@1 / DeepSeek-R1 pass@1, using teacher scores of 79.8 (AIME 2024), 97.3 (MATH-500) and 71.5 (GPQA Diamond). The 8B point is a Llama-3.1 student rather than Qwen, which is why it sits below the 7B Qwen student on AIME. The gap between the AIME and MATH-500 curves is the clearest published picture of the capacity gap. Sources: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B https://arxiv.org/html/2501.12948v1 CHART: Method adoption: arXiv mentions by sub-topic, 2020-2026 [id: method-adoption-by-year, type: line, unit: papers] Year | on-policy distillation (papers) 2020 | 9 2021 | 7 2022 | 9 2023 | 10 2024 | 17 2025 | 38 2026 | 359 Year | self-distillation (papers) 2020 | 28 2021 | 55 2022 | 75 2023 | 123 2024 | 132 2025 | 209 2026 | 427 Year | dataset distillation (papers) 2020 | 5 2021 | 2 2022 | 18 2023 | 51 2024 | 88 2025 | 104 2026 | 68 Year | chain-of-thought distillation (papers) 2020 | 0 2021 | 0 2022 | 0 2023 | 4 2024 | 4 2025 | 12 2026 | 14 Notes: arXiv API full-text ('all:') phrase counts per submission year, retrieved 2026-09-04. 2026 covers only 1 January to 4 September. The on-policy curve is the story of 2026: a 9.4x jump in eight months, matching its adoption in Qwen3, DeepSeek-V4 and Nemotron 3 Ultra. Phrase counting is a proxy — it over-counts passing mentions and misses papers that use different terminology. Sources: https://export.arxiv.org/api/query?search_query=all:%22on-policy%20distillation%22%20AND%20submittedDate:%5B202601010000%20TO%20202609042359%5D&max_results=1 https://arxiv.org/abs/2604.00626 CHART: Same base model, two training recipes [id: distill-vs-rl-bars, type: bar, unit: %] Setting | Reinforcement learning (%) Qwen2.5-32B (DeepSeek, AIME'24) | 47 Qwen3-8B (AIME'24) | 67.6 Qwen3-8B (AIME'25) | 55.5 Setting | Distillation (%) Qwen2.5-32B (DeepSeek, AIME'24) | 72.6 Qwen3-8B (AIME'24) | 74.4 Qwen3-8B (AIME'25) | 65.5 Notes: DeepSeek's comparison is offline SFT on 800k teacher traces versus large-scale RL on the identical Qwen2.5-32B base. Qwen3's is on-policy logit distillation versus RL, at 1,800 vs 17,920 GPU-hours. Sources: https://arxiv.org/html/2501.12948v1 https://arxiv.org/html/2505.09388v1 CHART: Citation counts of landmark distillation and distillation-adjacent papers [id: landmark-citations, type: bar, unit: citations] Paper | Semantic Scholar citations (2026-09-04) (citations) Hinton 2015 (KD) | 25899 DistilBERT 2019 | 10433 DeepSeek-R1 2025 | 5597 FitNets 2015 | 4873 KD Survey (Gou 2021) | 4570 Attention Transfer 2017 | 3209 Buciluă 2006 (Model Compression) | 2907 MiniLM 2020 | 2550 TinyBERT 2020 | 2504 Relational KD 2019 | 2038 Deep Mutual Learning 2018 | 2033 Speculative Decoding 2023 | 1950 Seq-Level KD 2016 | 1478 CRD 2020 | 1406 Born-Again NN 2018 | 1279 Distilling Step-by-Step 2023 | 1074 GKD 2024 | 732 Orca 2023 | 433 Notes: Retrieved from the Semantic Scholar Graph API batch endpoint on 2026-09-04. Citation counts move; treat these as a September 2026 snapshot. Hinton et al. alone accounts for more citations than the next four papers combined. Speculative Decoding (Leviathan et al. 2023) is included as distillation-adjacent: it is not a KD method, but it created the demand for draft-model distillation as a distinct research problem. Sources: https://api.semanticscholar.org/graph/v1/paper/arXiv:1503.02531?fields=title,year,citationCount https://www.semanticscholar.org/paper/0c908739fbff75f03469d13d4a1a07de3414ee19 CHART: Training samples needed to reach o1-preview-class math reasoning [id: sample-efficiency-collapse, type: bar, unit: samples] Recipe | Reasoning traces used (samples) OpenThoughts3 (Jun 2025) | 1200000 DeepSeek-R1-Distill (Jan 2025) | 800000 Alpaca (Mar 2023) | 52000 Sky-T1 (Jan 2025) | 17000 Bespoke-Stratos (Jan 2025) | 17000 s1 (Jan 2025) | 1000 LIMO (Feb 2025) | 817 Notes: These recipes do not all target the same capability or reach the same score, so this is a chart about data budgets, not a quality ranking. OpenThoughts3 uses more data because it targets a 7B student and a higher absolute ceiling; s1 and LIMO use a 32B base whose latent ability is largely already present. LIMO figures are from v3 (July 2025); v1 reported AIME24 57.1 from the same 817 samples. Sources: https://arxiv.org/html/2501.12948v1 https://novasky-ai.github.io/posts/sky-t1/ https://arxiv.org/abs/2501.19393 https://arxiv.org/abs/2502.03387 https://arxiv.org/abs/2506.04178 https://crfm.stanford.edu/2023/03/13/alpaca.html DETAIL RECORDS -------------- TABLE: Papers in the academic register [46 rows, from data/academic.json $.extras] Paper | Authors | Year | Venue | Category | Method | Citations (citations) | What it does | Why it matters | Source URL ----- | ------- | ---- | ----- | -------- | ------ | --------------------- | ------------ | -------------- | ---------- Model Compression | Cristian Buciluă, Rich Caruana, Alexandru Niculescu-Mizil | 2006 | KDD | origins | Ensemble compression via labelled synthetic transfer set | 2907 | Compress a large ensemble into a single small neural net by having the ensemble label unlabelled or synthetic data. | The original teacher-student result, nine years before the word 'distillation' was applied to it. | https://dl.acm.org/doi/10.1145/1150402.1150464 Distilling the Knowledge in a Neural Network | Geoffrey Hinton, Oriol Vinyals, Jeff Dean | 2015 | arXiv / NIPS 2014 Deep Learning Workshop | origins | Response/logit KD with temperature-softened soft targets | 25899 | Softened teacher output distributions carry 'dark knowledge' that one-hot labels discard. | The field's canonical reference; defines temperature, soft targets, and the KD loss that almost every later method modifies. | https://arxiv.org/abs/1503.02531 FitNets: Hints for Thin Deep Nets | Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, Yoshua Bengio | 2015 | ICLR | taxonomy-feature | Feature-based (hint layer + regressor) | 4873 | Match intermediate activations, not just outputs, so thin deep students can be trained at all. | Opened the feature-based branch of the taxonomy and showed students can outperform teachers. | https://arxiv.org/abs/1412.6550 Sequence-Level Knowledge Distillation | Yoon Kim, Alexander M. Rush | 2016 | EMNLP | sequence | Sequence-level KD on teacher-generated outputs | 1478 | For generation, train on the teacher's whole output sequences rather than token-level distributions. | The direct ancestor of every black-box LLM distillation recipe, including DeepSeek-R1-Distill. | https://arxiv.org/abs/1606.07947 Paying More Attention to Attention | Sergey Zagoruyko, Nikos Komodakis | 2017 | ICLR | taxonomy-feature | Attention-map transfer | 3209 | Transfer spatial attention maps between teacher and student CNNs. | Established that where a model looks is a transferable form of knowledge. | https://arxiv.org/abs/1612.03928 Deep Mutual Learning | Ying Zhang, Tao Xiang, Timothy M. Hospedales, Huchuan Lu | 2018 | CVPR | taxonomy-online | Online / mutual distillation between peers | 2033 | A cohort of students teach each other during training with no pretrained teacher at all. | Founding paper of online distillation; showed a fixed superior teacher is not required. | https://arxiv.org/abs/1706.00384 Born Again Neural Networks | Tommaso Furlanello, Zachary C. Lipton, Michael Tschannen, Laurent Itti, Anima Anandkumar | 2018 | ICML | taxonomy-self | Self-distillation across generations | 1279 | Distil a model into an identically-sized copy of itself, repeatedly, and it keeps improving. | Showed distillation is a regulariser, not only a compression technique. | https://arxiv.org/abs/1805.04770 On the Efficacy of Knowledge Distillation | Jang Hyun Cho, Bharath Hariharan | 2019 | ICCV | theory | Empirical study of capacity mismatch | 803 | Bigger teachers are often worse teachers; early-stopping the teacher mitigates the capacity gap. | The first rigorous statement of the capacity-gap problem that the 2026 CoT literature revisits. | https://arxiv.org/abs/1910.01348 Relational Knowledge Distillation | Wonpyo Park, Dongju Kim, Yan Lu, Minsu Cho | 2019 | CVPR | taxonomy-relation | Relation-based (distance-wise and angle-wise losses) | 2038 | Transfer the geometric relations between examples rather than their individual representations. | Defines the third arm of the standard taxonomy. | https://arxiv.org/abs/1904.05068 DistilBERT, a distilled version of BERT | Victor Sanh, Lysandre Debut, Julien Chaumond, Thomas Wolf | 2019 | arXiv / NeurIPS EMC^2 Workshop | encoder-compression | Pretraining-time triple loss (LM + distillation + cosine distance) | 10433 | 40% smaller, 60% faster, 97% of BERT's language-understanding capability retained. | The result that made distillation routine engineering practice in NLP. | https://arxiv.org/abs/1910.01108 Patient Knowledge Distillation for BERT Model Compression | Siqi Sun, Yu Cheng, Zhe Gan, Jingjing Liu | 2019 | EMNLP | encoder-compression | Multi-layer intermediate supervision | 994 | Learn from several of the teacher's intermediate layers instead of only its last one. | Established multi-layer supervision as the default for transformer distillation. | https://arxiv.org/abs/1908.09355 TinyBERT: Distilling BERT for Natural Language Understanding | Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, Qun Liu | 2020 | Findings of EMNLP | encoder-compression | Two-stage transformer distillation (pretraining + task-specific) | 2504 | A 4-layer student keeps >96.8% of BERT-base GLUE at 7.5x smaller and 9.4x faster. | Set the size/speed/quality frontier that pre-LLM compression work benchmarked against. | https://arxiv.org/abs/1909.10351 MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices | Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, Denny Zhou | 2020 | ACL | encoder-compression | Bottleneck architecture + progressive knowledge transfer from a custom IB-BERT-large teacher | 1078 | 4.3x smaller, 5.5x faster, 62 ms on a Pixel 4 — and beats the BERT-base baseline on SQuAD F1. | Frequently miscited as a student beating its teacher: MobileBERT beats the same-size BERT-base baseline on SQuAD, while its actual teacher is a custom IB-BERT-LARGE that scores above it. | https://arxiv.org/abs/2004.02984 MiniLM: Deep Self-Attention Distillation | Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, Ming Zhou | 2020 | NeurIPS | encoder-compression | Query-key and value-value relation distillation from the last self-attention layer | 2550 | Distil attention relations rather than activations, avoiding layer-mapping heuristics. | The most architecture-agnostic of the BERT-era recipes; still the base of many embedding models. | https://arxiv.org/abs/2002.10957 Dataset Distillation | Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, Alexei A. Efros | 2018 | arXiv | dataset-distillation | Gradient-based synthesis of a tiny training set | 386 | Compress a dataset, not a model: a handful of synthetic images that train a network to near-full accuracy. | Founded a parallel sub-field that reached ~104 arXiv papers in 2025. | https://arxiv.org/abs/1811.10959 Knowledge Distillation: A Survey | Jianping Gou, Baosheng Yu, Stephen J. Maybank, Dacheng Tao | 2021 | IJCV | survey | Survey | 4570 | Codifies the response/feature/relation and offline/online/self taxonomies the field still uses. | The reference taxonomy for the pre-LLM era. | https://arxiv.org/abs/2006.05525 Does Knowledge Distillation Really Work? | Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi, Andrew Gordon Wilson | 2021 | NeurIPS | theory | Empirical analysis of student-teacher fidelity | 309 | Students often improve in generalisation without ever matching the teacher's predictive distribution. | Separates 'distillation works' from 'distillation transfers the function' — still unresolved. | https://arxiv.org/abs/2106.05945 Fast Inference from Transformers via Speculative Decoding | Yaniv Leviathan, Matan Kalman, Yossi Matias | 2023 | ICML | inference | Draft-and-verify decoding with a small approximation model | 1950 | A small drafter proposes tokens that the large model verifies in parallel, with no quality loss. | Created the demand for draft-model distillation as a distinct research problem. | https://arxiv.org/abs/2211.17192 Distilling Step-by-Step! | Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, Tomas Pfister | 2023 | ACL Findings | cot-distillation | Multi-task training on labels plus extracted LLM rationales | 1074 | A 770M T5 beats few-shot 540B PaLM using 80% of the available data. | Founding paper of rationale/CoT distillation; a ~700x parameter reduction at higher accuracy. | https://arxiv.org/abs/2305.02301 Large Language Models Are Reasoning Teachers | Namgyu Ho, Laura Schmid, Se-Young Yun | 2023 | ACL | cot-distillation | Fine-tune-CoT: sample teacher rationales, filter by correctness, fine-tune a small student | 552 | Very large teachers can hand their chain-of-thought ability to models orders of magnitude smaller. | With Magister et al., one of the two papers that named CoT distillation as a method. | https://arxiv.org/abs/2212.10071 Teaching Small Language Models to Reason | Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, Aliaksei Severyn | 2023 | ACL | cot-distillation | CoT distillation with answer-consistency filtering | 439 | Distilling reasoning chains from a 540B teacher lifts small T5 students far above their scale. | Concurrent independent discovery of CoT distillation at Google. | https://arxiv.org/abs/2212.08410 Specializing Smaller Language Models towards Multi-Step Reasoning | Yao Fu, Hao Peng, Litu Ou, Ashish Sabharwal, Tushar Khot | 2023 | ICML | cot-distillation | Capability-targeted distillation with an explicit generality trade-off | 370 | Small models can be specialised into strong reasoners, but they pay for it in general ability. | First clean statement of the specialisation trade-off in distilled students. | https://arxiv.org/abs/2301.12726 Orca: Progressive Learning from Complex Explanation Traces of GPT-4 | Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, Ahmed Awadallah | 2023 | arXiv (Microsoft Research) | cot-distillation | Explanation-trace imitation with progressive learning from ChatGPT then GPT-4 | 433 | A 13B student trained on GPT-4 explanation traces beats Vicuna-13B by >100% on Big-Bench Hard. | Showed imitation of process, not just output, closes most of the gap to a frontier teacher. | https://arxiv.org/abs/2306.02707 MiniLLM: Knowledge Distillation of Large Language Models | Yuxian Gu, Li Dong, Furu Wei, Minlie Huang | 2024 | ICLR | objective | Reverse KL divergence with on-policy policy-gradient optimisation | 116 | Swap mode-covering forward KL for mode-seeking reverse KL, and the student stops hallucinating the teacher's tails. | Made reverse KL the default objective for generative distillation; scales 120M-13B. Citation count is unconfirmed — the Semantic Scholar record appears split across the arXiv and ICLR 2024 entries, so 116 likely understates the true figure. | https://arxiv.org/abs/2306.08543 On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD) | Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem | 2024 | ICLR | on-policy | Generalized KD: teacher feedback on student-generated sequences with a generalized JSD loss family | 732 | Train the student on its own outputs and let the teacher grade them, fixing the train/inference distribution mismatch. | The template for the on-policy distillation pipelines that dominate 2026 post-training. | https://arxiv.org/abs/2306.13649 Zephyr: Direct Distillation of LM Alignment | Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, Thomas Wolf | 2023 | arXiv (Hugging Face) | alignment | Distilled supervised fine-tuning followed by distilled DPO on AI-ranked preferences | 634 | Alignment itself can be distilled: MT-Bench 7.34 with no human annotation, beating Llama2-Chat-70B. | Extended distillation from capability to preference and safety behaviour. | https://arxiv.org/abs/2310.16944 DistillSpec: Improving Speculative Decoding via Knowledge Distillation | Yongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon, Afshin Rostamizadeh, Sanjiv Kumar, Jean-François Kagy, Rishabh Agarwal | 2024 | ICLR | inference | Distil the draft model to the target using on-policy data and a divergence chosen per task | 171 | Aligning the drafter to the target yields 10-45% speedups over standard speculative decoding. | Turns distillation into a latency technique rather than a capacity technique. | https://arxiv.org/abs/2310.08461 Compact Language Models via Pruning and Knowledge Distillation (Minitron) | Saurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jan Kautz, Pavlo Molchanov | 2024 | NeurIPS | pruning-distillation | Depth/width/attention/MLP pruning + KD-based retraining | 189 | Derive 8B and 4B models from a 15B parent with up to 40x fewer training tokens and <3% of the data. | Made prune-then-distil the standard way to produce a small-model family from one pretraining run. | https://arxiv.org/abs/2407.14679 Gemma 2: Improving Open Language Models at a Practical Size | Gemma Team, Google DeepMind | 2024 | arXiv | pretraining-kd | Distillation replaces next-token prediction during pretraining | 2389 | Train the 2B and 9B models on teacher distributions for >50x the compute-optimal token count. | Reframed distillation as a way to simulate training beyond the available data, not just to compress. | https://arxiv.org/html/2408.00118v1 DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning | DeepSeek-AI | 2025 | arXiv; Nature 645(8081) 2025 | reasoning-distillation | Sequence-level SFT on 800k rejection-sampled traces from a 671B MoE teacher | 5597 | Six open students from 1.5B to 70B, with the 32B beating large-scale RL on the same base by 25.6 AIME points. | The most consequential distillation release to date; first open-weight frontier LLM published after peer review. | https://arxiv.org/html/2501.12948v1 s1: Simple test-time scaling | Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, Tatsunori Hashimoto | 2025 | EMNLP | reasoning-distillation | SFT on 1,000 curated Gemini Flash Thinking traces plus budget forcing at inference | 1427 | AIME 2024 56.7 and MATH-500 93.0 from 1,000 samples and 26 minutes on 16 H100s. | The sample-efficiency floor of the 2025 wave; showed the base model already contains most of the capability. | https://arxiv.org/abs/2501.19393 LIMO: Less is More for Reasoning | Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, Pengfei Liu | 2025 | arXiv | reasoning-distillation | Hand-curated SFT on 817 reasoning chains | 508 | 817 samples beat datasets 100x larger, with a 45.8-point absolute out-of-distribution gain. | Argues reasoning is elicited from pretraining, not taught by the traces — the strongest form of the 'less is more' claim. | https://arxiv.org/abs/2502.03387 Distillation Scaling Laws | Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, Russ Webb | 2025 | ICML | theory | Controlled scaling study fitting student loss to the teacher/student compute split | 58 | Students 143M-12.6B, teachers spanning a similar range, up to 512B tokens — the first predictive law for distillation. | Gives compute-optimal recipes and bounds where distillation beats supervised learning. | https://arxiv.org/abs/2502.08606 Quantification of Large Language Model Distillation | Sunbowen Lee, Junting Zhou, Chang Ao, Kaige Li, Xinrun Du, Sirui He, Haihong Wu, Tianci Liu, Jiaheng Liu, Hamid Alinejad-Rokny, Min Yang, Yitao Liang, Zhoufutu Wen, Shiwen Ni | 2025 | ACL | measurement | Response Similarity Evaluation + Identity Consistency Evaluation | 11 | Two metrics that estimate how heavily a given model was distilled from another. | The first quantitative handle on provenance and homogenisation in the open-model ecosystem. | https://arxiv.org/abs/2501.12619 Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning | Minwu Kim, Anubhav Shrestha, Safal Shrestha, Aadim Nepal, Keith Ross | 2025 | arXiv | theory | Controlled pass@1 vs pass@k analysis of RLVR and distillation | 30 | RLVR raises accuracy but rarely capability; distillation raises both when it injects new knowledge. | The clearest theoretical account of why distillation beat RL in the DeepSeek and Qwen ablations. | https://arxiv.org/abs/2505.14216 Qwen3 Technical Report | Qwen Team, Alibaba | 2025 | arXiv | on-policy | Strong-to-weak distillation: off-policy trace distillation then on-policy logit KL | 7650 | 1,800 GPU-hours of on-policy distillation beats 17,920 GPU-hours of RL by 6.8 AIME'24 points. | The first frontier-lab report to publish a like-for-like GPU-hour comparison of distillation against RL. | https://arxiv.org/html/2505.09388v1 OpenThoughts: Data Recipes for Reasoning Models | Etash Guha, Ryan Marten, Sedrick Keh, Negin Raoof, Georgios Smyrnis and the OpenThoughts team | 2025 | arXiv | reasoning-distillation | 1,000+ controlled ablations over the trace-generation pipeline | 208 | OpenThoughts3-1.2M yields a 7B student beating DeepSeek-R1-Distill-Qwen-7B by 15-20 points across three benchmarks. | Turned reasoning-trace curation from folklore into a measured, reproducible recipe. | https://arxiv.org/abs/2506.04178 Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs | Nicolas Boizard, Kevin El-Haddad, Céline Hudelot, Pierre Colombo | 2024 | TMLR | cross-tokenizer | Optimal-transport loss over logits from mismatched vocabularies | 60 | Distil between model families whose tokenizers do not align. | Removed the same-tokenizer constraint that limited white-box distillation to a single family. | https://arxiv.org/abs/2402.12030 Universal Cross-Tokenizer Distillation via Approximate Likelihood Matching | Benjamin Minixhofer, Ivan Vulić, Edoardo Maria Ponti | 2025 | NeurIPS | cross-tokenizer | Approximate likelihood matching across tokenizers | 41 | A principled cross-tokenizer objective that substantially outperforms prior heuristic methods. | Current strongest published result on the tokenizer-mismatch problem. | https://arxiv.org/abs/2503.20083 AI models collapse when trained on recursively generated data | Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, Yarin Gal | 2024 | Nature 631(8022):755-759 | risk | Recursive-training analysis for LLMs, VAEs and Gaussian mixtures | 989 | Training generation after generation on model output irreversibly erases the tails of the data distribution. | The strongest published caution against an ecosystem in which most training data is teacher output. | https://www.nature.com/articles/s41586-024-07566-y A Survey of On-Policy Distillation for Large Language Models | Mingyang Song, Mao Zheng | 2026 | arXiv | survey | Survey; formalises OPD as f-divergence minimisation over student-sampled trajectories | 109 | Organises the on-policy literature along optimisation target, signal source and training stabilisation. | The reference text for the technique that defines 2026 post-training; revised to v4 within three months. | https://arxiv.org/abs/2604.00626 Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective | Tokio Kajitsuka, Ukyo Honda, Sho Takase | 2026 | arXiv | theory | Corrected evaluation protocol for CoT distillation across teacher-student pairs | 3 | Under a fair protocol, CoT distillation sometimes makes students worse than their own baseline. | Brings the 2019 capacity-gap result into the reasoning era and gives practical pairing rules. | https://arxiv.org/abs/2604.08880 Benchmarking Distilled Language Models: Performance and Efficiency in Resource-Constrained Settings | Sachin Gopal Wani, Eric Page, Ajay Dholakia, David Ellison | 2026 | TPC Technology Conference | measurement | Controlled performance/efficiency benchmark of distilled vs vanilla models | 1 | A distilled 8B is over 2,000x more compute-efficient than training its vanilla counterpart, with reasoning on par with models 10x its size. | Quantifies the claim that distillation is now a primary training strategy rather than a compression afterthought. | https://arxiv.org/abs/2602.20164 Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning | NVIDIA | 2026 | arXiv | on-policy | Multi-teacher on-policy distillation (MOPD) from >10 domain specialists | 14 | Consolidate more than ten domain-specialised teachers into one 550B/55B-active student via dense token-level grading of student rollouts. | The most elaborate published distillation pipeline to date; MOPD replaces a multi-domain RL stage entirely. | https://arxiv.org/abs/2606.15007 DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence | DeepSeek-AI | 2026 | arXiv | on-policy | Independent domain experts merged into one student via on-policy distillation | 710 | V4-Pro (1.6T, 49B active) and V4-Flash (284B, 13B active) use distillation, not mixed RL, as the capability-merging step. | Confirms that at the frontier, distillation has moved from compression to model composition. | https://arxiv.org/abs/2606.19348 Contrastive Representation Distillation | Yonglong Tian, Dilip Krishnan, Phillip Isola | 2020 | ICLR | feature | Contrastive objective maximising a lower bound on mutual information between teacher and student representations | 1406 | Treat distillation as representation learning: maximise mutual information between teacher and student features instead of matching logits. | Established the contrastive family of feature distillation, and its benchmark suite is widely reused for KD comparisons. | https://arxiv.org/abs/1910.10699 TABLE: Student-vs-teacher benchmark retention [24 rows, from data/academic.json $.extras] Student | Teacher | Method | Student params (B) | Teacher params (B) | Benchmark | Teacher score (%) | Student score (%) | Retention (%) | Source URL ------- | ------- | ------ | ------------------ | ------------------ | --------- | ----------------- | ----------------- | ------------- | ---------- DeepSeek-R1-Distill-Qwen-1.5B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 1.5 | 671 | AIME 2024 pass@1 | 79.8 | 28.9 | 36.2 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-7B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 7 | 671 | AIME 2024 pass@1 | 79.8 | 55.5 | 69.5 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Llama-8B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 8 | 671 | AIME 2024 pass@1 | 79.8 | 50.4 | 63.2 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-14B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 14 | 671 | AIME 2024 pass@1 | 79.8 | 69.7 | 87.3 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-32B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 32 | 671 | AIME 2024 pass@1 | 79.8 | 72.6 | 91 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Llama-70B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 70 | 671 | AIME 2024 pass@1 | 79.8 | 70 | 87.7 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-1.5B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 1.5 | 671 | MATH-500 pass@1 | 97.3 | 83.9 | 86.2 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-7B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 7 | 671 | MATH-500 pass@1 | 97.3 | 92.8 | 95.4 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-32B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 32 | 671 | MATH-500 pass@1 | 97.3 | 94.3 | 96.9 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Llama-70B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 70 | 671 | MATH-500 pass@1 | 97.3 | 94.5 | 97.1 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-32B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 32 | 671 | GPQA Diamond pass@1 | 71.5 | 62.1 | 86.9 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Llama-70B | DeepSeek-R1 | Sequence-level SFT on 800k teacher traces (black-box) | 70 | 671 | GPQA Diamond pass@1 | 71.5 | 65.2 | 91.2 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B Sky-T1-32B-Preview | QwQ-32B-Preview | SFT on 17k rejection-sampled teacher traces | 32 | 32 | MATH-500 | 85.4 | 82.4 | 96.5 | https://novasky-ai.github.io/posts/sky-t1/ Sky-T1-32B-Preview | QwQ-32B-Preview | SFT on 17k rejection-sampled teacher traces | 32 | 32 | AIME 2024 | 50 | 43.3 | 86.6 | https://novasky-ai.github.io/posts/sky-t1/ Bespoke-Stratos-32B | DeepSeek-R1 | SFT on 17k traces via a modified Sky-T1 pipeline | 32 | 671 | AIME 2024 | 79.8 | 63.3 | 79.3 | https://huggingface.co/bespokelabs/Bespoke-Stratos-32B Bespoke-Stratos-32B | DeepSeek-R1 | SFT on 17k traces via a modified Sky-T1 pipeline | 32 | 671 | MATH-500 | 97.3 | 93 | 95.6 | https://huggingface.co/bespokelabs/Bespoke-Stratos-32B s1.1-32B | DeepSeek-R1 | SFT on 1,000 curated traces + budget forcing | 32 | 671 | MATH-500 | 97.3 | 95.4 | 98 | https://arxiv.org/html/2501.19393v3 Qwen3-8B (on-policy distilled) | Qwen3-32B | On-policy logit KL on student rollouts | 8 | 32 | AIME 2024 pass@1 | 81.4 | 74.4 | 91.4 | https://arxiv.org/html/2505.09388v1 Qwen3-8B (on-policy distilled) | Qwen3-32B | On-policy logit KL on student rollouts | 8 | 32 | AIME 2025 pass@1 | 72.9 | 65.5 | 89.8 | https://arxiv.org/html/2505.09388v1 MobileBERT | IB-BERT-LARGE (teacher); BERT-base scores shown as reference baseline | Bottleneck architecture + progressive knowledge transfer | 0.025 | 0.11 | GLUE score | 78.3 | 77.7 | 99.2 | https://arxiv.org/abs/2004.02984 MobileBERT | IB-BERT-LARGE (teacher); BERT-base scores shown as reference baseline | Bottleneck architecture + progressive knowledge transfer | 0.025 | 0.11 | SQuAD v1.1 F1 | 88.5 | 90 | 101.7 | https://arxiv.org/abs/2004.02984 MobileBERT | IB-BERT-LARGE (teacher); BERT-base scores shown as reference baseline | Bottleneck architecture + progressive knowledge transfer | 0.025 | 0.11 | SQuAD v2.0 F1 | 77.1 | 79.2 | 102.7 | https://arxiv.org/abs/2004.02984 DistilBERT | BERT-base | Pretraining-time logit KD + cosine-distance loss | 0.066 | 0.11 | GLUE (normalised: teacher = 100) | 100 | 97 | 97 | https://arxiv.org/abs/1910.01108 TinyBERT (4 layers) | BERT-base | Two-stage transformer distillation | 0.0145 | 0.11 | GLUE (normalised: teacher = 100) | 100 | 96.8 | 96.8 | https://arxiv.org/abs/1909.10351 METHODOLOGY NOTES paperCounts: arXiv counts were taken from the public arXiv API on 2026-09-04 using search_query=abs:"knowledge distillation" (or all:"" for the sub-topic chart) intersected with a per-year submittedDate range, reading totalResults. Phrase counting over-counts passing mentions and misses papers using other terminology; treat the shape of the curve, not the absolute level, as the finding. citationCounts: Citation counts come from the Semantic Scholar Graph API batch endpoint (fields=title,year,citationCount,venue), retrieved 2026-09-04. They differ from Google Scholar, typically by 10-30% downward. retention: retention_pct = student_score / teacher_score x 100, using each source's own reported figures. Where student and teacher were evaluated in different papers or harnesses, that is noted on the relevant table. omissions: Any figure a primary source did not publish is marked 'undisclosed' rather than estimated. Several 2025-2026 recipes (Bespoke-Stratos, LIMO, OpenThoughts, DeepSeek-R1-Distill) never published a dollar or GPU-hour training cost. sourcesRegister: $.sources[] is a curated register of the principal works, not a complete bibliography: roughly 18 further URLs are cited inline via per-row `_source` and per-chart `sources` fields (for example Stanton et al. arXiv:2106.05945 and Minixhofer et al. arXiv:2503.20083). Any bibliography rendered from $.sources[] alone will omit those; render inline `_source` values as well for completeness. TIMELINE OF THIS PERSPECTIVE (37 EVENTS) ---------------------------------------- 2006-08-20 | research | Buciluă, Caruana & Niculescu-Mizil: Model Compression 2015-03-09 | research | Hinton, Vinyals & Dean: Distilling the Knowledge in a Neural Network 2015-03-27 | research | FitNets: hints from intermediate layers 2016-06-25 | research | Kim & Rush: Sequence-Level Knowledge Distillation 2018-11-27 | research | Dataset Distillation 2019-10-02 | research | DistilBERT 2019-10-03 | research | Cho & Hariharan: On the Efficacy of Knowledge Distillation 2020-04-06 | research | MobileBERT 2021-06-01 | research | Knowledge Distillation: A Survey (IJCV) 2023-03-13 | research | Stanford Alpaca 2023-05-03 | research | Distilling Step-by-Step 2023-06-05 | research | Orca 2023-06-14 | research | MiniLLM: reverse-KL distillation 2023-06-23 | research | GKD: on-policy distillation of language models 2023-10-25 | research | Zephyr-7B: distilled DPO 2024-07-19 | research | Minitron: pruning + distillation 2024-07-24 | research | Shumailov et al., Nature: model collapse 2024-07-31 | research | Gemma 2 makes distillation a pretraining objective 2024-09-25 | product | Llama 3.2 1B/3B: one-shot pruning plus logit distillation 2025-01-10 | research | Sky-T1-32B-Preview trained for under $450 2025-01-22 | research | DeepSeek-R1 and the R1-Distill series 2025-01-22 | research | Quantification of LLM Distillation 2025-01-22 | research | Bespoke-Stratos-32B: 17k samples, 47x less data 2025-01-28 | research | Hugging Face launches Open-R1 2025-01-31 | research | s1: 1,000 samples, 26 minutes, 16 H100s 2025-02-05 | research | LIMO: 817 samples 2025-02-12 | research | Apple publishes Distillation Scaling Laws 2025-03-12 | research | Gemma 3 refines pretraining distillation 2025-05-14 | research | Qwen3 formalises strong-to-weak distillation 2025-06-04 | research | OpenThoughts3-1.2M and OpenThinker3-7B 2025-09-17 | research | DeepSeek-R1 published in Nature 2025-10-27 | research | Thinking Machines Lab popularises on-policy distillation 2026-04-01 | research | A Survey of On-Policy Distillation for Large Language Models 2026-04-10 | research | Capacity gap revisited for chain-of-thought distillation 2026-04-26 | research | DeepSeek-V4 replaces mixed RL with on-policy distillation 2026-06-12 | research | Nemotron 3 Ultra introduces multi-teacher on-policy distillation 2026-07-08 | research | A community survey reports distillation across the 2026 frontier SOURCES CITED BY THIS SECTION (48) ---------------------------------- [1] Distilling the Knowledge in a Neural Network — arXiv (Hinton, Vinyals, Dean), 2015-03-09 (paper) https://arxiv.org/abs/1503.02531 [2] Model compression — ACM KDD (Buciluă, Caruana, Niculescu-Mizil), 2006 (paper) https://dl.acm.org/doi/10.1145/1150402.1150464 [3] FitNets: Hints for Thin Deep Nets — arXiv / ICLR, 2015-03-27 (paper) https://arxiv.org/abs/1412.6550 [4] Sequence-Level Knowledge Distillation — arXiv / EMNLP, 2016-06-25 (paper) https://arxiv.org/abs/1606.07947 [5] DistilBERT, a distilled version of BERT — arXiv (Hugging Face), 2019-10-02 (paper) https://arxiv.org/abs/1910.01108 [6] TinyBERT: Distilling BERT for Natural Language Understanding — arXiv / Findings of EMNLP, 2019-09-23 (paper) https://arxiv.org/abs/1909.10351 [7] MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices — arXiv / ACL, 2020-04-06 (paper) https://arxiv.org/abs/2004.02984 [8] MiniLM: Deep Self-Attention Distillation — arXiv / NeurIPS, 2020-02-25 (paper) https://arxiv.org/abs/2002.10957 [9] On the Efficacy of Knowledge Distillation — arXiv / ICCV, 2019-10-03 (paper) https://arxiv.org/abs/1910.01348 [10] Knowledge Distillation: A Survey — arXiv / IJCV, 2020-06-09 (paper) https://arxiv.org/abs/2006.05525 [11] A Survey on Knowledge Distillation of Large Language Models — arXiv, 2024-02-20 (paper) https://arxiv.org/abs/2402.13116 [12] Distilling Step-by-Step! — arXiv / ACL (Google), 2023-05-03 (paper) https://arxiv.org/abs/2305.02301 [13] Orca: Progressive Learning from Complex Explanation Traces of GPT-4 — arXiv (Microsoft Research), 2023-06-05 (paper) https://arxiv.org/abs/2306.02707 [14] MiniLLM: Knowledge Distillation of Large Language Models — arXiv / ICLR 2024, 2023-06-14 (paper) https://arxiv.org/abs/2306.08543 [15] On-Policy Distillation of Language Models (GKD) — arXiv / ICLR 2024 (Google DeepMind), 2023-06-23 (paper) https://arxiv.org/abs/2306.13649 [16] Zephyr: Direct Distillation of LM Alignment — arXiv (Hugging Face), 2023-10-25 (paper) https://arxiv.org/abs/2310.16944 [17] Compact Language Models via Pruning and Knowledge Distillation (Minitron) — arXiv / NeurIPS (NVIDIA), 2024-07-19 (paper) https://arxiv.org/abs/2407.14679 [18] Gemma 2: Improving Open Language Models at a Practical Size — arXiv (Google DeepMind), 2024-07-31 (paper) https://arxiv.org/html/2408.00118v1 [19] Gemma 3 Technical Report — arXiv (Google DeepMind), 2025-03-25 (paper) https://arxiv.org/html/2503.19786v1 [20] Llama 3.2: Revolutionizing edge AI and vision — Meta AI, 2024-09-25 (blog) https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ [21] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — arXiv (DeepSeek-AI), 2025-01-22 (paper) https://arxiv.org/html/2501.12948v1 [22] DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning — Nature 645(8081), 2025-09-17 (paper) https://www.nature.com/articles/s41586-025-09422-z [23] DeepSeek-R1-Distill-Qwen-32B model card — Hugging Face (DeepSeek-AI), 2025-01-20 (docs) https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B [24] Sky-T1: Train your own O1 preview model within $450 — NovaSky, UC Berkeley Sky Computing Lab, 2025-01-10 (blog) https://novasky-ai.github.io/posts/sky-t1/ [25] Bespoke-Stratos: The unreasonable effectiveness of reasoning distillation — Bespoke Labs, 2025-01-22 (blog) https://www.bespokelabs.ai/blog/bespoke-stratos-the-unreasonable-effectiveness-of-reasoning-distillation [26] Bespoke-Stratos-32B model card — Hugging Face (Bespoke Labs), 2025-01-22 (docs) https://huggingface.co/bespokelabs/Bespoke-Stratos-32B [27] s1: Simple test-time scaling — arXiv / EMNLP 2025 (Stanford, UW), 2025-01-31 (paper) https://arxiv.org/abs/2501.19393 [28] LIMO: Less is More for Reasoning — arXiv (SJTU / GAIR), 2025-02-05 (paper) https://arxiv.org/abs/2502.03387 [29] Distillation Scaling Laws — arXiv / ICML 2025 (Apple), 2025-02-12 (paper) https://arxiv.org/abs/2502.08606 [30] Distillation Scaling Laws (Apple ML Research page) — Apple Machine Learning Research, 2025-02-12 (blog) https://machinelearning.apple.com/research/distillation-scaling-laws [31] Qwen3 Technical Report — arXiv (Qwen Team, Alibaba), 2025-05-14 (paper) https://arxiv.org/html/2505.09388v1 [32] OpenThoughts: Data Recipes for Reasoning Models — arXiv (OpenThoughts consortium), 2025-06-04 (paper) https://arxiv.org/abs/2506.04178 [33] Open-R1: a fully open reproduction of DeepSeek-R1 — Hugging Face, 2025-01-28 (blog) https://huggingface.co/blog/open-r1 [34] On-Policy Distillation — Thinking Machines Lab, 2025-10-27 (blog) https://thinkingmachines.ai/blog/on-policy-distillation/ [35] A Survey of On-Policy Distillation for Large Language Models — arXiv (Song & Zheng), 2026-04-01 (paper) https://arxiv.org/abs/2604.00626 [36] Revisiting the Capacity Gap in Chain-of-Thought Distillation from a Practical Perspective — arXiv (Kajitsuka, Honda, Takase), 2026-04-10 (paper) https://arxiv.org/abs/2604.08880 [37] DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence — arXiv (DeepSeek-AI), 2026-04-26 (paper) https://arxiv.org/abs/2606.19348 [38] Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning — arXiv (NVIDIA), 2026-06-12 (paper) https://arxiv.org/abs/2606.15007 [39] Distillation in 2026 (so far): which frontier models use it and how — Hugging Face, 2026-07-08 (blog) https://huggingface.co/blog/sergiopaniego/distillation-2026 [40] Benchmarking Distilled Language Models: Performance and Efficiency in Resource-Constrained Settings — arXiv (Wani, Page, Dholakia, Ellison), 2026-01-28 (paper) https://arxiv.org/abs/2602.20164 [41] Quantification of Large Language Model Distillation — arXiv / ACL 2025, 2025-01-22 (paper) https://arxiv.org/abs/2501.12619 [42] AI models collapse when trained on recursively generated data — Nature 631(8022):755-759, 2024-07-24 (paper) https://www.nature.com/articles/s41586-024-07566-y [43] Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning — arXiv (Kim et al., NYU), 2025-05-20 (paper) https://arxiv.org/abs/2505.14216 [44] Alpaca: A Strong, Replicable Instruction-Following Model — Stanford CRFM, 2023-03-13 (blog) https://crfm.stanford.edu/2023/03/13/alpaca.html [45] Semantic Scholar Graph API (citation counts) — Allen Institute for AI, 2026-09-04 (docs) https://api.semanticscholar.org/graph/v1/paper/arXiv:1503.02531?fields=title,year,citationCount [46] arXiv API user manual (paper-count methodology) — arXiv / Cornell University, 2026-09-04 (docs) https://info.arxiv.org/help/api/user-manual.html [47] Open-R1 repository (evaluation tables) — Hugging Face, 2025-01-28 (docs) https://github.com/huggingface/open-r1 [48] Contrastive Representation Distillation — arXiv (MIT / Google Research), 2019-10-23 (paper) https://arxiv.org/abs/1910.10699 ================================================================================================ 4. FINANCIAL PERSPECTIVE — THE ECONOMICS OF AI DISTILLATION: PRICES, TRAINING COSTS, AND THE MARKET SHOCKS ================================================================================================ Page: https://global-distillation.com/financial Data: https://global-distillation.com/data/financial.json Updated: 2026-09-03 Content: 8 key figures, 8 tables, 7 charts, 30 dated events, 16 glossary terms, 63 sources SUMMARY -------- Distillation is, at bottom, an arbitrage: the capability embedded in a $40M-$500M frontier training run can be harvested through an API for a four- or five-figure query bill and re-trained into a small model for hundreds of dollars. That asymmetry is now visible in published price lists, where the gap between a vendor's flagship and its small tier runs 10x to 42x on output tokens, and in the research record, where Sky-T1-32B was distilled for under $450 and s1-32B for a reported ~$50. It became a macro event on 27 January 2025, when DeepSeek-R1 and its six open-weight distilled students wiped $589B off Nvidia's market capitalisation in a single session - the largest one-day loss in stock-market history - and it became a policy event in February 2026 when OpenAI told the House Select Committee on China that DeepSeek was running obfuscated distillation pipelines against its models. The counter-trend matters too: Epoch AI measures inference prices for a fixed capability level falling 9x-900x per year, which compresses the payback on self-hosting a distilled model to the point where, on our own TCO model, a single-GPU deployment only beats Claude Haiku 4.5 above ~658M output tokens a month and never beats the cheapest serverless open-model endpoints. And in August 2026 DeepSeek reversed the race to zero, raising V4 API prices by as much as 1,100% and introducing peak/off-peak rates - the first major signal that ultra-cheap distilled inference was being priced against capacity, not against marginal cost. KEY FIGURES ----------- - Nvidia single-day market-cap loss: 589 USD billions (-17% in one session) 27 Jan 2025, after DeepSeek-R1 and its distilled students shipped. Largest one-day loss in US market history. Source: https://www.cnbc.com/2025/01/27/nvidia-sheds-almost-600-billion-in-market-cap-biggest-drop-ever.html - Cost to train Sky-T1-32B by distillation: 450 USD (8x H100 for 19 hours) Matches o1-preview on Math500 and AIME24; teacher was QwQ-32B-Preview, base was Qwen2.5-32B-Instruct. Source: https://novasky-ai.github.io/posts/sky-t1/ - DeepSeek-R1 reinforcement-learning training cost: 294,000 USD (512 H800s x 80 hours) Disclosed in the peer-reviewed Nature paper, Sept 2025. Excludes the ~$5.6M base-model run and all R&D, data and infrastructure. Source: https://www.cnn.com/2025/09/19/business/deepseek-ai-training-cost-china-intl - Frontier-vs-small output price spread, OpenAI: 42 x ($50 vs $1.20 per MTok) gpt-6-astra output $50/MTok against gpt-5.6-luna at $1.20/MTok on the same list. Source: https://developers.openai.com/api/docs/pricing - Inference price decline for fixed capability: 900 x per year (upper bound) (range 9x-900x) Epoch AI, measuring the cheapest model clearing a fixed benchmark threshold. GPT-4-level fell from $37.50/MTok (Mar 2023) to $0.18/MTok (Feb 2025). Source: https://epoch.ai/data-insights/llm-inference-price-trends - Frontier training-run cost growth: 2.4 x per year (95% CI 2.0x-3.1x, since 2016) Epoch AI, amortised hardware + energy across 45 frontier models; cloud-rental method gives 2.6x/yr. Source: https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models - Teacher-query bill to rebuild an 800k-trace reasoning corpus: 40,000 USD (upper end) ($1,920 at the cheap end) Author calculation: 800k traces x 2,000 output tokens = 1.6B output tokens, priced at 2026 list rates from Claude Opus 5 ($25/MTok) down to gpt-5.6-luna ($1.20/MTok). Source: https://platform.claude.com/docs/en/about-claude/pricing - DeepSeek V4 API price increase: 1,100 % (maximum) (effective 16 Aug 2026) Cache-hit input tokens rose up to 1,100%; output tokens 127%-371%. Peak/off-peak schedule introduced for the first time. Source: https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html KEY FINDINGS ------------ 1. The distillation arbitrage is a five-order-of-magnitude gap, and it is widening Epoch AI puts the final training run of GPT-4 at roughly $40M on an amortised-hardware basis (the Stanford AI Index puts it at $78M on cloud-rental accounting), and frontier run costs have grown 2.4x per year since 2016. Against that, Sky-T1-32B was distilled for under $450 of GPU time and TinyZero reproduced R1-Zero-style behaviour for under $30. Even the teacher-query cost is small: at 2026 list prices, regenerating an 800k-trace reasoning corpus like the one behind DeepSeek's R1-Distill family costs $1,920 (gpt-5.6-luna) to $40,000 (Claude Opus 5). The ratio between building the capability and copying it is roughly 1,000:1 to 100,000:1. Sources: https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models https://novasky-ai.github.io/posts/sky-t1/ https://developers.openai.com/api/docs/pricing 2. Every major vendor now sells its own distillation discount as a product tier The frontier-to-small spread on output tokens is 10x at Anthropic (Claude Fable 5.1 at $50/MTok vs Haiku 4.5 at $5/MTok), 42x at OpenAI (gpt-6-astra $50 vs gpt-5.6-luna $1.20), 30x at Google (Gemini 3.1 Pro $12 vs Gemini 2.5 Flash-Lite $0.40) and 75x at Mistral (Medium 3.5 at $7.50 vs Ministral 3 3B at $0.10). Vendors capture the distillation margin internally rather than losing it to third parties. The commercial logic is that a customer who would otherwise self-distill can be retained at a price point the vendor still profits at. Sources: https://platform.claude.com/docs/en/about-claude/pricing https://developers.openai.com/api/docs/pricing https://ai.google.dev/gemini-api/docs/pricing https://mistral.ai/pricing/api 3. Self-hosting a distilled model is now the expensive option for most buyers On our TCO model (1x H100 SXM at $2.40/GPU-hour on demand, 24/7, plus 0.25 FTE of MLOps at $200k/yr fully loaded = $5,919/month), a self-hosted distilled model only undercuts Claude Opus 5 above ~132M output tokens a month, Claude Sonnet 5 above ~329M, and Claude Haiku 4.5 above ~658M. It never undercuts gpt-5.6-luna or DeepSeek V4-Flash within the throughput capacity of a single GPU. Serverless open-model endpoints have eaten the economic case for owning inference except at genuine scale or where data residency forces it. Sources: https://www.gmicloud.ai/en/blog/nvidia-h100-gpu-pricing-2026-rent-vs-buy-cost-analysis https://platform.claude.com/docs/en/about-claude/pricing https://api-docs.deepseek.com/quick_start/pricing/ 4. The January 2025 shock repriced compute, not software DeepSeek-R1 shipped on 20 January 2025 with six open-weight distilled students under MIT licence. Seven days later Nvidia fell 17% and lost $589B of market capitalisation, the Nasdaq 100 fell 3%, and the semiconductor index had its worst day since March 2020. The market read distillation as a claim that frontier capability could be reproduced without frontier capital expenditure. It was wrong on the timescale: Nvidia crossed $5T in October 2025 - the first company ever to do so - and stood at roughly $5.43T on 2 September 2026. Sources: https://www.cnbc.com/2025/01/27/nvidia-sheds-almost-600-billion-in-market-cap-biggest-drop-ever.html https://www.thenationalnews.com/future/technology/2026/04/25/will-deepseeks-new-ai-model-crash-nvidias-5tn-party/ https://stockanalysis.com/stocks/nvda/market-cap/ 5. DeepSeek's $5.6M figure was a marginal-cost number, and the dispute is about accounting The DeepSeek-V3 technical report states 2.788M H800 GPU-hours for full training; the widely circulated $5.576M is that figure multiplied by an assumed $2/GPU-hour rental rate. SemiAnalysis countered on 31 January 2025 that DeepSeek's server capex is around $1.6B across roughly 50,000 Hopper GPUs with about $944M of operating cost, and that the published number excludes R&D, data, failed runs and hardware total cost of ownership. The Nature paper on R1 (Sept 2025) is narrower still: $294,000 covers only the reinforcement-learning stage on top of an already-trained base. Sources: https://arxiv.org/abs/2412.19437 https://semianalysis.com/2025/01/31/deepseek-debates/ https://www.theregister.com/2025/09/19/deepseek_cost_train/ 6. The race to zero reversed in August 2026 DeepSeek warned on 6 August 2026 of a significant price increase and implemented it on 16 August: V4-Flash output went from $0.28/MTok to $0.66 off-peak and $1.32 at peak; V4-Pro output from $0.87 to $1.98/$3.96. Cache-hit input tokens rose by as much as 1,100%. Seventeen of twenty-four hours remain at the half-price off-peak rate, and peak hours are set on Beijing business time, so the increase falls hardest on domestic users and lightest on Western buyers. The signal is that ultra-cheap distilled inference was capacity-constrained, not structurally free. Sources: https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html https://api-docs.deepseek.com/quick_start/pricing/ https://www.caixinglobal.com/2026-08-14/tech-brief-aug-14-deepseek-launches-v4-pro-and-raises-api-prices-by-as-much-as-1100-102474222.html 7. Cloud vendors monetise distillation through fine-tuning and provisioned throughput, not through the distillation itself OpenAI's Model Distillation ships Stored Completions free and charges standard fine-tuning rates ($25/MTok training for gpt-4.1, $1.50/MTok for gpt-4.1-nano). Amazon Bedrock Model Distillation (GA 1 May 2025) charges for the teacher inference calls used to synthesise data, then bills the resulting custom model at $1.95/month storage plus Provisioned Throughput - there is no on-demand tier for a distilled model on Bedrock at any volume. AWS markets distilled models as up to 500% faster and up to 75% cheaper to run with under 2% accuracy loss on RAG. The workflow is free; the lock-in is in where the student runs. Sources: https://developers.openai.com/api/docs/pricing https://docs.aws.amazon.com/bedrock/latest/userguide/model-distillation.html https://press.aboutamazon.com/2024/12/aws-strengthens-amazon-bedrock-with-industry-first-ai-safeguard-new-agent-capability-and-model-customization 8. Distillation allegations moved from a commercial dispute to a congressional one Microsoft security researchers observed large-scale data exfiltration through OpenAI developer accounts they linked to DeepSeek in late 2024; the probe became public on 29 January 2025. The House Select Committee on the CCP concluded in its April 2025 report that it is highly likely DeepSeek used unlawful model distillation techniques. In February 2026 OpenAI submitted a memo to the same committee describing sophisticated, multi-stage distillation pipelines using obfuscated third-party routers to conceal origin. The financial stake is that a distillation attack converts a multi-hundred-million-dollar capital asset into a commodity a competitor can rent. Sources: https://www.bloomberg.com/news/articles/2025-01-29/microsoft-probing-if-deepseek-linked-group-improperly-obtained-openai-data https://www.techpolicy.press/us-house-select-committee-report-accuses-deepseek-of-spying-and-circumventing-export-controls-on-chips/ https://cdn.openai.com/pdf/045aa967-ee96-4a09-94ee-3098ddf6db2c/OpenAI-US-House-Select-Cmte-Update-%5B021226%5D.pdf 9. Model extraction is cheap enough to be an operating expense, not a capital project Carlini et al. recovered the exact hidden dimension of gpt-3.5-turbo and estimated the full embedding-projection matrix could be extracted for under $2,000 in API queries; a limited version of the attack cost under $200, and ada and babbage were fully extracted for under $20. Combined with the corpus-generation figures above, the total cash cost of a serious behavioural-cloning effort against a frontier model sits in the $10^3-$10^5 range against a $10^8 asset. No defensive spend scales down to that. Sources: https://arxiv.org/pdf/2403.06634 https://proceedings.mlr.press/v235/carlini24a.html 10. Capital followed the small-model thesis, but the exits were modest Arcee AI raised a $24M Series A led by Emergence Capital for domain-specific small language models. Predibase, which sold fine-tuning tooling for small open models, raised over $28M and was acquired by Rubrik in June 2025 for a reported $100M-$500M. Together AI, the largest pure-play open-model inference platform, raised $305M at $3.3B in February 2025 and $800M at $8.3B in July 2026. Mistral raised a EUR 1.7B Series C at EUR 11.7B in September 2025 with ASML taking 11%. The value accrued to inference capacity and to European sovereignty plays, not to distillation tooling as a standalone category. Sources: https://www.arcee.ai/blog/arcee-ai-secures-24m-series-a-to-transform-the-landscape-of-small-language-models https://techcrunch.com/2025/06/25/rubrik-acquires-predibase-to-accelerate-adoption-of-ai-agents/ https://techcrunch.com/2026/07/01/neocloud-together-ai-raises-800m-leaps-to-8-3b-valuation/ https://mistral.ai/news/mistral-ai-raises-1-7-b-to-accelerate-technological-progress-with-ai/ TABLES -------- TABLE: Frontier vs small-tier list prices, September 2026 [id: price-matrix-frontier-vs-small, 11 rows] Each vendor's flagship compared with its cheapest general-purpose tier, on the same published price list. Prices are USD per million tokens, standard (non-batch, non-cached) rates. Vendor | Frontier model | Frontier in (USD/MTok) | Frontier out (USD/MTok) | Small / distilled tier | Small in (USD/MTok) | Small out (USD/MTok) | Output price ratio (x) | Source URL ------ | -------------- | ---------------------- | ----------------------- | ---------------------- | ------------------- | -------------------- | ---------------------- | ---------- OpenAI | gpt-6-astra | 10 | 50 | gpt-5.6-luna | 0.2 | 1.2 | 41.7 | https://developers.openai.com/api/docs/pricing OpenAI | gpt-5.6-sol | 4 | 20 | gpt-5.4-nano | 0.2 | 1.25 | 16 | https://developers.openai.com/api/docs/pricing Anthropic | Claude Fable 5.1 | 10 | 50 | Claude Haiku 4.5 | 1 | 5 | 10 | https://platform.claude.com/docs/en/about-claude/pricing Anthropic | Claude Opus 5 | 5 | 25 | Claude Haiku 4.5 | 1 | 5 | 5 | https://platform.claude.com/docs/en/about-claude/pricing Google | Gemini 3.1 Pro Preview | 2 | 12 | Gemini 2.5 Flash-Lite | 0.1 | 0.4 | 30 | https://ai.google.dev/gemini-api/docs/pricing xAI | grok-4.6 (<200k) | 2 | 6 | grok-build-0.1 | 1 | 2 | 3 | https://docs.x.ai/docs/models DeepSeek | deepseek-v4-pro (peak) | 1.32 | 3.96 | deepseek-v4-flash (off-peak) | 0.22 | 0.66 | 6 | https://api-docs.deepseek.com/quick_start/pricing/ Mistral | Mistral Medium 3.5 | 1.5 | 7.5 | Ministral 3 (3B) | 0.1 | 0.1 | 75 | https://mistral.ai/pricing/api Alibaba (Qwen) | qwen3.8-max | 2 | 6 | qwen-turbo | 0.05 | 0.2 | 30 | https://www.alibabacloud.com/help/en/model-studio/model-pricing Meta (via Together AI) | Llama 3.3 70B | 1.04 | 1.04 | Llama 3 8B Instruct Lite | 0.14 | 0.14 | 7.4 | https://www.together.ai/pricing OpenAI open weights (via Groq) | gpt-oss-120b | 0.15 | 0.6 | gpt-oss-20b | 0.075 | 0.3 | 2 | https://console.groq.com/docs/models Notes: Ratios are output-token price ratios, computed from the listed figures. DeepSeek peak hours are 01:00-04:00 and 06:00-10:00 UTC Monday-Friday; the other 17 hours are half price. xAI and Anthropic rates shown are for prompts below the long-context threshold. Sources: https://developers.openai.com/api/docs/pricing https://platform.claude.com/docs/en/about-claude/pricing https://ai.google.dev/gemini-api/docs/pricing https://docs.x.ai/docs/models https://api-docs.deepseek.com/quick_start/pricing/ https://mistral.ai/pricing/api https://www.alibabacloud.com/help/en/model-studio/model-pricing https://www.together.ai/pricing https://console.groq.com/docs/models TABLE: What it costs to build a model vs what it costs to copy one [id: training-cost-ladder, 11 rows] Disclosed and estimated training costs, ordered from frontier pre-training down to hobbyist distillation. Note that these figures are not accounted on a common basis - read the note column. Model / run | Organisation | Reported cost (USD) | Compute | Date | What the number covers | Source URL ----------- | ------------ | ------------------- | ------- | ---- | ---------------------- | ---------- Gemini Ultra 1.0 | Google DeepMind | 191000000 | TPU v4, ~35 MW | 2023-12 | Stanford AI Index cloud-rental accounting; Epoch's amortised-hardware method gives ~$30M | https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models GPT-4 | OpenAI | 78000000 | A100 cluster | 2023-03 | Stanford AI Index cloud-rental accounting; Epoch amortised-hardware estimate ~$40M | https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models GPT-4 (amortised) | OpenAI | 40000000 | A100 cluster | 2023-03 | Epoch AI: depreciated hardware + energy for the final run only | https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models DeepSeek infrastructure (all-in estimate) | DeepSeek / High-Flyer | 1600000000 | ~50,000 Hopper GPUs | 2025-01 | SemiAnalysis estimate of total server capex; ~$944M of operating cost on top | https://semianalysis.com/2025/01/31/deepseek-debates/ DeepSeek-V3 (final run, claimed) | DeepSeek | 5576000 | 2.788M H800 GPU-hours | 2024-12 | GPU-hours are in the technical report; the dollar figure applies an assumed $2/GPU-hour rental rate | https://arxiv.org/abs/2412.19437 DeepSeek-R1 (RL stage) | DeepSeek | 294000 | 512 H800s x 80 hours (~41k GPU-hours) | 2025-09 | Disclosed in the Nature paper; excludes the V3 base model, data, energy, infrastructure and staff | https://www.cnn.com/2025/09/19/business/deepseek-ai-training-cost-china-intl DeepSeek-R1-Distill-Qwen-32B | DeepSeek | n/a | SFT on 800k R1-generated samples | 2025-01 | Cost undisclosed; the 800k-sample corpus is the expensive input, not the fine-tune | https://arxiv.org/html/2501.12948v1 Bespoke-Stratos-32B | Bespoke Labs | n/a | 8x H100 for 27 hours (216 GPU-hours) | 2025-01 | Cost undisclosed. 17k traces distilled from DeepSeek-R1 in 1.5 hours of teacher inference; 47x fewer examples than R1-Distill-Qwen-32B | https://www.bespokelabs.ai/blog/bespoke-stratos-the-unreasonable-effectiveness-of-reasoning-distillation Sky-T1-32B-Preview | NovaSky, UC Berkeley | 450 | 8x H100 for 19 hours (152 GPU-hours) | 2025-01 | Full disclosed cost of the fine-tune; 17k traces from QwQ-32B-Preview reformatted with GPT-4o-mini | https://novasky-ai.github.io/posts/sky-t1/ s1-32B | Stanford / University of Washington | 50 | 16x H100 for 26 minutes (~6.9 GPU-hours) | 2025-02 | GPU time is stated in the paper; the ~$50 figure comes from press coverage, not the paper itself. 1,000 traces distilled from Gemini Thinking Experimental | https://arxiv.org/html/2501.19393v2 TinyZero | UC Berkeley (Jiayi Pan et al.) | 30 | 3B Qwen base, RL on Countdown task | 2025-01 | Server cost for the experiments only; reproduces R1-Zero-style self-verification on a narrow task | https://www.dailycal.org/news/campus/research-and-ideas/campus-researchers-replicate-disruptive-chinese-ai-for-30/article_a1cc5cd0-dee4-11ef-b8ca-171526dfb895.html Notes: The costs in this table are NOT comparable on a like-for-like basis. Frontier rows are whole pre-training runs; distillation rows are fine-tunes that assume a free base model and a paid-for teacher. The honest comparison is the ratio between the top of the table and the bottom, which is roughly 10^5 to 10^6. Sources: https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models https://semianalysis.com/2025/01/31/deepseek-debates/ https://arxiv.org/abs/2412.19437 https://novasky-ai.github.io/posts/sky-t1/ https://arxiv.org/html/2501.19393v2 TABLE: What it costs to buy a teacher's reasoning traces at 2026 list prices [id: distillation-corpus-cost, 7 rows] Author calculation. Assumes an average reasoning trace of 2,000 output tokens and ignores input-token cost (prompts are short relative to reasoning traces). Two corpus sizes: the 17k traces used by Sky-T1 and Bespoke-Stratos, and the 800k samples DeepSeek used for its R1-Distill family. Teacher model | Output price (USD/MTok) | 17k traces (34M out-tok) (USD) | 800k traces (1.6B out-tok) (USD) | Source URL ------------- | ----------------------- | ------------------------ | ------------------------ | ---------- Claude Opus 5 | 25 | 850 | 40000 | https://platform.claude.com/docs/en/about-claude/pricing gpt-5.6-sol | 20 | 680 | 32000 | https://developers.openai.com/api/docs/pricing Gemini 3.5 Flash | 9 | 306 | 14400 | https://ai.google.dev/gemini-api/docs/pricing grok-4.6 | 6 | 204 | 9600 | https://docs.x.ai/docs/models deepseek-v4-pro (off-peak) | 1.98 | 67 | 3168 | https://api-docs.deepseek.com/quick_start/pricing/ gpt-5.6-luna | 1.2 | 41 | 1920 | https://developers.openai.com/api/docs/pricing gpt-oss-120b (via Groq) | 0.6 | 20 | 960 | https://console.groq.com/docs/models Notes: Applying the Batch API discount (50% at Anthropic and OpenAI) halves every figure again. These are list prices for legitimate API use; a distillation programme that violates terms of service would face the same compute bill plus the cost of evading detection, which OpenAI's February 2026 memo to Congress describes as obfuscated third-party routers. Sources: https://platform.claude.com/docs/en/about-claude/pricing https://developers.openai.com/api/docs/pricing https://api-docs.deepseek.com/quick_start/pricing/ TABLE: TCO: teacher API vs distilled API vs self-hosted, by monthly volume [id: tco-scenarios, 7 rows] Author-computed model. Assumptions: 4 input tokens per output token (typical RAG-style workload); monthly cost = output_MTok x (4 x input_price + output_price). Self-hosted row = 1x H100 SXM rented on demand at $2.40/GPU-hour for 730 hours ($1,752/mo) plus 0.25 FTE of MLOps at $200,000/yr fully loaded ($4,167/mo) = $5,919/mo flat, with capacity of ~1,051M output tokens/month at a sustained 400 output tok/s. Output tokens/month (millions) | Claude Opus 5 ($5/$25) (USD/mo) | Claude Sonnet 5 ($2/$10) (USD/mo) | Claude Haiku 4.5 ($1/$5) (USD/mo) | gpt-5.6-luna ($0.20/$1.20) (USD/mo) | DeepSeek V4-Flash off-peak ($0.22/$0.66) (USD/mo) | Self-hosted distilled 8B, 1x H100 (USD/mo) | Source URL ------------------------ | ------------------------ | ------------------------ | ------------------------ | ------------------------ | ------------------------ | ------------------------ | ---------- 1 | 45 | 18 | 9 | 2 | 2 | 5919 | https://platform.claude.com/docs/en/about-claude/pricing 10 | 450 | 180 | 90 | 20 | 15 | 5919 | https://platform.claude.com/docs/en/about-claude/pricing 50 | 2250 | 900 | 450 | 100 | 77 | 5919 | https://platform.claude.com/docs/en/about-claude/pricing 100 | 4500 | 1800 | 900 | 200 | 154 | 5919 | https://platform.claude.com/docs/en/about-claude/pricing 200 | 9000 | 3600 | 1800 | 400 | 308 | 5919 | https://platform.claude.com/docs/en/about-claude/pricing 500 | 22500 | 9000 | 4500 | 1000 | 770 | 5919 | https://platform.claude.com/docs/en/about-claude/pricing 1000 | 45000 | 18000 | 9000 | 2000 | 1540 | 5919 | https://platform.claude.com/docs/en/about-claude/pricing Notes: Breakevens for the self-hosted option: 132M output tokens/month vs Opus 5, 329M vs Sonnet 5, 658M vs Haiku 4.5, 2,959M vs gpt-5.6-luna and 3,843M vs DeepSeek V4-Flash off-peak. The last two exceed single-GPU capacity, so at these prices a one-GPU deployment never pays back against the cheapest hosted endpoints. Excludes the one-off distillation cost ($450-$40,000, see other tables), redundancy, and the fact that measured self-hosted cost per output MTok on identical H100 hardware spans $0.21 to $15.25 depending purely on request concurrency. Sources: https://platform.claude.com/docs/en/about-claude/pricing https://developers.openai.com/api/docs/pricing https://api-docs.deepseek.com/quick_start/pricing/ https://www.gmicloud.ai/en/blog/nvidia-h100-gpu-pricing-2026-rent-vs-buy-cost-analysis https://arxiv.org/html/2606.11690v1 TABLE: Cloud vendor distillation products and how they charge [id: vendor-distillation-products, 7 rows] Managed distillation workflows offered by the major platforms, with the actual revenue mechanism. Vendor | Product | Launched | How it charges | Stated benefit | Source URL ------ | ------- | -------- | -------------- | -------------- | ---------- OpenAI | Model Distillation (Stored Completions + Evals) | 2024-10 | Stored Completions free; standard fine-tuning rates apply ($25/MTok training for gpt-4.1, $5 for gpt-4.1-mini, $1.50 for gpt-4.1-nano, $100/hr for o4-mini) | Train smaller cost-efficient models on frontier outputs for a specific task | https://developers.openai.com/api/docs/pricing Amazon Web Services | Amazon Bedrock Model Distillation | 2024-12 preview, GA 2025-05-01 | Teacher inference charged at on-demand rates when Bedrock synthesises data; custom model storage $1.95/model/month; inference only via Provisioned Throughput | Up to 500% faster and up to 75% less expensive to run, with under 2% accuracy loss on RAG | https://docs.aws.amazon.com/bedrock/latest/userguide/model-distillation.html Google Cloud | Vertex AI supervised tuning / distillation | 2024 | Per training token; tuned model endpoints billed at 1.5x the base model rate | Task-specific tuned Gemini Flash and Flash-Lite students | https://cloud.google.com/vertex-ai/pricing Microsoft | Azure OpenAI / Microsoft Foundry stored completions and distillation | 2024-10 | Standard Azure fine-tuning and inference rates; distillation requires a minimum of 10 stored completions | Turn production traffic against a large model into a fine-tuning set for a small one | https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/stored-completions?view=foundry-classic Anthropic | No first-party distillation product | n/a | Tiered model line (Fable/Opus/Sonnet/Haiku) plus Batch API 50% discount and prompt caching down to 0.025x input price | Vendor captures the cost-reduction margin internally rather than selling a distillation pipeline | https://platform.claude.com/docs/en/about-claude/pricing Together AI | Serverless open-model inference and fine-tuning | 2023 | Per token for serverless; dedicated H100 at $3.99/hr promotional ($5.49 regular), B200 at $8.99/hr | Host the distilled student without owning hardware | https://www.together.ai/pricing Fireworks AI | Serverless and on-demand GPU deployments | 2023 | H100 80GB and H200 141GB at $8.00/hr from 1 Sep 2026 (previously $7.00); B200 $13.00/hr; GB300 $20.00/hr; 1.5x premium for region-restricted deployments | Dedicated capacity for custom and distilled models | https://fireworks.ai/pricing Notes: Bedrock's constraint is the commercially interesting one: a distilled model cannot be served on demand, so the customer trades a per-token bill for an hourly Provisioned Throughput commitment, which reverses the economics for low-volume users. Sources: https://developers.openai.com/api/docs/pricing https://docs.aws.amazon.com/bedrock/latest/userguide/model-distillation.html https://aws.amazon.com/bedrock/pricing/ https://www.together.ai/pricing https://fireworks.ai/pricing TABLE: Price of a fixed capability level over time [id: inference-price-decline, 5 rows] Measured price of the cheapest model clearing a fixed benchmark threshold, from Epoch AI and a16z. This is the deflation that distillation both causes and competes with. Capability level | First available | Price then (USD/MTok) | Cheapest as of | Price then (USD/MTok) | Decline (x) | Source URL ---------------- | --------------- | --------------------- | -------------- | --------------------- | ----------- | ---------- GPT-3 level (MMLU 42) | 2021-11 | 60 | 2024-11 | 0.06 | 1000 | https://a16z.com/llmflation-llm-inference-cost/ GPT-3.5 level (MMLU >= 64.8) | 2022-11 | 20 | 2024-10 | 0.07 | 286 | https://epoch.ai/data-insights/llm-inference-price-trends GPT-4 level (MMLU >= 86.0) | 2023-03 | 37.5 | 2025-02 | 0.18 | 208 | https://epoch.ai/data-insights/llm-inference-price-trends PhD-level science (GPQA Diamond >= 50) | 2023-11 | 15 | 2024-12 | 0.12 | 125 | https://epoch.ai/data-insights/llm-inference-price-trends MMLU 83 (GPT-4 launch level) | 2023-03 | 30 | 2024-11 | 0.48 | 62 | https://a16z.com/llmflation-llm-inference-cost/ Notes: Epoch's headline range across six benchmarks is a 9x-900x annual decline, with a median around 50x and roughly 200x for models released since 2024. a16z's 'LLMflation' framing is a 10x decline per year and 1,000x over three years. The MMLU 83 start price is a16z's stated ~62x reduction applied to the GPT-4 launch price; treat it as derived rather than directly quoted. Sources: https://epoch.ai/data-insights/llm-inference-price-trends https://a16z.com/llmflation-llm-inference-cost/ TABLE: Capital raised against the small-model / distillation thesis [id: vc-funding-small-models, 8 rows] Funding rounds and exits for companies whose pitch is small, distilled or open models. Company | Event | Amount (USD millions) | Valuation | Date | Thesis | Source URL ------- | ----- | --------------------- | --------- | ---- | ------ | ---------- Arcee AI | Seed | 5.5 | undisclosed | 2023-12 | Domain-specific small language models | https://www.arcee.ai/blog/arcee-ai-secures-24m-series-a-to-transform-the-landscape-of-small-language-models Arcee AI | Series A (Emergence Capital) | 24 | undisclosed | 2024-07 | Model merging and Spectrum training to cut SLM training cost | https://venturebeat.com/ai/small-language-models-rising-as-arcee-ai-lands-24m-series-a Predibase | Total VC raised before exit | 28 | undisclosed | 2024 | Fine-tuning tooling for small open models (Llama, Mistral) | https://www.techtarget.com/searchdatabackup/news/366626870/Rubrik-pivots-to-generative-AI-with-Predibase-acquisition Predibase | Acquired by Rubrik | 300 | reported $100M-$500M range | 2025-06-25 | Agentic AI needs cheap task-specific models | https://techcrunch.com/2025/06/25/rubrik-acquires-predibase-to-accelerate-adoption-of-ai-agents/ Together AI | Series B (General Catalyst, Prosperity7) | 305 | $3.3B | 2025-02-20 | End-to-end platform for building with 200+ open-source models | https://siliconangle.com/2025/02/20/together-ai-raises-305m-ai-optimized-public-cloud/ Together AI | Series C | 800 | $8.3B | 2026-07-01 | Neocloud inference capacity for open and distilled models | https://techcrunch.com/2026/07/01/neocloud-together-ai-raises-800m-leaps-to-8-3b-valuation/ Mistral AI | Series C (ASML lead, EUR 1.3B of EUR 1.7B) | 1900 | EUR 11.7B post-money (~$14B) | 2025-09-09 | European open-weight model family from 3B Ministral up to Large | https://mistral.ai/news/mistral-ai-raises-1-7-b-to-accelerate-technological-progress-with-ai/ Mistral AI | Reported raise in progress | 3200 | reported EUR 20B target | 2026-06 | Rumoured EUR 3B round; unconfirmed | https://techcrunch.com/2026/06/12/mistral-is-rumored-to-be-raising-e3b-at-e20-valuation/ Notes: Amounts converted to USD millions where the original is in euros, at approximately 1.13 USD/EUR for the September 2025 round (CNBC reported the valuation as ~$14B). The Predibase acquisition amount is the midpoint of a reported $100M-$500M range and should be treated as an estimate, not a disclosed figure. The Mistral 2026 round is rumoured and unconfirmed. Sources: https://www.arcee.ai/blog/arcee-ai-secures-24m-series-a-to-transform-the-landscape-of-small-language-models https://techcrunch.com/2025/06/25/rubrik-acquires-predibase-to-accelerate-adoption-of-ai-agents/ https://techcrunch.com/2026/07/01/neocloud-together-ai-raises-800m-leaps-to-8-3b-valuation/ https://www.cnbc.com/2025/09/09/ai-firm-mistral-valued-at-14-billion-as-chip-giant-asml-takes-major-stake.html TABLE: The cost-of-theft asymmetry, line by line [id: cost-of-theft-asymmetry, 14 rows] Every published figure that bears on the question: what does it cost to build frontier capability, and what does it cost to take it? Line item | Side | Cost (USD) | Basis | Source URL --------- | ---- | ---------- | ----- | ---------- Gemini Ultra 1.0 final training run | Build | 191000000 | Stanford AI Index cloud-rental accounting | https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models GPT-4 final training run | Build | 78000000 | Stanford AI Index cloud-rental accounting | https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models GPT-4 final training run (amortised) | Build | 40000000 | Epoch AI hardware depreciation + energy | https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models DeepSeek-V3 pre-training (claimed) | Build | 5576000 | 2.788M H800 GPU-hours at an assumed $2/hr | https://arxiv.org/abs/2412.19437 DeepSeek-R1 RL stage | Build | 294000 | Nature paper: 512 H800s x 80 hours | https://www.cnn.com/2025/09/19/business/deepseek-ai-training-cost-china-intl 800k-trace reasoning corpus from Claude Opus 5 | Copy | 40000 | Author calculation at $25/MTok output, 2,000 tokens/trace | https://platform.claude.com/docs/en/about-claude/pricing 800k-trace reasoning corpus from gpt-5.6-luna | Copy | 1920 | Author calculation at $1.20/MTok output | https://developers.openai.com/api/docs/pricing Full projection-matrix extraction of gpt-3.5-turbo (estimated) | Copy | 2000 | Carlini et al., estimated query cost | https://arxiv.org/pdf/2403.06634 Partial model-stealing attack on gpt-3.5 | Copy | 200 | Carlini et al., executed attack | https://arxiv.org/pdf/2403.06634 Bespoke-Stratos 17k-trace corpus generation | Copy | n/a | 1.5 hours of DeepSeek-R1 inference; dollar cost undisclosed | https://www.bespokelabs.ai/blog/bespoke-stratos-the-unreasonable-effectiveness-of-reasoning-distillation Sky-T1-32B student fine-tune | Copy | 450 | 8x H100 for 19 hours | https://novasky-ai.github.io/posts/sky-t1/ s1-32B student fine-tune | Copy | 50 | 16x H100 for 26 minutes; dollar figure from press coverage | https://arxiv.org/html/2501.19393v2 Full projection-matrix extraction of ada and babbage | Copy | 20 | Carlini et al., executed attack | https://arxiv.org/pdf/2403.06634 TinyZero R1-Zero-style reproduction | Copy | 30 | Server cost for the experiments, narrow task only | https://www.dailycal.org/news/campus/research-and-ideas/campus-researchers-replicate-disruptive-chinese-ai-for-30/article_a1cc5cd0-dee4-11ef-b8ca-171526dfb895.html Notes: The build side and copy side are not substitutes: a distilled student inherits behaviour on the distribution it was distilled over, not the teacher's full capability surface. But for the specific task a buyer cares about, the copy side is 3-6 orders of magnitude cheaper, and no legal or technical defence currently scales down to that price point. Sources: https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models https://arxiv.org/pdf/2403.06634 https://novasky-ai.github.io/posts/sky-t1/ CHART DATA ---------- CHART: Output price per million tokens: frontier tier vs small tier [id: price-per-mtok-frontier-vs-small, type: bar, unit: USD/MTok] Vendor | Frontier tier (USD/MTok) OpenAI (gpt-6-astra) | 50 Anthropic (Fable 5.1) | 50 Anthropic (Opus 5) | 25 Google (Gemini 3.1 Pro) | 12 Mistral (Medium 3.5) | 7.5 xAI (grok-4.6) | 6 Alibaba (qwen3.8-max) | 6 DeepSeek (V4-Pro peak) | 3.96 Meta (Llama 3.3 70B) | 1.04 Vendor | Small / distilled tier (USD/MTok) OpenAI (gpt-5.6-luna) | 1.2 Anthropic (Haiku 4.5) | 5 Anthropic (Haiku 4.5) | 5 Google (Gemini 2.5 Flash-Lite) | 0.4 Mistral (Ministral 3 3B) | 0.1 xAI (grok-build-0.1) | 2 Alibaba (qwen-turbo) | 0.2 DeepSeek (V4-Flash off-peak) | 0.66 Meta (Llama 3 8B Lite) | 0.14 Notes: Standard list prices as of 3 September 2026. Anthropic appears twice because Opus 5 and Fable 5.1 sit at different points on the same line. Meta prices are Together AI's serverless rates, since Meta does not sell a first-party API for these models. Sources: https://developers.openai.com/api/docs/pricing https://platform.claude.com/docs/en/about-claude/pricing https://ai.google.dev/gemini-api/docs/pricing https://mistral.ai/pricing/api https://docs.x.ai/docs/models https://www.alibabacloud.com/help/en/model-studio/model-pricing https://api-docs.deepseek.com/quick_start/pricing/ https://www.together.ai/pricing CHART: Price of a fixed capability level, 2021-2026 [id: capability-price-decline, type: line, unit: USD/MTok] Date | GPT-3 level (MMLU 42) - a16z (USD/MTok) 2021-11 | 60 2024-11 | 0.06 Date | GPT-3.5 level (MMLU >= 64.8) - Epoch (USD/MTok) 2022-11 | 20 2024-10 | 0.07 Date | GPT-4 level (MMLU >= 86) - Epoch (USD/MTok) 2023-03 | 37.5 2025-02 | 0.18 Date | GPQA Diamond >= 50 - Epoch (USD/MTok) 2023-11 | 15 2024-12 | 0.12 Notes: Plot on a log y-axis. Each series has only the two endpoints published by the source; the intermediate path was not disclosed as a series. Epoch's aggregate finding across six benchmarks is a 9x-900x annual decline with a median near 50x. Sources: https://epoch.ai/data-insights/llm-inference-price-trends https://a16z.com/llmflation-llm-inference-cost/ CHART: Monthly bill vs monthly volume: when does self-hosting a distilled model win? [id: tco-breakeven, type: line, unit: USD] Output tokens per month (millions) | Claude Opus 5 (teacher) (USD) 1 | 45 10 | 450 50 | 2250 100 | 4500 200 | 9000 500 | 22500 1000 | 45000 Output tokens per month (millions) | Claude Sonnet 5 (USD) 1 | 18 10 | 180 50 | 900 100 | 1800 200 | 3600 500 | 9000 1000 | 18000 Output tokens per month (millions) | Claude Haiku 4.5 (vendor small tier) (USD) 1 | 9 10 | 90 50 | 450 100 | 900 200 | 1800 500 | 4500 1000 | 9000 Output tokens per month (millions) | gpt-5.6-luna (vendor nano tier) (USD) 1 | 2 10 | 20 50 | 100 100 | 200 200 | 400 500 | 1000 1000 | 2000 Output tokens per month (millions) | DeepSeek V4-Flash off-peak (USD) 1 | 2 10 | 15 50 | 77 100 | 154 200 | 308 500 | 770 1000 | 1540 Output tokens per month (millions) | Self-hosted distilled 8B, 1x H100 + 0.25 FTE (USD) 1 | 5919 10 | 5919 50 | 5919 100 | 5919 200 | 5919 500 | 5919 1000 | 5919 Notes: Author-computed. Assumes 4 input tokens per output token. Self-hosted line is flat at $5,919/month ($1,752 GPU + $4,167 staffing) up to ~1,051M output tokens/month capacity. Crossings: 132M vs Opus 5, 329M vs Sonnet 5, 658M vs Haiku 4.5. The gpt-5.6-luna and DeepSeek lines never cross within capacity. Sources: https://platform.claude.com/docs/en/about-claude/pricing https://developers.openai.com/api/docs/pricing https://api-docs.deepseek.com/quick_start/pricing/ https://www.gmicloud.ai/en/blog/nvidia-h100-gpu-pricing-2026-rent-vs-buy-cost-analysis CHART: Training cost: frontier runs vs distilled students [id: training-cost-ladder-chart, type: bar, unit: USD] Model / run | Reported training cost (USD) DeepSeek infrastructure (SemiAnalysis est.) | 1600000000 Gemini Ultra 1.0 (AI Index) | 191000000 GPT-4 (AI Index) | 78000000 GPT-4 (Epoch amortised) | 40000000 DeepSeek-V3 final run (claimed) | 5576000 DeepSeek-R1 RL stage (Nature) | 294000 Sky-T1-32B | 450 s1-32B | 50 TinyZero | 30 Notes: Log scale spans eight orders of magnitude. The bars are not accounted on a common basis - see the training-cost table notes. The point of the chart is the shape of the ladder, not a like-for-like comparison. Sources: https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models https://semianalysis.com/2025/01/31/deepseek-debates/ https://arxiv.org/abs/2412.19437 https://novasky-ai.github.io/posts/sky-t1/ https://www.cnn.com/2025/09/19/business/deepseek-ai-training-cost-china-intl CHART: Teacher-query bill to assemble a distillation corpus, 2026 list prices [id: corpus-cost-by-teacher, type: bar, unit: USD] Teacher model | 17k traces (Sky-T1 / Bespoke-Stratos scale) (USD) Claude Opus 5 | 850 gpt-5.6-sol | 680 Gemini 3.5 Flash | 306 grok-4.6 | 204 deepseek-v4-pro (off-peak) | 67 gpt-5.6-luna | 41 gpt-oss-120b (Groq) | 20 Teacher model | 800k traces (DeepSeek R1-Distill scale) (USD) Claude Opus 5 | 40000 gpt-5.6-sol | 32000 Gemini 3.5 Flash | 14400 grok-4.6 | 9600 deepseek-v4-pro (off-peak) | 3168 gpt-5.6-luna | 1920 gpt-oss-120b (Groq) | 960 Notes: Author calculation: 2,000 output tokens per trace, input cost ignored. Halve every bar again if the Batch API 50% discount applies. Sources: https://platform.claude.com/docs/en/about-claude/pricing https://developers.openai.com/api/docs/pricing https://ai.google.dev/gemini-api/docs/pricing https://api-docs.deepseek.com/quick_start/pricing/ https://console.groq.com/docs/models CHART: Nvidia market capitalisation around the distillation shocks [id: nvidia-market-cap-impact, type: bar, unit: USD billions] Event | Market-cap change (USD billions) 27 Jan 2025: DeepSeek-R1 shock (one session) | -589 14 May - 8 Jul 2026: drawdown from peak | -1000 Event | Market-cap level (USD billions) Oct 2025: first $5T company | 5060 2 Sep 2026 | 5430 Notes: The 27 January 2025 move was a 17% single-session fall and the largest one-day market-cap loss in US stock-market history; the Nasdaq 100 fell 3% and the S&P 500 1.5% the same day. The 2026 drawdown was attributed to rotation into memory and storage semiconductors rather than to distillation news. Sources: https://www.cnbc.com/2025/01/27/nvidia-sheds-almost-600-billion-in-market-cap-biggest-drop-ever.html https://finance.yahoo.com/markets/stocks/articles/nvidia-stock-valuation-falls-pre-132505643.html https://stockanalysis.com/stocks/nvda/market-cap/ CHART: DeepSeek V4 pricing before and after 16 August 2026 [id: deepseek-price-reversal, type: bar, unit: USD/MTok] Model and token type | Before 16 Aug 2026 (USD/MTok) V4-Flash input (cache miss) | 0.14 V4-Flash output | 0.28 V4-Pro input (cache miss) | 0.435 V4-Pro output | 0.87 Model and token type | After, off-peak (USD/MTok) V4-Flash input (cache miss) | 0.22 V4-Flash output | 0.66 V4-Pro input (cache miss) | 0.66 V4-Pro output | 1.98 Model and token type | After, peak (USD/MTok) V4-Flash input (cache miss) | 0.44 V4-Flash output | 1.32 V4-Pro input (cache miss) | 1.32 V4-Pro output | 3.96 Notes: Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday - Beijing business hours - so 17 of 24 hours stay at the off-peak rate and Western buyers are largely insulated. Cache-hit input tokens, not shown here, rose by as much as 1,100%. Sources: https://api-docs.deepseek.com/quick_start/pricing/ https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html DETAIL RECORDS -------------- TABLE: List prices per model, September 2026 [63 rows, from data/financial.json $.extras] Model | Vendor | Tier | Input (USD/MTok) | Output (USD/MTok) | Params (B) | Release | Source URL ----- | ------ | ---- | ---------------- | ----------------- | ---------- | ------- | ---------- gpt-6-astra | OpenAI | frontier | 10 | 50 | n/a | n/a | https://developers.openai.com/api/docs/pricing gpt-5.6-sol | OpenAI | frontier | 4 | 20 | n/a | 2026-07 | https://developers.openai.com/api/docs/pricing gpt-5.6-terra | OpenAI | frontier | 2 | 12 | n/a | 2026-07 | https://developers.openai.com/api/docs/pricing gpt-5.6-luna | OpenAI | distilled | 0.2 | 1.2 | n/a | 2026-07 | https://developers.openai.com/api/docs/pricing gpt-5.5 | OpenAI | frontier | 5 | 30 | n/a | n/a | https://developers.openai.com/api/docs/pricing gpt-5.5-pro | OpenAI | frontier | 30 | 180 | n/a | n/a | https://developers.openai.com/api/docs/pricing gpt-5.4 | OpenAI | frontier | 2.5 | 15 | n/a | n/a | https://developers.openai.com/api/docs/pricing gpt-5.4-mini | OpenAI | distilled | 0.75 | 4.5 | n/a | n/a | https://developers.openai.com/api/docs/pricing gpt-5.4-nano | OpenAI | distilled | 0.2 | 1.25 | n/a | n/a | https://developers.openai.com/api/docs/pricing gpt-5 | OpenAI | frontier | 1.25 | 10 | n/a | 2025-08 | https://developers.openai.com/api/docs/pricing gpt-5-mini | OpenAI | distilled | 0.25 | 2 | n/a | 2025-08 | https://developers.openai.com/api/docs/pricing gpt-5-nano | OpenAI | distilled | 0.05 | 0.4 | n/a | 2025-08 | https://developers.openai.com/api/docs/pricing gpt-4o | OpenAI | frontier | 2.5 | 10 | n/a | 2024-05 | https://developers.openai.com/api/docs/pricing gpt-4o-mini | OpenAI | distilled | 0.15 | 0.6 | n/a | 2024-07 | https://developers.openai.com/api/docs/pricing o1 | OpenAI | frontier | 15 | 60 | n/a | 2024-12 | https://developers.openai.com/api/docs/pricing o3 | OpenAI | frontier | 2 | 8 | n/a | 2025-04 | https://developers.openai.com/api/docs/pricing o4-mini | OpenAI | distilled | 1.1 | 4.4 | n/a | 2025-04 | https://developers.openai.com/api/docs/pricing Claude Fable 5.1 | Anthropic | frontier | 10 | 50 | n/a | n/a | https://platform.claude.com/docs/en/about-claude/pricing Claude Opus 5 | Anthropic | frontier | 5 | 25 | n/a | n/a | https://platform.claude.com/docs/en/about-claude/pricing Claude Opus 4.1 | Anthropic | frontier | 15 | 75 | n/a | 2025-08 | https://platform.claude.com/docs/en/about-claude/pricing Claude Sonnet 5 | Anthropic | frontier | 2 | 10 | n/a | n/a | https://platform.claude.com/docs/en/about-claude/pricing Claude Sonnet 4.6 | Anthropic | frontier | 3 | 15 | n/a | n/a | https://platform.claude.com/docs/en/about-claude/pricing Claude Haiku 4.5 | Anthropic | distilled | 1 | 5 | n/a | 2025-10 | https://platform.claude.com/docs/en/about-claude/pricing Claude Haiku 3.5 | Anthropic | distilled | 0.8 | 4 | n/a | 2024-11 | https://platform.claude.com/docs/en/about-claude/pricing Gemini 3.1 Pro Preview | Google | frontier | 2 | 12 | n/a | n/a | https://ai.google.dev/gemini-api/docs/pricing Gemini 3.8 Flash | Google | distilled | 0.75 | 3.75 | n/a | n/a | https://ai.google.dev/gemini-api/docs/pricing Gemini 3.5 Flash | Google | distilled | 1.5 | 9 | n/a | n/a | https://ai.google.dev/gemini-api/docs/pricing Gemini 3.5 Flash-Lite | Google | distilled | 0.3 | 2.5 | n/a | n/a | https://ai.google.dev/gemini-api/docs/pricing Gemini 3.1 Flash-Lite | Google | distilled | 0.25 | 1.5 | n/a | n/a | https://ai.google.dev/gemini-api/docs/pricing Gemini 2.5 Pro | Google | frontier | 1.25 | 10 | n/a | 2025-03 | https://ai.google.dev/gemini-api/docs/pricing Gemini 2.5 Flash | Google | distilled | 0.3 | 2.5 | n/a | 2025-04 | https://ai.google.dev/gemini-api/docs/pricing Gemini 2.5 Flash-Lite | Google | distilled | 0.1 | 0.4 | n/a | 2025-06 | https://ai.google.dev/gemini-api/docs/pricing deepseek-v4-pro (peak) | DeepSeek | frontier | 1.32 | 3.96 | n/a | 2026-08 | https://api-docs.deepseek.com/quick_start/pricing/ deepseek-v4-pro (off-peak) | DeepSeek | frontier | 0.66 | 1.98 | n/a | 2026-08 | https://api-docs.deepseek.com/quick_start/pricing/ deepseek-v4-flash (peak) | DeepSeek | distilled | 0.44 | 1.32 | n/a | 2026-07 | https://api-docs.deepseek.com/quick_start/pricing/ deepseek-v4-flash (off-peak) | DeepSeek | distilled | 0.22 | 0.66 | n/a | 2026-07 | https://api-docs.deepseek.com/quick_start/pricing/ deepseek-v4-flash (pre-16 Aug 2026) | DeepSeek | distilled | 0.14 | 0.28 | n/a | 2026-07 | https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html deepseek-v4-pro (pre-16 Aug 2026) | DeepSeek | frontier | 0.435 | 0.87 | n/a | 2025-10 | https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html Mistral Medium 3.5 | Mistral AI | frontier | 1.5 | 7.5 | n/a | n/a | https://mistral.ai/pricing/api Mistral Large 3 | Mistral AI | frontier | 0.5 | 1.5 | n/a | n/a | https://mistral.ai/pricing/api Mistral Small 4 | Mistral AI | distilled | 0.15 | 0.6 | n/a | n/a | https://mistral.ai/pricing/api Ministral 3 (3B) | Mistral AI | distilled | 0.1 | 0.1 | 3 | n/a | https://mistral.ai/pricing/api Ministral 3 (8B) | Mistral AI | distilled | 0.15 | 0.15 | 8 | n/a | https://mistral.ai/pricing/api Ministral 3 (14B) | Mistral AI | distilled | 0.2 | 0.2 | 14 | n/a | https://mistral.ai/pricing/api Codestral | Mistral AI | distilled | 0.3 | 0.9 | n/a | n/a | https://mistral.ai/pricing/api grok-4.6 (<200k) | xAI | frontier | 2 | 6 | n/a | 2026-08 | https://docs.x.ai/docs/models grok-4.5 (<200k) | xAI | frontier | 2 | 6 | n/a | n/a | https://docs.x.ai/docs/models grok-4.3 (<200k) | xAI | frontier | 1.25 | 2.5 | n/a | n/a | https://docs.x.ai/docs/models grok-build-0.1 (<200k) | xAI | distilled | 1 | 2 | n/a | n/a | https://docs.x.ai/docs/models qwen3.8-max | Alibaba | frontier | 2 | 6 | n/a | n/a | https://www.alibabacloud.com/help/en/model-studio/model-pricing qwen3.7-plus | Alibaba | frontier | 0.4 | 1.6 | n/a | n/a | https://www.alibabacloud.com/help/en/model-studio/model-pricing qwen3.8-flash | Alibaba | distilled | 0.15 | 0.47 | n/a | n/a | https://www.alibabacloud.com/help/en/model-studio/model-pricing qwen-turbo | Alibaba | distilled | 0.05 | 0.2 | n/a | n/a | https://www.alibabacloud.com/help/en/model-studio/model-pricing qwen3.8-27b | Alibaba | open | 0.5 | 3 | 27 | n/a | https://www.alibabacloud.com/help/en/model-studio/model-pricing qwen3-8b | Alibaba | open | 0.18 | 0.7 | 8 | 2025-04 | https://www.alibabacloud.com/help/en/model-studio/model-pricing Llama 3.3 70B (Together AI) | Meta / Together AI | open | 1.04 | 1.04 | 70 | 2024-12 | https://www.together.ai/pricing Llama 3 8B Instruct Lite (Together AI) | Meta / Together AI | open | 0.14 | 0.14 | 8 | 2024-04 | https://www.together.ai/pricing Qwen2.5 7B Instruct Turbo (Together AI) | Alibaba / Together AI | open | 0.3 | 0.3 | 7 | 2024-09 | https://www.together.ai/pricing gpt-oss-120b (Together AI) | OpenAI / Together AI | open | 0.15 | 0.6 | 116.8 | 2025-08 | https://www.together.ai/pricing gpt-oss-120b (Groq) | OpenAI / Groq | open | 0.15 | 0.6 | 116.8 | 2025-08 | https://console.groq.com/docs/models gpt-oss-20b (Groq) | OpenAI / Groq | open | 0.075 | 0.3 | 20.9 | 2025-08 | https://console.groq.com/docs/models Qwen3.8-27B (Groq) | Alibaba / Groq | open | 0.8 | 4 | 27 | n/a | https://console.groq.com/docs/models GLM-5.3-Flash (Together AI) | Zhipu / Together AI | open | 0.15 | 0.5 | n/a | n/a | https://www.together.ai/pricing TABLE: Documented training runs [13 rows, from data/financial.json $.extras] Run | Organisation | Cost (USD) | GPU-hours (hours) | Hardware | Date | Note | Source URL --- | ------------ | ---------- | ----------------- | -------- | ---- | ---- | ---------- Gemini Ultra 1.0 | Google DeepMind | 191000000 | n/a | TPU v4, ~35 MW | 2023-12 | Stanford AI Index cloud-rental accounting. Epoch's amortised-hardware method gives ~$30M for the same run. | https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models GPT-4 | OpenAI | 78000000 | n/a | NVIDIA A100 cluster | 2023-03 | Stanford AI Index cloud-rental figure; Epoch amortised estimate is ~$40M. Hardware is 47-67% of total development cost, R&D staff 29-49%, energy 2-6%. | https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models DeepSeek total infrastructure | DeepSeek / High-Flyer | 1600000000 | n/a | ~50,000 NVIDIA Hopper GPUs | 2025-01 | SemiAnalysis estimate of server capex, plus ~$944M of operating cost. Disputes DeepSeek's $5.6M framing as marginal-cost-only. | https://semianalysis.com/2025/01/31/deepseek-debates/ DeepSeek-V3 (full training run) | DeepSeek | 5576000 | 2788000 | NVIDIA H800 | 2024-12 | The technical report states 2.788M H800 GPU-hours. The dollar figure is that multiplied by an assumed $2/GPU-hour; the paper itself does not state a USD cost. | https://arxiv.org/abs/2412.19437 DeepSeek-R1 (reinforcement-learning stage) | DeepSeek | 294000 | 40960 | 512x NVIDIA H800 for 80 hours | 2025-09 | Disclosed in the peer-reviewed Nature paper. Covers GPU cost of the final RL run only; excludes the V3 base model, data, energy, infrastructure and staff. | https://www.cnn.com/2025/09/19/business/deepseek-ai-training-cost-china-intl DeepSeek-R1-Distill-Qwen-32B | DeepSeek | n/a | n/a | undisclosed | 2025-01 | Supervised fine-tune of Qwen2.5-32B on 800k R1-generated samples. Cost undisclosed. Scores 72.6 AIME 2024, 94.3 MATH-500, 62.1 GPQA Diamond. | https://arxiv.org/html/2501.12948v1 DeepSeek-R1-Distill-Llama-70B | DeepSeek | n/a | n/a | undisclosed | 2025-01 | Same 800k-sample corpus applied to Llama-3.3-70B. Scores 70.0 AIME 2024 pass@1 (86.7 cons@64) and 94.5 MATH-500 pass@1. Released under MIT licence. | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B Bespoke-Stratos-32B | Bespoke Labs | n/a | n/a | undisclosed | 2025-01-22 | Distilled from DeepSeek-R1 on 17k traces (47x fewer than R1-Distill-Qwen-32B); traces generated in 1.5 hours of teacher inference. GPT-4o-mini filtering raised the retained-correct rate from 25% to 73%. Neither dollar cost nor training GPU-hours are disclosed in the post; only the 1.5 h of teacher trace generation is given. | https://www.bespokelabs.ai/blog/bespoke-stratos-the-unreasonable-effectiveness-of-reasoning-distillation Sky-T1-32B-Preview | NovaSky, UC Berkeley | 450 | 152 | 8x NVIDIA H100 for 19 hours, DeepSpeed ZeRO-3 offload | 2025-01-10 | Base Qwen2.5-32B-Instruct, teacher QwQ-32B-Preview, 17k examples (10k math, 5k code, 1k science/puzzles). Math500 82.4 vs o1-preview 81.4; AIME24 43.3 vs 40.0. | https://novasky-ai.github.io/posts/sky-t1/ s1-32B | Stanford / University of Washington | 50 | 6.9 | 16x NVIDIA H100 for 26 minutes, PyTorch FSDP | 2025-02-01 | 1,000 traces (s1K) distilled from Gemini Thinking Experimental into Qwen2.5-32B-Instruct. Exceeds o1-preview on competition maths by up to 27% with budget forcing. The ~$50 figure comes from press coverage, not the paper. | https://arxiv.org/html/2501.19393v2 TinyZero | UC Berkeley (Jiayi Pan et al.) | 30 | n/a | undisclosed | 2025-01-24 | RL on a 3B Qwen base for Countdown and multiplication. Reproduces R1-Zero-style self-verification but only on very restricted task types. The $30 is server cost for the experiments. | https://www.dailycal.org/news/campus/research-and-ideas/campus-researchers-replicate-disruptive-chinese-ai-for-30/article_a1cc5cd0-dee4-11ef-b8ca-171526dfb895.html gpt-oss-120b | OpenAI | n/a | n/a | undisclosed | 2025-08-05 | 116.8B total / 5.1B active parameters, Apache 2.0. Training cost undisclosed. Matches or exceeds o4-mini on competition coding; now the price floor for hosted open reasoning models at $0.15/$0.60 per MTok. | https://openai.com/index/gpt-oss-model-card/ Frontier training run, 2026 class | Multiple | n/a | n/a | 1e26-1e27 FLOP class | 2026 | Epoch AI projects that the largest training runs will exceed $1B by 2027, extrapolating from 2.4x/year growth since 2016 (95% CI 2.0x-3.1x). No verified 2026 per-model figure has been published by any lab. | https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models TABLE: Market events [21 rows, from data/financial.json $.extras] Date | Event | Impact | Figure | Source URL ---- | ----- | ------ | ------ | ---------- 2024-10-01 | OpenAI ships Model Distillation (Stored Completions + Evals) in the API | Distillation becomes a first-party product; Stored Completions free, fine-tuning at standard rates | $1.50-$25 per MTok training depending on student | https://developers.openai.com/api/docs/pricing 2024-12-03 | AWS announces Amazon Bedrock Model Distillation at re:Invent | Marketed as up to 500% faster and up to 75% cheaper with under 2% accuracy loss on RAG | 75% cost reduction claim | https://press.aboutamazon.com/2024/12/aws-strengthens-amazon-bedrock-with-industry-first-ai-safeguard-new-agent-capability-and-model-customization 2025-01-20 | DeepSeek releases R1 plus six open-weight distilled students under MIT licence | Frontier-class reasoning becomes free to download and self-host | R1-Distill-Llama-70B: 86.7 AIME 2024 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B 2025-01-27 | Nvidia posts the largest single-day market-cap loss in US history | Nvidia -17%, Nasdaq 100 -3%, S&P 500 -1.5%, semiconductor index worst day since March 2020 | -$589 billion | https://www.cnbc.com/2025/01/27/nvidia-sheds-almost-600-billion-in-market-cap-biggest-drop-ever.html 2025-01-29 | Microsoft and OpenAI confirm an investigation into DeepSeek-linked accounts | Distillation reframed from a research technique into an IP-theft allegation with policy consequences | Large-scale API exfiltration observed in late 2024 | https://www.bloomberg.com/news/articles/2025-01-29/microsoft-probing-if-deepseek-linked-group-improperly-obtained-openai-data 2025-01-31 | SemiAnalysis publishes 'DeepSeek Debates', disputing the $5.6M cost claim | Reframes the cheap-training narrative as a marginal-cost accounting artefact | $1.6B server capex, ~50,000 Hopper GPUs, ~$944M opex | https://semianalysis.com/2025/01/31/deepseek-debates/ 2025-02-20 | Together AI raises $305M Series B | Capital flows into serving infrastructure for open and distilled models rather than into distillation tooling | $3.3B valuation | https://siliconangle.com/2025/02/20/together-ai-raises-305m-ai-optimized-public-cloud/ 2025-04-16 | House Select Committee on the CCP publishes 'DeepSeek Unmasked' | Congressional finding that unlawful distillation was highly likely; feeds into export-control and procurement policy | 85% of chatbot replies found filtered to CCP narratives | https://www.techpolicy.press/us-house-select-committee-report-accuses-deepseek-of-spying-and-circumventing-export-controls-on-chips/ 2025-05-01 | Amazon Bedrock Model Distillation reaches general availability | Distilled models must run on Provisioned Throughput, imposing an hourly floor instead of per-token billing | $1.95/model/month storage plus PT hourly rates | https://docs.aws.amazon.com/bedrock/latest/userguide/model-distillation.html 2025-06-10 | OpenAI cuts o3 pricing by ~80% | Demonstrates that inference-stack optimisation, not distillation, can deliver an order-of-magnitude price move on an unchanged model | $10/$40 to $2/$8 per MTok | https://venturebeat.com/ai/openai-announces-80-price-drop-for-o3-its-most-powerful-reasoning-model 2025-06-25 | Rubrik acquires Predibase | The clearest exit for a small-model fine-tuning tooling company; modest relative to inference-infrastructure valuations | Reported $100M-$500M | https://techcrunch.com/2025/06/25/rubrik-acquires-predibase-to-accelerate-adoption-of-ai-agents/ 2025-08-05 | OpenAI releases gpt-oss-120b and gpt-oss-20b under Apache 2.0 | Sets a new floor price for hosted open reasoning models and removes much of the incentive to distill OpenAI models illicitly | $0.15/$0.60 and $0.075/$0.30 per MTok on Groq | https://openai.com/index/gpt-oss-model-card/ 2025-09-09 | Mistral raises EUR 1.7B Series C with ASML leading | European sovereign-AI capital backs a full small-to-large open-weight model ladder | EUR 11.7B post-money, ASML ~11% fully diluted | https://mistral.ai/news/mistral-ai-raises-1-7-b-to-accelerate-technological-progress-with-ai/ 2025-09-17 | Nature publishes DeepSeek-R1 with a $294,000 training-cost disclosure | First peer-reviewed cost figure for a frontier-class reasoning model; immediately contested as marginal-cost-only | $294,000, 512 H800s x 80 hours | https://www.cnn.com/2025/09/19/business/deepseek-ai-training-cost-china-intl 2025-10-29 | Nvidia becomes the first $5 trillion company; DeepSeek launches V4-Pro the same day | The distillation-driven efficiency narrative fails to dent the compute build-out | $5.06 trillion close | https://www.thenationalnews.com/future/technology/2026/04/25/will-deepseeks-new-ai-model-crash-nvidias-5tn-party/ 2026-02-12 | OpenAI submits a memo on DeepSeek distillation to the House Select Committee on China | Escalates distillation from a civil-IP question to a national-security one; describes obfuscated routers and unauthorised resellers | n/a | https://cdn.openai.com/pdf/045aa967-ee96-4a09-94ee-3098ddf6db2c/OpenAI-US-House-Select-Cmte-Update-%5B021226%5D.pdf 2026-07-01 | Together AI raises $800M at an $8.3B valuation | Confirms that the value in the open/distilled model stack accrues to inference capacity | 2.5x valuation increase in 17 months | https://techcrunch.com/2026/07/01/neocloud-together-ai-raises-800m-leaps-to-8-3b-valuation/ 2026-07-08 | Nvidia has shed roughly $1 trillion from its 14 May 2026 peak | Attributed to rotation into memory and storage semiconductors, not to model-efficiency news | -16% from peak; ~97% server GPU share retained | https://finance.yahoo.com/markets/stocks/articles/nvidia-stock-valuation-falls-pre-132505643.html 2026-07-31 | DeepSeek launches V4-Flash at $0.14/$0.28 per MTok | Described as accelerating the AI industry's race to zero; the price holds for sixteen days | $0.14 in / $0.28 out per MTok | https://www.axios.com/2026/08/01/deepseek-model-cheap-ai-price-war 2026-08-16 | DeepSeek raises V4 API prices by up to 1,100% and introduces peak/off-peak rates | The first major reversal of the cheap-inference trend; signals that ultra-low prices were capacity-constrained rather than structural | V4-Pro output $0.87 to $1.98 off-peak / $3.96 peak | https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html 2026-09-02 | Nvidia market capitalisation at ~$5.43 trillion, up ~28% year on year | About 9x the value erased on 27 January 2025 has since been added | $5.43 trillion | https://stockanalysis.com/stocks/nvda/market-cap/ TABLE: GPU rental rates [11 rows, from data/financial.json $.extras] GPU | Provider | Price (USD/hour) | Type | Source URL --- | -------- | ---------------- | ---- | ---------- NVIDIA H100 80GB | Vast.ai | 1.73 | on-demand (cheapest listed) | https://getdeploying.com/gpus/nvidia-h100 NVIDIA H100 PCIe | RunPod | 1.99 | on-demand | https://getdeploying.com/gpus/nvidia-h100 NVIDIA H100 PCIe | GMI Cloud | 2 | on-demand | https://www.gmicloud.ai/en/blog/nvidia-h100-gpu-pricing-2026-rent-vs-buy-cost-analysis NVIDIA H100 SXM | GMI Cloud | 2.4 | on-demand | https://www.gmicloud.ai/en/blog/nvidia-h100-gpu-pricing-2026-rent-vs-buy-cost-analysis NVIDIA H100 | Market median (~40 providers) | 3.39 | on-demand median, 3 Sep 2026 | https://getdeploying.com/gpus/nvidia-h100 NVIDIA HGX H100 | Together AI | 3.99 | dedicated, promotional (regular $5.49) | https://www.together.ai/pricing NVIDIA H100 80GB | Fireworks AI | 8 | on-demand from 1 Sep 2026 (was $7.00) | https://fireworks.ai/pricing NVIDIA HGX B200 | Together AI | 8.99 | dedicated | https://www.together.ai/pricing NVIDIA B200 180GB | Fireworks AI | 13 | on-demand from 1 Sep 2026 (was $10.00) | https://fireworks.ai/pricing NVIDIA GB300 288GB | Fireworks AI | 20 | on-demand from 1 Sep 2026 (was $18.00) | https://fireworks.ai/pricing NVIDIA H100 (8-GPU node) | AWS / GCP / Azure | 48 | on-demand midpoint of $32-$64/hr range | https://getdeploying.com/gpus/nvidia-h100 TOTAL-COST-OF-OWNERSHIP ASSUMPTIONS (every input behind the self-hosting model) description: Every assumption behind the TCO scenario table and breakeven chart, stated so a reader can substitute their own. workloadShape: 4 input tokens per output token, typical of RAG and summarisation. Costs are expressed per million OUTPUT tokens: blended = 4 x input_price + output_price. gpuHourlyRate: 2.4 gpuHourlyRateNote: NVIDIA H100 SXM on-demand at GMI Cloud. The market median across ~40 providers is $3.39/GPU-hour; the cheapest listed on-demand is Vast.ai at $1.73; AWS, GCP and Azure charge $4.00-$8.00. hoursPerMonth: 730 monthlyGpuCost: 1752 staffingFte: 0.25 staffingAnnualFullyLoaded: 200000 monthlyStaffingCost: 4167 monthlyTotalSelfHosted: 5919 assumedThroughputTokensPerSec: 400 throughputNote: 400 output tok/s is the sourced anchor for a 7B-class dense model in FP8 on one H100 with vLLM continuous batching. Real cost per output MTok on identical H100 hardware has been measured spanning $0.21 to $15.25 depending purely on request concurrency, so this is the single largest source of error in the model. monthlyCapacityMtokOut: 1051 breakevenMtokOutPerMonth: vs_claude_opus_5 = 132; vs_claude_sonnet_5 = 329; vs_claude_haiku_4_5 = 658; vs_gpt_5_6_luna = 2959; vs_deepseek_v4_flash_offpeak = 3843 excluded: One-off distillation cost ($450-$40,000 depending on corpus size and teacher); Redundancy, autoscaling headroom and failover capacity; Data egress and storage; Batch API discounts (50% at OpenAI and Anthropic) which would push the breakevens further out; Prompt caching, which at Anthropic cuts cache-read input to 0.1x base (0.025x on Fable 5.1); Quality loss from the distilled student relative to the teacher sources: https://www.gmicloud.ai/en/blog/nvidia-h100-gpu-pricing-2026-rent-vs-buy-cost-analysis; https://getdeploying.com/gpus/nvidia-h100; https://arxiv.org/html/2606.11690v1; https://platform.claude.com/docs/en/about-claude/pricing MARKET CONTEXT openrouterTokenShare: note = OpenRouter's 100-trillion-token study and follow-up reporting. Treat the mid-2026 figures as press-reported rather than peer-reviewed.; usModelShareJune2025_pct = 70; usModelShareMid2026_pct = 30; openWeightShareLate2025_pct = 33; openWeightShareMid2026 = majority (>50%); chineseOpenWeightShareMay2026_pct = 61; priceElasticity = A 10% price decrease corresponds to only ~0.5-0.7% more usage, implying demand is driven by capability rather than price.; smallModelTrend = Models under 15B parameters are losing API share while 15B-70B models gain, though the study cautions that small models are disproportionately self-hosted and therefore under-counted.; sources = https://arxiv.org/html/2601.10088v1; https://openrouter.ai/state-of-ai; https://fourweekmba.com/ai-openrouter-us-models-token-share-deepseek-volume-revenue-spl/ smallLanguageModelMarket: note = Analyst-firm estimates, not primary data. Included for order of magnitude only; the figures diverge widely between firms.; size2025_usd_bn = 9.16; size2026_usd_bn = 10.99; size2030_usd_bn = 22.45; cagr_pct = 19.6; sources = https://www.researchandmarkets.com/reports/6076406/small-language-model-market-report; https://www.marketsandmarkets.com/Market-Reports/small-language-model-market-4008452.html TIMELINE OF THIS PERSPECTIVE (30 EVENTS) ---------------------------------------- 2024-10-01 | product | OpenAI ships Model Distillation in the API 2024-11 | market | a16z publishes 'LLMflation' 2024-12-03 | product | AWS announces Amazon Bedrock Model Distillation at re:Invent 2024-12-26 | research | DeepSeek-V3 technical report discloses 2.788M H800 GPU-hours 2025-01-10 | research | Sky-T1-32B-Preview trained for under $450 2025-01-20 | product | DeepSeek releases R1 plus six open-weight distilled students under MIT licence 2025-01-22 | research | Bespoke-Stratos-32B distilled from DeepSeek-R1 on 17k traces 2025-01-24 | research | TinyZero reproduces R1-Zero-style behaviour for under $30 2025-01-27 | market | Nvidia loses $589B of market cap in one session 2025-01-29 | legal | Microsoft and OpenAI investigate DeepSeek-linked accounts 2025-01-31 | market | SemiAnalysis disputes the $5.6M figure 2025-02-01 | research | s1-32B: 1,000 traces, 26 minutes, reported ~$50 2025-02-20 | market | Together AI raises $305M Series B at $3.3B 2025-04-16 | policy | House Select Committee publishes 'DeepSeek Unmasked' 2025-05-01 | product | Amazon Bedrock Model Distillation reaches general availability 2025-06-10 | market | OpenAI cuts o3 pricing by ~80% 2025-06-25 | market | Rubrik acquires Predibase 2025-08-05 | product | OpenAI releases gpt-oss-120b and gpt-oss-20b under Apache 2.0 2025-09-09 | market | Mistral raises EUR 1.7B Series C at EUR 11.7B, ASML leads 2025-09-17 | research | Nature publishes the DeepSeek-R1 paper with a $294,000 cost figure 2025-10-29 | market | Nvidia becomes the first $5 trillion company 2026-01-06 | market | One year on, DeepSeek no longer moves markets 2026-02-12 | policy | OpenAI memo to the House Select Committee on China 2026-02-23 | policy | Frontier Model Forum publishes an issue brief on adversarial distillation 2026-07-01 | market | Together AI raises $800M at $8.3B 2026-07-08 | market | Nvidia sheds roughly $1T from its 14 May 2026 peak 2026-07-31 | product | DeepSeek V4-Flash launches at $0.14/$0.28 per MTok 2026-08-06 | market | DeepSeek warns of a significant API price increase 2026-08-16 | market | DeepSeek raises V4 prices by up to 1,100% and introduces peak/off-peak rates 2026-09-02 | market | Nvidia market capitalisation stands at ~$5.43T SOURCES CITED BY THIS SECTION (63) ---------------------------------- [1] OpenAI API Pricing — OpenAI, 2026-09-03 (pricing) https://developers.openai.com/api/docs/pricing [2] Anthropic Claude Pricing — Anthropic, 2026-09-03 (pricing) https://platform.claude.com/docs/en/about-claude/pricing [3] Gemini API Pricing — Google, 2026-09-03 (pricing) https://ai.google.dev/gemini-api/docs/pricing [4] DeepSeek API Pricing — DeepSeek, 2026-09-03 (pricing) https://api-docs.deepseek.com/quick_start/pricing/ [5] Mistral API Pricing — Mistral AI, 2026-09-03 (pricing) https://mistral.ai/pricing/api [6] xAI Models and Pricing — xAI, 2026-09-03 (pricing) https://docs.x.ai/docs/models [7] Alibaba Cloud Model Studio model pricing — Alibaba Cloud, 2026-09-03 (pricing) https://www.alibabacloud.com/help/en/model-studio/model-pricing [8] Together AI Pricing — Together AI, 2026-09-03 (pricing) https://www.together.ai/pricing [9] Fireworks AI Pricing — Fireworks AI, 2026-09-03 (pricing) https://fireworks.ai/pricing [10] Groq supported models — Groq, 2026-09-03 (pricing) https://console.groq.com/docs/models [11] Amazon Bedrock Pricing — Amazon Web Services, 2026-09-03 (pricing) https://aws.amazon.com/bedrock/pricing/ [12] Customize a model with distillation in Amazon Bedrock — Amazon Web Services, 2026 (docs) https://docs.aws.amazon.com/bedrock/latest/userguide/model-distillation.html [13] AWS Strengthens Amazon Bedrock with Industry-First AI Safeguard, New Agent Capability and Model Customization — Amazon, 2024-12-03 (news) https://press.aboutamazon.com/2024/12/aws-strengthens-amazon-bedrock-with-industry-first-ai-safeguard-new-agent-capability-and-model-customization [14] Model Distillation in the API — OpenAI, 2024-10-01 (blog) https://openai.com/index/api-model-distillation/ [15] How to use stored completions and distillation in Azure OpenAI — Microsoft, 2026 (docs) https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/stored-completions?view=foundry-classic [16] Vertex AI pricing — Google Cloud, 2026 (pricing) https://cloud.google.com/vertex-ai/pricing [17] Sky-T1: Train your own O1 preview model within $450 — NovaSky, UC Berkeley, 2025-01-10 (blog) https://novasky-ai.github.io/posts/sky-t1/ [18] Researchers open source Sky-T1, a reasoning AI model that can be trained for less than $450 — TechCrunch, 2025-01-11 (news) https://techcrunch.com/2025/01/11/researchers-open-source-sky-t1-a-reasoning-ai-model-that-can-be-trained-for-less-than-450/ [19] s1: Simple test-time scaling — arXiv (Muennighoff et al.), 2025-02-01 (paper) https://arxiv.org/html/2501.19393v2 [20] Bespoke-Stratos: The unreasonable effectiveness of reasoning distillation — Bespoke Labs, 2025-01-22 (blog) https://www.bespokelabs.ai/blog/bespoke-stratos-the-unreasonable-effectiveness-of-reasoning-distillation [21] Campus researchers replicate disruptive Chinese AI for $30 — The Daily Californian, 2025-01-31 (news) https://www.dailycal.org/news/campus/research-and-ideas/campus-researchers-replicate-disruptive-chinese-ai-for-30/article_a1cc5cd0-dee4-11ef-b8ca-171526dfb895.html [22] DeepSeek-V3 Technical Report — arXiv (DeepSeek-AI), 2024-12-26 (paper) https://arxiv.org/abs/2412.19437 [23] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — arXiv (DeepSeek-AI), 2025-01-22 (paper) https://arxiv.org/html/2501.12948v1 [24] DeepSeek-R1-Distill-Llama-70B model card — Hugging Face, 2025-01-20 (docs) https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B [25] China's DeepSeek shook the tech world. Its developer just revealed the cost — CNN, 2025-09-19 (news) https://www.cnn.com/2025/09/19/business/deepseek-ai-training-cost-china-intl [26] DeepSeek didn't really train its flagship model for $294,000 — The Register, 2025-09-19 (news) https://www.theregister.com/2025/09/19/deepseek_cost_train/ [27] DeepSeek Debates: Chinese Leadership On Cost, True Training Cost — SemiAnalysis, 2025-01-31 (blog) https://semianalysis.com/2025/01/31/deepseek-debates/ [28] Nvidia sheds almost $600 billion in market cap, biggest drop ever — CNBC, 2025-01-27 (news) https://www.cnbc.com/2025/01/27/nvidia-sheds-almost-600-billion-in-market-cap-biggest-drop-ever.html [29] Nvidia stock plummets, loses record $589 billion as DeepSeek prompts questions over AI spending — Yahoo Finance, 2025-01-27 (news) https://finance.yahoo.com/news/nvidia-stock-plummets-loses-record-589-billion-as-deepseek-prompts-questions-over-ai-spending-135105824.html [30] Nvidia market cap falls by $1 trillion as stock swoons — Yahoo Finance, 2026-07-08 (news) https://finance.yahoo.com/markets/stocks/articles/nvidia-stock-valuation-falls-pre-132505643.html [31] NVIDIA (NVDA) Market Cap & Net Worth — StockAnalysis, 2026-09-02 (filing) https://stockanalysis.com/stocks/nvda/market-cap/ [32] Will DeepSeek's new AI model crash Nvidia's $5tn party? — The National, 2026-04-25 (news) https://www.thenationalnews.com/future/technology/2026/04/25/will-deepseeks-new-ai-model-crash-nvidias-5tn-party/ [33] Why DeepSeek didn't cause an investor frenzy again in 2025 — CNBC, 2026-01-06 (news) https://www.cnbc.com/2026/01/06/why-deepseek-didnt-cause-an-investor-frenzy-again-in-2025.html [34] Microsoft Probing If DeepSeek-Linked Group Improperly Obtained OpenAI Data — Bloomberg, 2025-01-29 (news) https://www.bloomberg.com/news/articles/2025-01-29/microsoft-probing-if-deepseek-linked-group-improperly-obtained-openai-data [35] US House Select Committee Report Accuses DeepSeek of Spying and Circumventing Export Controls on Chips — Tech Policy Press, 2025-04-17 (news) https://www.techpolicy.press/us-house-select-committee-report-accuses-deepseek-of-spying-and-circumventing-export-controls-on-chips/ [36] OpenAI update to the US House Select Committee on the CCP — OpenAI, 2026-02-12 (filing) https://cdn.openai.com/pdf/045aa967-ee96-4a09-94ee-3098ddf6db2c/OpenAI-US-House-Select-Cmte-Update-%5B021226%5D.pdf [37] Issue Brief: Adversarial Distillation — Frontier Model Forum, 2026-02-23 (blog) https://www.frontiermodelforum.org/issue-briefs/issue-brief-adversarial-distillation/ [38] LLM inference prices have fallen rapidly but unequally across tasks — Epoch AI, 2025 (blog) https://epoch.ai/data-insights/llm-inference-price-trends [39] How much does it cost to train frontier AI models? — Epoch AI, 2024-06 (paper) https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models [40] Welcome to LLMflation - LLM inference cost is going down fast — Andreessen Horowitz, 2024-11 (blog) https://a16z.com/llmflation-llm-inference-cost/ [41] Stealing Part of a Production Language Model — arXiv (Carlini et al.), 2024-03-11 (paper) https://arxiv.org/pdf/2403.06634 [42] Stealing part of a production language model (ICML 2024) — PMLR, 2024 (paper) https://proceedings.mlr.press/v235/carlini24a.html [43] OpenAI announces 80% price drop for o3 — VentureBeat, 2025-06-10 (news) https://venturebeat.com/ai/openai-announces-80-price-drop-for-o3-its-most-powerful-reasoning-model [44] gpt-oss-120b & gpt-oss-20b Model Card — OpenAI, 2025-08-05 (docs) https://openai.com/index/gpt-oss-model-card/ [45] DeepSeek raises some V4 prices by more than 10x as AI demand strains capacity — InfoWorld, 2026-08-14 (news) https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html [46] Tech Brief: DeepSeek Launches V4-Pro and Raises API Prices by as Much as 1,100% — Caixin Global, 2026-08-14 (news) https://www.caixinglobal.com/2026-08-14/tech-brief-aug-14-deepseek-launches-v4-pro-and-raises-api-prices-by-as-much-as-1100-102474222.html [47] DeepSeek signals significant price hike amid surge in demand for low-cost AI models — South China Morning Post, 2026-08-06 (news) https://www.scmp.com/tech/tech-trends/article/3363129/deepseek-signals-significant-price-hike-amid-surge-demand-low-cost-ai-models [48] DeepSeek's new bargain model accelerates AI's race to zero — Axios, 2026-08-01 (news) https://www.axios.com/2026/08/01/deepseek-model-cheap-ai-price-war [49] Arcee AI secures $24M Series A to transform the landscape of small language models — Arcee AI, 2024-07 (blog) https://www.arcee.ai/blog/arcee-ai-secures-24m-series-a-to-transform-the-landscape-of-small-language-models [50] Small language models rising as Arcee AI lands $24M Series A — VentureBeat, 2024-07 (news) https://venturebeat.com/ai/small-language-models-rising-as-arcee-ai-lands-24m-series-a [51] Rubrik acquires Predibase to accelerate adoption of AI agents — TechCrunch, 2025-06-25 (news) https://techcrunch.com/2025/06/25/rubrik-acquires-predibase-to-accelerate-adoption-of-ai-agents/ [52] Rubrik pivots to generative AI with Predibase acquisition — TechTarget, 2025-06-26 (news) https://www.techtarget.com/searchdatabackup/news/366626870/Rubrik-pivots-to-generative-AI-with-Predibase-acquisition [53] Together AI raises $305M for its AI-optimized public cloud — SiliconANGLE, 2025-02-20 (news) https://siliconangle.com/2025/02/20/together-ai-raises-305m-ai-optimized-public-cloud/ [54] Neocloud Together AI raises $800M, leaps to $8.3B valuation — TechCrunch, 2026-07-01 (news) https://techcrunch.com/2026/07/01/neocloud-together-ai-raises-800m-leaps-to-8-3b-valuation/ [55] Mistral AI raises EUR 1.7B to accelerate technological progress with AI — Mistral AI, 2025-09-09 (blog) https://mistral.ai/news/mistral-ai-raises-1-7-b-to-accelerate-technological-progress-with-ai/ [56] AI firm Mistral valued at $14 billion as chip giant ASML takes major stake — CNBC, 2025-09-09 (news) https://www.cnbc.com/2025/09/09/ai-firm-mistral-valued-at-14-billion-as-chip-giant-asml-takes-major-stake.html [57] Mistral is rumored to be raising EUR 3B at EUR 20B valuation — TechCrunch, 2026-06-12 (news) https://techcrunch.com/2026/06/12/mistral-is-rumored-to-be-raising-e3b-at-e20-valuation/ [58] NVIDIA H100 GPU Pricing: 2026 Rent vs. Buy Cost Analysis — GMI Cloud, 2026 (pricing) https://www.gmicloud.ai/en/blog/nvidia-h100-gpu-pricing-2026-rent-vs-buy-cost-analysis [59] H100 Cloud Pricing: Compare 53+ Providers (2026) — GetDeploying, 2026-09-03 (pricing) https://getdeploying.com/gpus/nvidia-h100 [60] Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Serving Cost — arXiv, 2026-06 (paper) https://arxiv.org/html/2606.11690v1 [61] State of AI: An Empirical 100 Trillion Token Study with OpenRouter — arXiv / OpenRouter, 2026-01 (paper) https://arxiv.org/html/2601.10088v1 [62] OpenRouter State of AI — OpenRouter, 2026-01 (blog) https://openrouter.ai/state-of-ai [63] Small Language Model Market Report 2026 — Research and Markets, 2026 (news) https://www.researchandmarkets.com/reports/6076406/small-language-model-market-report ================================================================================================ 5. POLITICAL PERSPECTIVE — POLITICAL AND GEOPOLITICAL VIEW OF AI DISTILLATION ================================================================================================ Page: https://global-distillation.com/political Data: https://global-distillation.com/data/political.json Updated: 2026-09-04 Content: 8 key figures, 6 tables, 5 charts, 47 dated events, 16 glossary terms, 100 sources SUMMARY -------- Between January 2025 and September 2026, model distillation went from an obscure machine-learning technique to a named object of US national security policy. It began when OpenAI and Microsoft said they had evidence a DeepSeek-linked group had pulled data through OpenAI's API, and White House AI czar David Sacks said there was 'substantial evidence' DeepSeek had distilled OpenAI models. It escalated through OpenAI's March 2025 OSTP filing calling DeepSeek 'state-subsidized' and 'state-controlled,' Anthropic's September 2025 ban on entities majority-controlled from China, and a February 2026 wave of disclosures in which OpenAI, Google and Anthropic each published evidence of large-scale extraction campaigns. It became formal policy in April 2026 with OSTP memorandum NSTM-4 on 'adversarial distillation,' a State Department demarche cable, and H.R. 8283, which would create a public 'AI Model Extraction Attackers List' — followed by NSPM-11 in June, Anthropic's letter alleging a 28.8-million-exchange Alibaba campaign, and Treasury Secretary Bessent's July 2026 sanctions threat. Underneath the geopolitics sits an unresolved legal question: the strongest theory against distillation is breach of contract, not copyright or trade secret, and critics from ITIF to the Institute for Law & AI warn that building export controls and sanctions on top of private terms of service is a shaky foundation. KEY FIGURES ----------- - Claude exchanges in largest disclosed campaign: 28,800,000 exchanges (vs 16M disclosed in Feb 2026) Anthropic's June 10, 2026 letter to Senate Banking alleges Alibaba/Qwen-affiliated operators ran 28.8M+ exchanges between Apr 22 and Jun 5, 2026 Source: https://d1e00ek4ebabms.cloudfront.net/production/uploaded-files/Anthropic%20letter%20Alibaba%E3%83%BBJun%2010%202026%20-ba307cab-ccb4-43f1-88c0-515f29694678.pdf - Fraudulent accounts alleged (Alibaba campaign): 25,000 accounts (24,000 in the Feb 2026 DeepSeek/Moonshot/MiniMax disclosure) Anthropic's June 10, 2026 letter to the Senate Banking Committee alleges almost 25,000 fraudulent accounts violating its terms of service and regional access restrictions Source: https://d1e00ek4ebabms.cloudfront.net/production/uploaded-files/Anthropic%20letter%20Alibaba%E3%83%BBJun%2010%202026%20-ba307cab-ccb4-43f1-88c0-515f29694678.pdf - House Foreign Affairs vote on H.R. 8283: 43 yeas (43-0) (unanimous) Deterring American AI Model Theft Act of 2026 ordered reported April 22, 2026; not yet passed the full House as of Sept 2026 Source: https://www.congress.gov/119/meeting/house/119191/documents/HMKP-119-FA00-20260422-SD002.pdf - US states restricting DeepSeek on state devices: 14 states (0 before Jan 31, 2025) Texas, New York, Virginia, Iowa, South Dakota, North Carolina, Nebraska, Tennessee, Arkansas, North Dakota, Oklahoma, Alabama, Kansas and Georgia, as of April 2025; likely higher by Sept 2026 Source: https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek - Nvidia/AMD China revenue share to US government: 15 % of China chip revenue (new condition, Aug 2025) Reported arrangement tied to resumed H20 / MI308 export licences Source: https://fortune.com/2025/08/10/nvidia-amd-chips-h20-mi308-china-sales-revenue-trump-export-license/ - Section 232 tariff on advanced AI chips imported into the US (incl. those routed to China): 25 % (new; paired with BIS case-by-case review replacing presumption of denial) Proclamation signed Jan 14, 2026, effective Jan 15, 2026; separately BIS moved H200 / MI325X to case-by-case review for China and Macau Source: https://www.whitecase.com/insight-alert/president-trump-orders-narrowly-targeted-25-section-232-tariff-certain-advanced - Entities on Pentagon 1260H list after June 2026 update: 188 entities (+65 added June 2026) Alibaba, Baidu, BYD and Unitree added; DeepSeek notably not listed Source: https://www.cnbc.com/2026/06/09/alibaba-baidu-byd-named-on-pentagons-china-military-list-.html - Maximum EU AI Act fine for GPAI providers: 15,000,000 EUR or 3% of global turnover (enforceable from Aug 2, 2026) Whichever is higher; AI Office gained investigation and model-access powers on that date Source: https://ec.europa.eu/commission/presscorner/detail/en/ip_26_1714 KEY FINDINGS ------------ 1. Distillation is now a named category in US national security policy — but naming has not produced sanctions OSTP memorandum NSTM-4, 'Adversarial Distillation of American AI Models,' issued by OSTP under Director Michael Kratsios on April 23, 2026, is the first US policy instrument to formally classify systematic capability extraction from frontier models as a national security threat. It found that foreign entities, principally in China, are running 'deliberate, industrial-scale campaigns' using tens of thousands of proxy accounts, and committed the executive branch to threat-intelligence sharing with industry, joint defensive best practices, and exploration of accountability measures. NSPM-11, signed June 5, 2026, then directed the national security enterprise to help secure US models against distillation attacks. Naming has not yet become enforcement: no Chinese AI lab has been sanctioned or Entity Listed specifically for distillation as of early September 2026. In June 2026 the administration held off publishing an expanded blacklist covering DeepSeek and 100+ other flagged firms, reportedly to avoid escalating with Beijing, and DeepSeek was absent from the Pentagon's June 2026 1260H expansion that added Alibaba and Baidu. Treasury Secretary Bessent's July 21, 2026 warning that Washington could sanction overseas models found to have stolen from US firms remained a threat. Sources: https://www.nextgov.com/artificial-intelligence/2026/04/white-house-accuses-china-deliberate-industrial-scale-campaigns-steal-us-ai-models/413083/ https://www.whitehouse.gov/presidential-actions/2026/06/national-security-presidential-memorandum-nspm-11/ https://www.justsecurity.org/137498/diagnosis-deterrence-us-response-distillation/ https://www.cnbc.com/2026/06/17/us-deepseek-blacklist-cxmt-national-security-risks-.html https://www.cnbc.com/2026/07/21/bessent-china-ai-sanctions.html https://www.cnbc.com/2026/06/09/alibaba-baidu-byd-named-on-pentagons-china-military-list-.html 2. The legal core of the accusation is breach of contract, not IP theft — and that is a weaker peg than the rhetoric implies US copyright law does not protect purely machine-generated outputs, and OpenAI's own terms assign output rights to the user, so a copyright claim against a distiller is difficult. Trade secret theory is available but unsettled: it depends on whether querying an API counts as acquisition by 'improper means.' The theory that actually fits is contract — OpenAI's terms bar using Output 'to develop models that compete with OpenAI' — plus, where fraudulent accounts and evaded geo-restrictions are involved, the Computer Fraud and Abuse Act. H.R. 8283 mirrors this: its definition of a 'model extraction attack' turns on circumventing access controls, using fraudulent credentials, or violating terms of service. The symmetry argument has never gone away either: because copyright does not protect bare machine outputs and terms of service are private contracts, critics — including a Cornell tip sheet published within days of OpenAI's January 2025 statement — note that training on publishers' content also violated those publishers' terms. The asymmetry the labs rely on is jurisdictional and contractual rather than moral, which is why the fight migrated to export controls, sanctions and procurement bans rather than the courts. Sources: https://www.govinfo.gov/content/pkg/BILLS-119hr8283ih/pdf/BILLS-119hr8283ih.pdf https://law.asia/openai-deepseek-ai-distillation/ https://www.justsecurity.org/134124/costs-china-ai-distillation/ https://www.copyright.gov/newsnet/2025/1060.html https://news.cornell.edu/media-relations/tip-sheets/ironic-hypocritical-big-tech-call-out-deepseek https://futurism.com/openai-mockery-stole-work-deepseek 3. Building federal sanctions on top of private terms of service is the central critique of the 2026 bills ITIF's July 28, 2026 analysis of H.R. 8283 argues the bill leans too heavily on terms-of-service violations — private contracts that vary provider to provider — as the trigger for federal designation, and that an 'AI Model Extraction Attackers List' built on unverified corporate disclosures raises due-process concerns. It recommends narrowing the definition to intentional account fraud, requiring public evidentiary summaries, and adding safe harbours for open-source development, academic research and security testing. The Institute for Law & AI makes a parallel argument: policymakers should first ask how much distillation actually contributes to the capability gap before locking in restrictions. Sources: https://itif.org/publications/2026/07/28/how-to-fix-the-ai-model-theft-bill-before-it-becomes-law/ https://law-ai.org/responding-to-ai-distillation-without-panic/ https://www.lawfaremedia.org/article/responding-to-ai-distillation-without-panic 4. Disclosure moved from anecdote to numbers — and the numbers keep growing January 2025 accusations were qualitative: Microsoft security researchers observed suspected DeepSeek-linked individuals exfiltrating data via the OpenAI API, and David Sacks cited 'substantial evidence' without detailing it. By February 2026 the labs published counts: Anthropic attributed 150,000+ exchanges to DeepSeek, 3.4 million to Moonshot AI and 13 million to MiniMax across ~24,000 fraudulent accounts; Google's Threat Intelligence Group disrupted a cluster of 100,000+ prompts aimed at coercing Gemini reasoning traces. By June 2026 Anthropic put a single Alibaba-linked campaign at 28.8 million exchanges. The trend line of disclosed volume is the single most concrete input to the policy debate. Sources: https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://techinformed.com/google-disrupts-gemini-model-extraction-attempts/ https://d1e00ek4ebabms.cloudfront.net/production/uploaded-files/Anthropic%20letter%20Alibaba%E3%83%BBJun%2010%202026%20-ba307cab-ccb4-43f1-88c0-515f29694678.pdf https://cloud.google.com/blog/topics/threat-intelligence/distillation-experimentation-integration-ai-adversarial-use 5. Export controls and distillation policy are the same argument in two registers Anthropic's public position is that distillation 'reinforces the rationale for export controls,' because harvested exchanges are only useful if the distiller has compute to train on them. Dario Amodei made the compute-asymmetry argument in January 2025. The counter-current is commercial: the Biden AI Diffusion Rule (Jan 15, 2025) was rescinded before its May 15, 2025 enforcement date; the GAIN AI Act, which would have required US customers be served first, passed the Senate as an NDAA amendment in October 2025 but was dropped from the final FY2026 NDAA; and in January 2026 H200 and MI325X exports moved from presumption of denial to case-by-case review with a 25% tariff. Sources: https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://darioamodei.com/post/on-deepseek-and-export-controls https://www.wiley.law/alert-BIS-Rescinds-AI-Diffusion-Rule https://www.nextgov.com/policy/2025/12/bill-prioritizing-american-customers-ai-chips-not-expected-make-it-final-ndaa-sources-say/409920/ 6. Private corporate policy has become de facto foreign policy OpenAI blocked API traffic from unsupported regions including mainland China on July 9, 2024, then introduced Verified Organization ID checks for frontier API access in April 2025. Anthropic went furthest: on September 5, 2025 it barred service to companies more than 50% owned by entities in China, Russia, Iran and North Korea regardless of where those subsidiaries operate, accepting a revenue hit it described as in the low hundreds of millions of dollars. These access rules, not statutes, are what H.R. 8283 would convert into a designation trigger — which is precisely why critics call the mechanism circular. Sources: https://www.semafor.com/article/09/05/2025/anthropic-blocks-ai-sales-in-china https://www.rte.ie/news/business/2025/0905/1531954-us-ai-giant-anthropic-bars-chinese-owned-entities/ https://help.openai.com/en/articles/10910291-api-organization-verification https://restofworld.org/2024/exporter-openai-china-api-access/ 7. Evidence quality is contested, and at least one White House accusation drew expert pushback On July 22, 2026 OSTP Director Kratsios accused Moonshot AI of distilling Anthropic's Fable to build Kimi K3 and of obtaining export-controlled Nvidia servers. Researchers including Nathan Lambert (Allen Institute for AI) and Braden Hancock (Laude Institute) publicly doubted that distillation alone could explain K3's capabilities on that timeline — Fable had been publicly available only since July 1 — and no supporting evidence was published. Earlier, US officials alleged DeepSeek trained V4 on smuggled Blackwell GPUs; Nvidia called the claim farfetched and DeepSeek's V4 preview shipped optimised for Huawei Ascend silicon. Accusation-without-published-evidence is a recurring pattern. Sources: https://cyberscoop.com/white-house-accuses-moonshot-ai-anthropic-model-distillation/ https://techcrunch.com/2026/07/23/experts-say-exploiting-anthropics-fable-isnt-how-kimi-k3-got-so-good/ https://www.technologyreview.com/2026/04/24/1136422/why-deepseeks-v4-matters/ 8. China's response has moved from denial to mirror-imaging The Chinese Embassy in Washington called US allegations groundless and framed them as attacks on China's AI development. On July 27, 2026 the Ministry of Commerce went further, accusing 'many American AI enterprises' of distilling Chinese models — naming none and supplying no evidence — calling US actions 'double standards' and 'AI hegemony,' and promising 'all necessary measures' if Chinese firms were sanctioned. Separately, FT reported that MOFCOM was consulting Alibaba, ByteDance and Zhipu on adding model weights, key training data and chip designs to China's technology export catalogue, which would make Chinese open weights themselves a licensed export. Sources: https://www.implicator.ai/china-says-us-firms-distilled-chinese-models/ https://www.techrepublic.com/article/news-apac-china-ai-model-export-controls/ https://www.cnbc.com/2026/04/25/us-global-warning-alleged-china-ai-theft.html 9. The EU regulates capability, not provenance — which leaves distillation largely untouched EU AI Act GPAI obligations began applying August 2, 2025, backed by a Code of Practice published July 10, 2025 with Transparency, Copyright, and Safety & Security chapters; Commission enforcement powers, including fines up to EUR 15 million or 3% of global turnover, went live August 2, 2026. None of this creates a distillation-specific offence. Europe's actual friction with Chinese models has run through GDPR instead — Italy's Garante ordered DeepSeek's chatbot blocked in early 2025, and Berlin's commissioner found its transfer practices unlawful — while self-hosted open weights on EU servers sidestep the transfer question entirely. Sources: https://artificialintelligenceact.eu/code-of-practice-overview/ https://ec.europa.eu/commission/presscorner/detail/en/ip_26_1714 https://www.privacylaws.com/news/italy-and-south-korea-ban-deepseek-and-start-investigation/ https://www.pinsentmasons.com/out-law/analysis/eu-ai-act-gpai-deepseek-review 10. A grey market for API access is the practical enforcement problem Extraction at the scale the labs describe requires access the labs have formally denied. Reporting on China's 'transfer station' economy describes tens of thousands of internet-facing servers running reseller billing panels that proxy OpenAI, Anthropic, Google and other Western models into China, sometimes at roughly a tenth of list price. Anthropic says a single proxy network managed more than 20,000 fraudulent accounts. H.R. 8283 responds by defining a 'fraudulent account network provider' as a designation target in its own right — with a carve-out for services that enable internet access for freedom of expression. Sources: https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens-in https://www.deeplearning.ai/the-batch/inside-the-gray-market-for-llm-access https://www.govinfo.gov/content/pkg/BILLS-119hr8283ih/pdf/BILLS-119hr8283ih.pdf TABLES -------- TABLE: Policy matrix: how each jurisdiction treats distillation and Chinese models [id: policy-matrix-jurisdiction, 12 rows] Binding instruments, government-device restrictions and the specific stance on model extraction, as of September 2026. Jurisdiction | Primary instrument | Status | Key date | Distillation stance | Chinese-model stance | Source URL ------------ | ------------------ | ------ | -------- | ------------------- | -------------------- | ---------- United States (executive) | OSTP NSTM-4; NSPM-11; AI Action Plan | In force | 2026-04-23 | Named national security threat; intel sharing with labs; accountability measures 'explored' | Federal device bans proposed; export controls; no model sanctions yet | https://www.nextgov.com/artificial-intelligence/2026/04/white-house-accuses-china-deliberate-industrial-scale-campaigns-steal-us-ai-models/413083/ United States (Congress) | H.R. 8283 Deterring American AI Model Theft Act | Reported by committee 43-0; not enacted | 2026-04-22 | Would create public 'AI Model Extraction Attackers List' + IEEPA/Entity List authorities | PRC, Hong Kong, Macau and Russia are 'countries of concern' by statute | https://www.congress.gov/119/meeting/house/119191/documents/HMKP-119-FA00-20260422-SD002.pdf United States (states) | Executive directives; CA SB 53; NY RAISE Act | In force | 2026-01-01 | No state distillation offence; frontier-model transparency only | 7+ states bar DeepSeek on state devices and networks | https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek European Union | AI Act GPAI obligations + Code of Practice | Applying; enforcement powers live | 2026-08-02 | No distillation-specific rule; downstream fine-tuners can become providers | Capability-based, provenance-neutral; friction runs through GDPR | https://ec.europa.eu/commission/presscorner/detail/en/ip_26_1714 Italy | Garante order under GDPR | In force | 2025-01-30 | Not addressed | DeepSeek chatbot blocked for the general public, not just government | https://www.privacylaws.com/news/italy-and-south-korea-ban-deepseek-and-start-investigation/ United Kingdom | Sovereign AI Unit (DSIT); AI Security Institute | Non-statutory | 2025 | No dedicated instrument; treated as a security-research question | No ban; AISI evaluations flag DeepSeek jailbreak and censorship behaviour | https://oecd.ai/en/dashboards/policy-initiatives/uk-sovereign-ai-unit China | Global AI Governance Action Plan; WAICO; export catalogue review | Announced / under consultation | 2025-07-26 | Defends distillation as an industry-wide technique; counter-accuses US firms | Promotes open-weight release; weighing export controls on weights and training data | https://www.techrepublic.com/article/news-apac-china-ai-model-export-controls/ South Korea | AI Basic Act (Framework Act) | In force | 2026-01-22 | Not addressed | PIPC suspended DeepSeek app downloads Feb 2025 pending compliance | https://www.cooley.com/news/insight/2026/2026-01-27-south-koreas-ai-basic-act-overview-and-key-takeaways Japan | AI Promotion Act (May 2025) | In force; promotional, light-touch | 2025-05 | Not addressed | No ban; policy focus on domestic R&D capacity and competitiveness | https://www.whitecase.com/insight-our-thinking/ai-watch-global-regulatory-tracker-japan Australia | Government device directive | In force | 2025-02-04 | Not addressed | DeepSeek prohibited on government devices | https://www.privacylaws.com/news/italy-and-south-korea-ban-deepseek-and-start-investigation/ Taiwan | Government agency guidance | In force | 2025-02 | Not addressed | DeepSeek restricted in public sector over cross-border data transfer risk | https://tech.co/news/which-countries-have-banned-deepseek-already Gulf states (UAE, Saudi Arabia) | Bilateral compute and security agreements | Negotiated, deal-by-deal | 2026 | Not addressed directly; governed via US access conditions | UAE aligned G42 away from Chinese tech; Saudi retains Huawei links | https://www.iiss.org/publications/strategic-comments/2026/06/gulf-ai-infrastructure-and-the-limits-of-technological-sovereignty/ Notes: Coding reflects publicly reported instruments only. 'Not addressed' means no distillation-specific rule, not that generic IP or computer-misuse law is unavailable. Sources: https://www.congress.gov/bill/119th-congress/house-bill/8283/text https://ec.europa.eu/commission/presscorner/detail/en/ip_26_1714 https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek https://www.congress.gov/119/meeting/house/119191/documents/HMKP-119-FA00-20260422-SD002.pdf TABLE: US federal bills touching distillation, Chinese AI and chip flows (119th Congress) [id: us-bills-tracker, 8 rows] Every bill tracked here is from the 119th Congress (2025-2026). Status as of September 2026. Bill | Number | Lead sponsor | Introduced | Furthest stage | Distillation relevance | Source URL ---- | ------ | ------------ | ---------- | -------------- | ---------------------- | ---------- Deterring American AI Model Theft Act of 2026 | H.R. 8283 | Rep. Huizenga (R-MI) | 2026-04-15 | Reported by House Foreign Affairs 43-0 | Direct: defines 'model extraction attack', creates public attackers list, authorises IEEPA sanctions and Entity Listing | https://www.congress.gov/bill/119th-congress/house-bill/8283/text Decoupling America's AI Capabilities from China Act | S. 321 | Sen. Hawley (R-MO) | 2025-01-29 | Referred to Judiciary | Indirect: would bar import/export of AI tech and IP to/from China; penalties up to 20 years | https://www.congress.gov/bill/119th-congress/senate-bill/321 No DeepSeek on Government Devices Act | H.R. 1121 | Rep. Gottheimer (D-NJ) | 2025-02-07 | Referred to committee | Indirect: federal device ban aimed at the model most associated with distillation claims | https://www.congress.gov/bill/119th-congress/house-bill/1121/all-info No Adversarial AI Act | S. 2177 / H.R. 4142 | Sen. Scott (R-FL) [S. 2177]; Rep. Moolenaar (R-MI) [H.R. 4142] | 2025-06-25 | Referred to committee | Indirect: FASC list of foreign-adversary AI; bans agency use with narrow research carve-outs | https://www.congress.gov/bill/119th-congress/senate-bill/2177/text Chip Security Act (Senate) | S. 1705 | Sen. Cotton (R-AR) | 2025-05-08 | Referred to Banking | Upstream: location verification on exported AI chips limits compute available for distillation training | https://www.congress.gov/bill/119th-congress/senate-bill/1705 Chip Security Act (House) | H.R. 3447 | Rep. Huizenga (R-MI) | 2025-05-15 | Reported by House Foreign Affairs 42-0 (2026-03-26) | Upstream: same location-verification mandate; not enacted | https://www.congress.gov/bill/119th-congress/house-bill/3447 GAIN AI Act of 2025 | S. 3150 (also H.R. 5885) | Sen. Banks (R-IN) | 2025-11-06 | Passed Senate as NDAA amendment; dropped from final FY26 NDAA | Upstream: would require US customers be prioritised before advanced chip sales abroad | https://www.congress.gov/bill/119th-congress/senate-bill/3150 AI Security and Innovation Act | H.R. 9363 | Rep. Obernolte (R-CA) | 2026-06-18 | Reported by House Science, Space and Technology 29-0 (2026-06-25) | Peripheral: establishes an AI evaluation/security center under the National AI Initiative Act; no distillation provisions | https://science.house.gov/2026/6/h-r-9363-ai-security-and-innovation-act Notes: GAIN AI Act status confirmed by reporting that the final FY2026 NDAA, signed 2025-12-18, excluded it. Sources: https://www.congress.gov/bill/119th-congress/house-bill/8283/text https://www.nextgov.com/policy/2025/12/bill-prioritizing-american-customers-ai-chips-not-expected-make-it-final-ndaa-sources-say/409920/ https://www.congress.gov/bill/119th-congress/house-bill/3447 https://science.house.gov/2026/6/h-r-9363-ai-security-and-innovation-act https://www.govinfo.gov/app/details/BILLS-119hr4142ih TABLE: Terms-of-service and licence clauses across labs: can you train on the outputs? [id: tos-clause-comparison, 9 rows] The contractual layer that US policy now treats as a designation trigger. Wording summarised, not quoted in full. Provider / model family | Access model | Clause on training competing models | Who owns outputs | Jurisdictional restriction | Source URL ----------------------- | ------------ | ------------------------ | ---------------- | ------------------------ | ---------- OpenAI | Closed API + apps | Prohibits using Output to develop models that compete with OpenAI | Assigned to the user | API traffic blocked from unsupported regions incl. mainland China since 2024-07-09; Verified Organization ID checks for frontier models since 2025-04 | https://openai.com/policies/row-terms-of-use/ Anthropic (Claude) | Closed API + apps | Prohibits using the Services to develop competing products, including to train any AI/ML models | Anthropic assigns its rights, if any, in Outputs to the user | Since 2025-09-05 no service to entities >50% owned from China, Russia, Iran, North Korea, worldwide | https://www.anthropic.com/legal/commercial-terms Google (Gemini) | Closed API + apps | Prohibits using outputs to develop competing models | User-facing rights per service terms | Regional availability limits; GTIG disrupted extraction clusters in Feb 2026 | https://techinformed.com/google-disrupts-gemini-model-extraction-attempts/ Meta Llama 2 / Llama 3 | Open weights, community licence | Prohibited using Llama materials or outputs to improve any other LLM | Licensee | Acceptable use policy only | https://www.llama.com/llama3/license/ Meta Llama 3.1 and later | Open weights, community licence | Permitted: outputs may be used for synthetic data generation and distillation with attribution | Licensee | Acceptable use policy only | https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct xAI | Closed API + apps | Restricts competitive training; distillation risk addressed in its Risk Management Framework (2025-08-20) | Per service terms | Regional availability limits | https://docs.house.gov/meetings/ZS/ZS00/20260416/119165/HHRG-119-ZS00-Wstate-MahmoodY-20260416.pdf DeepSeek | Open weights + hosted API | Permissive open-weight licensing; V3/R1 weights on Hugging Face | Licensee | Hosted service blocked or restricted in Italy, South Korea, Australia, Taiwan and 7+ US states | https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek Alibaba (Qwen) | Open weights + hosted API | Permissive open-weight licensing | Licensee | Alibaba added to Pentagon 1260H list 2026-06; barred from Anthropic services under the 2025 China policy | https://www.cnbc.com/2026/06/09/alibaba-baidu-byd-named-on-pentagons-china-military-list-.html Moonshot AI (Kimi) | Open weights + hosted API | Permissive open-weight licensing | Licensee | Named in Anthropic Feb 2026 disclosure and State Department April 2026 cable | https://www.cnbc.com/2026/04/25/us-global-warning-alleged-china-ai-theft.html Notes: The asymmetry is structural: closed US labs restrict output-based training by contract, while the leading Chinese labs release weights under permissive licences. That is why an extraction-attack statute keyed to terms of service applies in one direction only. Sources: https://openai.com/policies/row-terms-of-use/ https://www.anthropic.com/legal/consumer-terms https://www.llama.com/llama3/license/ https://www.anthropic.com/legal/commercial-terms https://www.anthropic.com/legal/usage-policy TABLE: Legal theories for attacking distillation, and how strong each is [id: legal-theories, 8 rows] Strength ratings reflect the weight of published legal commentary cited, not a court ruling — no distillation case has been litigated to judgment. Theory | Source of law | Who can bring it | Strength | Principal weakness | Source URL ------ | ------------- | ---------------- | -------- | ------------------ | ---------- Breach of contract (terms of service) | State contract law | Model owner against the account holder | Strongest | Standard-form contract enforceability; jurisdiction and enforcement against foreign entities; privity if a proxy reseller holds the account | https://law.asia/openai-deepseek-ai-distillation/ Computer Fraud and Abuse Act | 18 U.S.C. 1030 | DOJ; private civil action | Strong where fraudulent credentials used | Post-Van Buren narrowing of 'exceeds authorized access'; requires proving the credentials were fraudulent, not merely ToS-violating | https://www.justsecurity.org/134124/costs-china-ai-distillation/ Trade secret misappropriation | Defend Trade Secrets Act; state UTSA | Model owner | Contested | Outputs served to any paying user are hard to characterise as secret; turns on whether API querying is 'improper means' | https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5121745 Economic Espionage Act | 18 U.S.C. 1831-1839 | DOJ | Weak in practice | Same secrecy problem as DTSA, plus foreign-defendant enforcement; commentators call it unsettled footing | https://www.justsecurity.org/134124/costs-china-ai-distillation/ Copyright infringement in outputs | 17 U.S.C. | Model owner | Weak | US Copyright Office and DC Circuit hold purely AI-generated outputs are not copyrightable; OpenAI assigns output rights to the user anyway | https://www.copyright.gov/newsnet/2025/1060.html Unfair competition / unjust enrichment | State law; Lanham Act adjacent | Model owner | Weak to moderate | Hard to establish deception where data was obtained through a public paid API rather than intrusion | https://law.asia/openai-deepseek-ai-distillation/ Export control / sanctions designation | ECRA, EAR Entity List, IEEPA | US government | Most likely operative route | Political, not judicial; requires attribution evidence agencies may not want to publish; escalation risk with Beijing | https://www.justsecurity.org/134124/costs-china-ai-distillation/ Statutory model-extraction designation (proposed) | H.R. 8283 if enacted | State Dept / Commerce | Untested | Anchored to private terms of service; ITIF flags due-process concerns and lack of research safe harbours | https://itif.org/publications/2026/07/28/how-to-fix-the-ai-model-theft-bill-before-it-becomes-law/ Notes: Commentators consulted include Joe Khawam (Law Reform Institute) on national security authorities, Camilla Hrdy on trade secrecy and generative AI, and Bahrad A. Sokhansanj (Institute for Law & AI) on proportionality. Sources: https://www.justsecurity.org/134124/costs-china-ai-distillation/ https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5121745 https://law-ai.org/responding-to-ai-distillation-without-panic/ https://www.copyright.gov/newsnet/2025/1060.html TABLE: US export-control milestones that frame the distillation debate [id: export-control-ledger, 12 rows] Compute access is the other half of the argument: distilled data is only useful with chips to train on. Date | Action | Scope | Effect / response | Source URL ---- | ------ | ----- | ----------------- | ---------- 2022-10-07 | BIS advanced computing and semiconductor rule | A100/H100-class GPUs to China | Nvidia introduced China-specific A800/H800 with reduced interconnect | https://www.congress.gov/crs-product/R48642 2023-10-17 | BIS October 2023 update | Captures A800/H800 and similar workarounds | Nvidia announced H20, L20, L2 for China | https://cset.georgetown.edu/article/bis-2023-update-explainer/ 2025-01-15 | AI Diffusion Rule published | Worldwide tiered licensing for advanced computing | Enforcement set for 2025-05-15; triggered allied and industry objections | https://www.wiley.law/alert-BIS-Rescinds-AI-Diffusion-Rule 2025-04 | H20 licence requirement imposed | Nvidia H20 to China | Nvidia forecast a $5.5bn charge in April 2025 and recorded $4.5bn in Q1 FY2026 | https://techcrunch.com/2025/05/28/nvidia-expects-to-lose-billions-in-revenue-due-to-h20-chip-licensing-requirements/ 2025-05-13 | BIS announces rescission of the AI Diffusion Rule | Global framework withdrawn before enforcement | Non-enforcement instruction pending formal rescission; replacement rule promised | https://www.wiley.law/alert-BIS-Rescinds-AI-Diffusion-Rule 2025-08 | H20 / MI308 licences resume with revenue-share arrangement | Nvidia and AMD China sales | Reported 15% of China chip revenue to the US government | https://fortune.com/2025/08/10/nvidia-amd-chips-h20-mi308-china-sales-revenue-trump-export-license/ 2025-10-09 | Senate passes NDAA including GAIN AI Act | US-customer-first allocation of advanced chips | Opposed by Nvidia, SIA and the White House AI adviser | https://www.nextgov.com/artificial-intelligence/2025/10/ai-export-control-bill-passes-senate-ndaa-amendment/408762/ 2025-12-18 | FY2026 NDAA signed without the GAIN AI Act | Chip allocation mandate dropped | Removed the main statutory brake on advanced chip exports | https://www.nextgov.com/policy/2025/12/bill-prioritizing-american-customers-ai-chips-not-expected-make-it-final-ndaa-sources-say/409920/ 2026-01-15 | BIS moves H200 / MI325X to case-by-case review | China and Macau destinations | Presumption of denial replaced; 25% tariff proclamation signed 2026-01-14 | https://www.cnbc.com/2026/01/14/trump-nvidia-h200-china-ai-chips.html 2026-03-26 | Chip Security Act reported out of House Foreign Affairs 42-0 | Location verification for exported AI chips | Bipartisan support; not enacted as of Sept 2026 | https://www.congress.gov/bill/119th-congress/house-bill/3447 2026-06-17 | Expanded Entity List publication held back | DeepSeek, CXMT and 100+ flagged firms | Reported delay to avoid escalation ahead of talks; DeepSeek remained unlisted | https://www.cnbc.com/2026/06/17/us-deepseek-blacklist-cxmt-national-security-risks-.html 2026-07-21 | China consults industry on AI export controls | Model weights, key training data, chip designs | Would make Chinese open weights a licensed export; still under review | https://www.techrepublic.com/article/news-apac-china-ai-model-export-controls/ Notes: Nvidia flagged an anticipated $5.5bn H20 charge in its April 15, 2025 8-K; the charge actually recorded in its Q1 FY2026 results (May 28, 2026 reporting) was $4.5bn. The 15% revenue-share figure is press-reported and has not been published as a formal rule. Sources: https://www.congress.gov/crs-product/R48642 https://www.wiley.law/alert-BIS-Rescinds-AI-Diffusion-Rule https://www.cnbc.com/2026/01/14/trump-nvidia-h200-china-ai-chips.html https://techcrunch.com/2025/05/28/nvidia-expects-to-lose-billions-in-revenue-due-to-h20-chip-licensing-requirements/ https://www.hpcwire.com/off-the-wire/nvidia-announces-financial-results-for-1st-quarter-fiscal-2026/ TABLE: Government restrictions on DeepSeek, by jurisdiction [id: restrictions-on-deepseek, 13 rows] The first Chinese model to be treated as a national security object rather than a product. Jurisdiction | Date | Scope | Stated rationale | Source URL ------------ | ---- | ----- | ---------------- | ---------- Italy | 2025-01-30 | Public block of the chatbot nationwide | Garante found privacy-policy and data-transfer disclosures inadequate | https://www.privacylaws.com/news/italy-and-south-korea-ban-deepseek-and-start-investigation/ Texas | 2025-01-31 | All state-owned devices | First US state ban; data harvesting and CCP-linkage concerns | https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek Taiwan | 2025-02-02 | Public sector agencies | Cross-border data transmission and leakage risk | https://www.taipeitimes.com/News/front/archives/2025/02/02/2003831193 Australia | 2025-02-04 | Government devices | Security concerns | https://www.thecable.ng/south-korea-joins-italy-australia-in-banning-deepseek-over-security-concerns/ New York State | 2025-02-10 | Government networks and devices | Foreign surveillance and censorship risk | https://www.nbcnews.com/tech/new-york-state-bans-deepseek-government-devices-rcna191510 Virginia | 2025-02-11 | State devices and networks | Third US state to act | https://natlawreview.com/article/three-states-ban-deepseek-use-state-devices-and-networks South Korea | 2025-02-17 | App-store downloads suspended | PIPC found non-compliance with Korean data protection law | https://www.privacylaws.com/news/italy-and-south-korea-ban-deepseek-and-start-investigation/ Iowa | 2025-02-19 | State devices, alongside other Chinese apps | Governor's directive | https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek South Dakota | 2025-03 | Government-issued devices and contractors | Bundled with RedNote restriction | https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek Oklahoma | 2025-03-21 | All state-owned devices | Governor Stitt executive action citing data security | https://oklahoma.gov/governor/newsroom/newsroom/2025/-governor-stitt-bans-deepseek-on-all-state-owned-devices-due-to-.html North Carolina | 2025-03 | State devices | Followed peer states | https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek US Commerce Department | 2025-02 | Department devices | Reported internal prohibition ahead of any statute | https://www.pymnts.com/artificial-intelligence-2/2025/report-commerce-department-bans-use-of-deepseek-on-government-devices Germany (Berlin DPA) | 2025-06-27 | Finding of unlawful processing under GDPR | Could not demonstrate EU-equivalent protection for data transferred to China | https://www.datenschutz-berlin.de/fileadmin/user_upload/pdf/pressemitteilungen/2025/20250627-BlnBDI-Press-Release_DeepSeek.pdf Notes: Dates for South Dakota and North Carolina are month-level in the underlying reporting. This is a floor count of publicly reported restrictions, not an exhaustive census; StateTech lists 14 US states restricting DeepSeek as of April 2025. Sources: https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek https://www.privacylaws.com/news/italy-and-south-korea-ban-deepseek-and-start-investigation/ https://tech.co/news/which-countries-have-banned-deepseek-already https://www.datenschutz-berlin.de/fileadmin/user_upload/pdf/pressemitteilungen/2025/20250627-BlnBDI-Press-Release_DeepSeek.pdf https://www.taipeitimes.com/News/front/archives/2025/02/02/2003831193 CHART DATA ---------- CHART: Distillation-adjacent policy and enforcement actions per quarter, 2022-2026 [id: policy-actions-per-quarter, type: bar, unit: actions] Quarter | Events catalogued in this dashboard's timeline (actions) 2022 Q4 | 1 2023 Q4 | 1 2024 Q3 | 1 2024 Q4 | 1 2025 Q1 | 9 2025 Q2 | 4 2025 Q3 | 6 2025 Q4 | 2 2026 Q1 | 5 2026 Q2 | 11 2026 Q3 | 6 Notes: Counts are computed strictly from the entries in this dashboard's own timeline[] — one point per event, no other inclusion rule — so every bar can be reproduced by filtering the timeline by quarter. Not an exhaustive census of AI policy activity. Quarters with zero catalogued events are omitted. 2026 Q3 runs only to September 4, 2026. Sources: https://www.justsecurity.org/137498/diagnosis-deterrence-us-response-distillation/ https://www.congress.gov/bill/119th-congress/house-bill/8283/text CHART: Jurisdiction stance comparison: which policy tools are actually in place [id: jurisdiction-policy-mix, type: stackedBar, unit: binary indicator] Jurisdiction | Distillation-specific policy instrument (binary indicator) United States | 1 European Union | 0 United Kingdom | 0 China | 0 South Korea | 0 Japan | 0 Australia | 0 Taiwan | 0 Jurisdiction | Binding law on general-purpose / frontier AI (binary indicator) United States | 0 European Union | 1 United Kingdom | 0 China | 1 South Korea | 1 Japan | 0 Australia | 0 Taiwan | 0 Jurisdiction | Restricts Chinese AI apps on government devices (binary indicator) United States | 1 European Union | 0 United Kingdom | 0 China | 0 South Korea | 1 Japan | 0 Australia | 1 Taiwan | 1 Jurisdiction | Unilateral controls on advanced AI chip / tech exports (binary indicator) United States | 1 European Union | 0 United Kingdom | 0 China | 1 South Korea | 0 Japan | 1 Australia | 0 Taiwan | 1 Notes: US row scores 0 on binding GPAI law at the federal level; California SB 53 and the New York RAISE Act are state instruments. EU scores 0 on export controls because those are member-state and Wassenaar instruments (e.g. the Netherlands), not bloc-level AI chip controls. China's binding-law score reflects its generative AI and labelling measures; its export-control score reflects existing materials controls plus the 2026 consultation on model weights. Sources: https://ec.europa.eu/commission/presscorner/detail/en/ip_26_1714 https://www.cooley.com/news/insight/2026/2026-01-27-south-koreas-ai-basic-act-overview-and-key-takeaways https://www.techrepublic.com/article/news-apac-china-ai-model-export-controls/ https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek CHART: Exchanges with Claude attributed to each accused lab, as disclosed by Anthropic [id: disclosed-extraction-volume, type: bar, unit: exchanges] Accused lab | February 23, 2026 disclosure (exchanges) DeepSeek | 150000 Moonshot AI | 3400000 MiniMax | 13000000 Accused lab | June 10, 2026 letter to Senate Banking (exchanges) Alibaba / Qwen-affiliated operators | 28800000 Notes: Figures are Anthropic's own attributions and have not been independently verified or adjudicated. The February set totals over 16 million exchanges across approximately 24,000 fraudulent accounts; the Alibaba campaign is dated April 22 to June 5, 2026 and used nearly 25,000 accounts. Alibaba denies using proprietary model outputs to train its models. Sources: https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://d1e00ek4ebabms.cloudfront.net/production/uploaded-files/Anthropic%20letter%20Alibaba%E3%83%BBJun%2010%202026%20-ba307cab-ccb4-43f1-88c0-515f29694678.pdf https://www.cnbc.com/2026/06/24/anthropic-alibaba-distillation-campaign.html CHART: How far each US bill has actually travelled [id: bill-progress, type: bar, unit: stage] Bill | Furthest stage reached as of 2026-09-04 (stage) H.R. 8283 Model Theft | 2 H.R. 3447 Chip Security | 2 S. 1705 Chip Security | 1 S. 3150 GAIN AI | 3 S. 321 Decoupling | 1 S. 2177 No Adversarial AI | 1 H.R. 4142 No Adversarial AI | 1 H.R. 1121 No DeepSeek on Gov Devices | 1 Notes: S. 3150 scores 3 because the GAIN AI Act passed the Senate as an amendment to the FY2026 NDAA on October 9, 2025 — but it was stripped in conference and the enacted NDAA excludes it, so no bill in this set has reached stage 4. Sources: https://www.congress.gov/bill/119th-congress/house-bill/8283/text https://www.congress.gov/bill/119th-congress/house-bill/3447 https://www.nextgov.com/policy/2025/12/bill-prioritizing-american-customers-ai-chips-not-expected-make-it-final-ndaa-sources-say/409920/ CHART: Cumulative restrictions on DeepSeek among the jurisdictions this dashboard tracks, first quarter after R1 [id: deepseek-restrictions-cumulative, type: line, unit: jurisdictions] Month | Cumulative jurisdictions listed in this dashboard (states, national governments, agencies) (jurisdictions) 2025-01 | 2 2025-02 | 9 2025-03 | 12 2025-04 | 12 Notes: Counts only the restrictions listed in the 'Government restrictions on DeepSeek' table: Italy and Texas in January; Taiwan, Australia, New York, Virginia, South Korea, Iowa and the US Commerce Department in February; South Dakota, Oklahoma and North Carolina in March; no further additions in April within this table. This is deliberately narrower than the full picture — StateTech lists 14 US states restricting DeepSeek by April 2025, including Nebraska, Tennessee, Arkansas, North Dakota, Alabama, Georgia and Kansas, which this table does not enumerate individually. Sources: https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek https://www.privacylaws.com/news/italy-and-south-korea-ban-deepseek-and-start-investigation/ DETAIL RECORDS -------------- POLICY INSTRUMENTS (25) - United States — OSTP Memorandum NSTM-4, Adversarial Distillation of American AI Models [In force, 2026-04-23] What it says: Memorandum issued by OSTP under Director Michael Kratsios, finding that foreign entities principally based in China are running deliberate, industrial-scale campaigns to distill US frontier AI systems using tens of thousands of proxy accounts and jailbreaking techniques, and that the resulting models deliberately strip security protocols. Distillation relevance: The first US policy instrument to name distillation as a national security threat. Commits agencies to intelligence sharing with AI companies, joint defensive best practices, and exploring accountability measures. Source: https://www.nextgov.com/artificial-intelligence/2026/04/white-house-accuses-china-deliberate-industrial-scale-campaigns-steal-us-ai-models/413083/ - United States — National Security Presidential Memorandum NSPM-11, AI in the National Security Enterprise [In force, 2026-06-05] What it says: Organises national security AI policy around adoption, adaptation, assurance and accountability; establishes governance guardrails, a talent reserve, and controls preventing deployed national-security AI from being disabled without federal approval. Rescinds and replaces NSM-25. Distillation relevance: Section 4(c) directs leaders to secure advanced AI systems including against malicious distillation attacks, and to partner with private-sector AI companies through threat-intelligence sharing and joint red-teaming. Source: https://www.whitehouse.gov/presidential-actions/2026/06/national-security-presidential-memorandum-nspm-11/ - United States — H.R. 8283, Deterring American AI Model Theft Act of 2026 [Reported by House Foreign Affairs 43-0; not enacted, 2026-04-15] What it says: Defines 'model extraction attack', 'closed-source AI model', 'entity of concern', 'country of concern' (PRC including Hong Kong and Macau, Russia, and designated Country Group D:5 states) and 'fraudulent account network provider'. Requires an executive assessment within 180 days, creates a public AI Model Extraction Attackers List maintained by the Secretary of State, and authorises IEEPA sanctions and Entity List designation. Distillation relevance: The core proposed statute. Its Sense of Congress expressly preserves authorized model training that adheres to terms of service as legitimate research, distinguishing it from extraction attacks. Source: https://www.govinfo.gov/content/pkg/BILLS-119hr8283ih/pdf/BILLS-119hr8283ih.pdf - United States — Winning the Race: America's AI Action Plan [In force, 2025-07-23] What it says: More than 90 federal actions across innovation, infrastructure and international pillars, including promotion of open-source and open-weight models and export of the full American AI stack to allied countries. Distillation relevance: Called for an industry information-sharing centre that became the channel for the April 2026 Frontier Model Forum distillation intelligence sharing. Its open-weights promotion also sits in tension with restricting capability diffusion. Source: https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf - United States — AI Diffusion Rule (Framework for Artificial Intelligence Diffusion) [Rescinded before enforcement, 2025-01-15] What it says: Worldwide tiered licensing framework for advanced computing and model weights, due to be enforced from May 15, 2025. Distillation relevance: OpenAI's March 2025 OSTP filing proposed banning PRC-produced models across the rule's Tier 1 country group. BIS announced rescission on May 13, 2025 and instructed non-enforcement pending formalisation. Source: https://www.wiley.law/alert-BIS-Rescinds-AI-Diffusion-Rule - United States — BIS advanced computing export controls (Oct 2022 and Oct 2023) [In force as amended, 2022-10-07] What it says: Restricts export of high-end AI accelerators to China; the October 2023 update captured the China-specific A800 and H800 designed around the original thresholds. Distillation relevance: Establishes the compute asymmetry the distillation argument depends on: harvested exchanges only convert into capability if the recipient has training compute. Source: https://www.congress.gov/crs-product/R48642 - United States — H200 / MI325X case-by-case licensing plus 25% tariff [In force, 2026-01-15] What it says: Replaces presumption of denial with case-by-case review for China and Macau destinations, subject to third-party technical testing and a limit on the China share relative to US customers. A proclamation signed January 14, 2026 imposes a 25% duty on advanced computing chips at the same performance thresholds. Distillation relevance: Loosens exactly the constraint that Anthropic and CNAS argue is the binding limit on distillation's usefulness. Reported volume allowances were contested in press accounts. Source: https://www.cnbc.com/2026/01/14/trump-nvidia-h200-china-ai-chips.html - United States — Chip Security Act (H.R. 3447 / S. 1705) [H.R. 3447 reported 42-0 on 2026-03-26; not enacted, 2025-05-08] What it says: Would require Commerce, within 180 days, to mandate chip security mechanisms implementing location verification on covered integrated circuits before export, reexport or in-country transfer, with reporting when a chip is detected outside its licensed location. Distillation relevance: Targets smuggling, the alleged compute channel behind DeepSeek V4 and Moonshot's infrastructure claims. Source: https://www.congress.gov/bill/119th-congress/house-bill/3447 - United States — GAIN AI Act of 2025 (S. 3150 / H.R. 5885) [Passed Senate as an NDAA amendment; excluded from the enacted FY2026 NDAA, 2025-11-06] What it says: Would require US chipmakers to prioritise American customers before selling advanced AI chips to China and other arms-embargoed countries. Distillation relevance: The main proposed statutory brake on the compute flows that make distilled data useful. Opposed by Nvidia, SIA and the White House AI adviser; dropped in conference. Source: https://www.congress.gov/bill/119th-congress/senate-bill/3150 - United States — No Adversarial AI Act (S. 2177 / H.R. 4142) [Referred to committee, 2025-06-25] What it says: Would direct the Federal Acquisition Security Council to identify and publish a list of AI developed by companies associated with China, Russia, Iran and North Korea, and ban executive agency use with narrow research, testing and mission-critical exceptions requiring written justification to Congress and OMB. Distillation relevance: Procurement-side response: keeps allegedly distilled models out of federal systems rather than punishing the extraction itself. Source: https://www.congress.gov/bill/119th-congress/senate-bill/2177/text - United States — Decoupling America's Artificial Intelligence Capabilities from China Act (S. 321) [Referred to Senate Judiciary; not advanced, 2025-01-29] What it says: Would prohibit US persons from exporting AI or generative AI technology or IP to China or importing Chinese-developed AI, bar US-China joint AI research, and prohibit financing China-linked AI R&D, with penalties up to 20 years imprisonment. Distillation relevance: The maximalist response, introduced the day the DeepSeek distillation story broke. Its breadth — potentially reaching individuals downloading Chinese models — drew criticism across the political spectrum. Source: https://www.congress.gov/bill/119th-congress/senate-bill/321 - United States — Section 1260H Chinese Military Companies List (June 2026 update) [In force; procurement prohibitions from 2026-06-30, 2026-06-08] What it says: Adds 65 entities including Alibaba, Baidu, BYD and Unitree, bringing the list to 188. Designations cite indirect SASAC affiliation and military-civil fusion contributions. Distillation relevance: Anthropic's June 10 letter cites the Alibaba listing to argue that distillation feeds PLA-relevant capability. DeepSeek was not added despite State Department claims about its military support. Source: https://www.cnbc.com/2026/06/09/alibaba-baidu-byd-named-on-pentagons-china-military-list-.html - United States (states) — State bans on DeepSeek for government devices [In force in 14+ states, 2025-01-31] What it says: Executive directives barring DeepSeek from state-owned devices and networks: Texas (Jan 31, 2025), New York (Feb 10), Virginia (Feb 11), Iowa (Feb 19), South Dakota, Nebraska, Tennessee, Arkansas, North Dakota, Oklahoma (Mar 21), Alabama, Georgia and North Carolina (March), and Kansas (April) — 14 states as of April 2025 per StateTech. Some states, such as Florida, restricted at agency level only. Distillation relevance: Not distillation-specific — the stated grounds are data collection and CCP linkage — but the state layer is where restriction on allegedly distilled models actually binds today. Source: https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek - United States (states) — California SB 53 (TFAIA) and New York RAISE Act [In force, 2026-01-01] What it says: California's Transparency in Frontier Artificial Intelligence Act took effect January 1, 2026 as the first state law requiring standardized safety and transparency disclosures from frontier developers. New York's RAISE Act was substantially amended on March 27, 2026 to align with it. Distillation relevance: Neither addresses distillation. They matter because they fill the federal gap on frontier-model regulation while distillation policy runs entirely through national security channels. Source: https://www.cooley.com/news/insight/2026/2026-03-31-new-yorks-frontier-ai-law-gets-a-california-makeover-with-some-key-differences - European Union — EU AI Act, general-purpose AI obligations [Applying since 2025-08-02; enforceable since 2026-08-02, 2025-08-02] What it says: Obligations on GPAI providers covering technical documentation, downstream information, copyright policy and a public training-data summary, with additional systemic-risk duties above a compute threshold. Fines up to EUR 15 million or 3% of worldwide turnover. Distillation relevance: Provenance-neutral: nothing in the Act makes extracting another model's capabilities an offence. Commission officials have signalled that DeepSeek-style compute efficiency may force a rethink of the systemic-risk compute threshold. Source: https://ec.europa.eu/commission/presscorner/detail/en/ip_26_1714 - European Union — General-Purpose AI Code of Practice [Published 2025-07-10; approved 2025-08-01; voluntary, 2025-07-10] What it says: Three chapters — Transparency, Copyright, Safety and Security — offering a presumption-of-conformity route for GPAI providers. Distillation relevance: The Copyright chapter is the closest EU analogue to the distillation debate, but it governs what goes into training rather than where the training data came from. Source: https://artificialintelligenceact.eu/code-of-practice-overview/ - Italy — Garante order blocking DeepSeek [In force, 2025-01-30] What it says: The Italian data protection authority ordered DeepSeek to block its chatbot in Italy after it failed to address concerns about its privacy policy and data transfers to China. Distillation relevance: Shows the European route runs through GDPR, not IP or security law. Self-hosted open weights on EU infrastructure avoid the transfer question entirely. Source: https://www.privacylaws.com/news/italy-and-south-korea-ban-deepseek-and-start-investigation/ - South Korea — AI Basic Act (Framework Act on AI Development and Trust) [In force, 2026-01-22] What it says: Comprehensive framework with generative-AI and high-impact AI obligations, a triennial AI Basic Plan, a National AI Committee, an AI Policy Center and an AI Safety Research Institute. Distillation relevance: Silent on distillation. Korea's operative action against Chinese models was the February 2025 PIPC suspension of DeepSeek app downloads on data-protection grounds. Source: https://www.cooley.com/news/insight/2026/2026-01-27-south-koreas-ai-basic-act-overview-and-key-takeaways - Japan — Act on Promotion of Research and Development and Application of AI-Related Technologies [In force, 2025-05] What it says: Japan's first comprehensive AI legislation. Promotional and light-touch: establishes an AI Strategy Headquarters chaired by the Prime Minister and a Basic Plan, with principles including maintaining domestic R&D capability and national security. Distillation relevance: No distillation provisions and no restrictions on Chinese models. Japan's relevance to the distillation chain is upstream, through its semiconductor equipment export controls. Source: https://www.whitecase.com/insight-our-thinking/ai-watch-global-regulatory-tracker-japan - United Kingdom — Sovereign AI Unit and AI Security Institute [Non-statutory; operating, 2025] What it says: The Sovereign AI Unit was established within DSIT in 2025 under the AI Opportunities Action Plan, alongside roughly GBP 2 billion in commitments, compute infrastructure plans and AI Growth Zones. The safety institute's remit has expanded toward standards and assurance. Distillation relevance: The UK has no distillation instrument and no ban on Chinese models. Its contribution is evaluation evidence — published findings on DeepSeek R1 jailbreak susceptibility and embedded political bias. Source: https://oecd.ai/en/dashboards/policy-initiatives/uk-sovereign-ai-unit - China — Global AI Governance Action Plan and World AI Cooperation Organization [Announced 2025-07; organisation launched 2026-07 with 29 founding members, 2025-07-26] What it says: A thirteen-point roadmap on AI safety, infrastructure, data standards and sustainable development, plus an International Open Source AI Cooperation Initiative and a proposed cooperation body headquartered in Shanghai. Distillation relevance: China's counter-narrative: open capability diffusion as a global public good, positioned directly against US restriction of model access. Source: https://technode.com/2025/07/29/china-proposes-new-global-ai-cooperation-organization-headquarter-planned-in-shanghai/ - China — MOFCOM consultation on AI export controls [Under consultation, 2026-07-21] What it says: Reported consultation with Alibaba, ByteDance, Zhipu and chip firms on adding model weights, key training data and chip designs to the technology export catalogue, potentially restricting foreign download of leading Chinese model weights while keeping hosted access open, plus limits on foreign fabrication of Chinese chip designs. Distillation relevance: Would end the asymmetry that makes the current US framing possible: if Chinese weights become licensed exports, both sides regulate model diffusion. Source: https://www.techrepublic.com/article/news-apac-china-ai-model-export-controls/ - Corporate (United States) — Anthropic policy barring entities controlled from China, Russia, Iran and North Korea [In force, 2025-09-05] What it says: Terms of service update prohibiting service to companies more than 50% owned by entities in those jurisdictions, regardless of where the subsidiary operates. Cited legal compulsion to share data with intelligence services; estimated revenue impact in the low hundreds of millions of dollars. Distillation relevance: Private access policy that H.R. 8283 would convert into a legal designation trigger. Anthropic's June 2026 letter cites it as the reason Alibaba's access was unauthorized. Source: https://www.semafor.com/article/09/05/2025/anthropic-blocks-ai-sales-in-china - Corporate (United States) — OpenAI geographic API blocking and Verified Organization checks [In force, 2024-07-09] What it says: From July 9, 2024 OpenAI blocked API traffic from unsupported regions including mainland China. From April 2025 it required government-ID Verified Organization status for access to certain frontier models, one ID per organisation per 90 days. Distillation relevance: The access-control layer whose circumvention is what H.R. 8283's definition of a model extraction attack actually turns on. Source: https://help.openai.com/en/articles/10910291-api-organization-verification - Multilateral (industry) — Frontier Model Forum distillation threat-intelligence sharing [Operating since 2026-04, 2026-04-07] What it says: OpenAI, Anthropic and Google agreed to share indicators of extraction campaigns through the Frontier Model Forum, responding to the AI Action Plan's call for an industry information-sharing centre. Distillation relevance: The main non-governmental countermeasure. Its principal obstacle is antitrust: Anthropic has asked Congress to clarify that sharing tactics between competing labs is lawful. Source: https://www.techbrew.com/stories/openai-anthropic-google-distillation-collab PUBLIC ACCUSATIONS AND DISPUTES (14) - 2025-01-28 — David Sacks, White House AI and crypto czar vs DeepSeek Claim: DeepSeek distilled knowledge out of OpenAI models to build its own systems Evidence: Asserted 'substantial evidence' in a Fox News interview; no evidence detailed publicly Outcome: Set the political frame for everything that followed; no legal action Source: https://www.bloomberg.com/news/articles/2025-01-28/ai-czar-sacks-says-evidence-deepseek-leaned-on-openai-s-models - 2025-01-29 — OpenAI and Microsoft vs A group believed linked to DeepSeek Claim: Large-scale unauthorized data exfiltration through the OpenAI API, violating terms of service Evidence: Microsoft security researchers observed the activity in autumn 2024; OpenAI said it had seen evidence of distillation and referenced obfuscated methods Outcome: Accounts terminated; no lawsuit filed as of September 2026; matter escalated to Congress and the executive branch instead Source: https://www.bloomberg.com/news/articles/2025-01-29/microsoft-probing-if-deepseek-linked-group-improperly-obtained-openai-data - 2025-03-13 — OpenAI (OSTP filing) vs DeepSeek and PRC-produced models generally Claim: DeepSeek is state-subsidized and state-controlled, is insecure because Chinese law compels data disclosure, and has attempted distillation of US frontier models including through new obfuscated methods Evidence: Platform activity described in the filing and in a parallel assessment provided to the House Select Committee Outcome: OpenAI later said it was proposing export-rule changes rather than usage bans; no PRC model ban was adopted Source: https://techcrunch.com/2025/03/13/openai-calls-deepseek-state-controlled-calls-for-bans-on-prc-produced-models/ - 2025-04-16 — House Select Committee on the CCP (bipartisan) vs DeepSeek Claim: Profound national security threat: data routed through China Mobile-linked infrastructure, likely unlawful distillation of US models, and use of restricted Nvidia chips Evidence: Committee investigation; report 'DeepSeek Unmasked'; questions posed to Nvidia about chip use Outcome: Recommendations to expand and better enforce export controls; no enforcement action against DeepSeek followed directly Source: https://chinaselectcommittee.house.gov/media/press-releases/moolenaar-krishnamoorthi-unveil-explosive-report-on-chinese-ai-firm-deepseek-demand-answers-from-nvidia-over-chip-use - 2026-02-12 — OpenAI vs Several major Chinese LLM providers and some university research labs; also actors in Russia Claim: Sophisticated, multi-stage pipelines used to distill American AI capabilities; attackers now also use the models to filter data and simulate human task feedback Evidence: Memo to the House Select Committee, 'Updated Stakes for American-Led, Democratic AI' Outcome: Fed directly into the April 2026 hearing and NSTM-4; no named-entity enforcement Source: https://cdn.openai.com/pdf/045aa967-ee96-4a09-94ee-3098ddf6db2c/OpenAI-US-House-Select-Cmte-Update-%5B021226%5D.pdf - 2026-02-12 — Google Threat Intelligence Group vs Unattributed actors worldwide, including researchers and private companies Claim: Model extraction activity targeting Gemini, including one cluster of more than 100,000 prompts attempting to coerce reasoning behaviour usable for replication Evidence: Google Threat Intelligence Group publication, February 12, 2026; activity disrupted Outcome: Notably framed as global rather than China-specific, complicating the geopolitical narrative Source: https://cloud.google.com/blog/topics/threat-intelligence/distillation-experimentation-integration-ai-adversarial-use - 2026-02-23 — Anthropic vs DeepSeek, Moonshot AI, MiniMax Claim: Industrial-scale distillation attacks: over 16 million exchanges via approximately 24,000 fraudulent accounts, in violation of terms of service and regional access restrictions Evidence: Internal detection classifiers and account analysis; per-lab attribution of 150,000+ (DeepSeek), 3.4M (Moonshot) and 13M (MiniMax) exchanges Outcome: Accounts terminated; detection classifiers, access controls and model-level countermeasures deployed; Anthropic called for coordinated industry and government response and for sustained chip export controls Source: https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks - 2026-04-23 — White House OSTP (Michael Kratsios) vs Foreign entities principally based in China Claim: Deliberate, industrial-scale campaigns to distill US frontier AI systems using tens of thousands of proxy accounts and jailbreaking; resulting models strip security protocols and undo neutrality mechanisms Evidence: Memorandum NSTM-4; underlying evidence drawn from lab disclosures rather than published independently Outcome: Directed agency intelligence sharing with industry and exploration of accountability measures; issued weeks before a planned Trump-Xi meeting Source: https://www.nextgov.com/artificial-intelligence/2026/04/white-house-accuses-china-deliberate-industrial-scale-campaigns-steal-us-ai-models/413083/ - 2026-04-24 — US State Department vs DeepSeek, Moonshot AI, MiniMax Claim: Surreptitious, unauthorized distillation campaigns producing models that appear comparable on select benchmarks at a fraction of the cost but do not replicate full performance and deliberately strip security protocols Evidence: Diplomatic cable to posts worldwide instructing staff to raise the issue with foreign counterparts Outcome: Chinese Embassy rejected the allegations as groundless and a deliberate attack on China's AI development Source: https://www.cnbc.com/2026/04/25/us-global-warning-alleged-china-ai-theft.html - 2026-06-10 — Anthropic (Sarah Heck, Head of Policy) vs Alibaba and Alibaba Qwen-affiliated operators Claim: Largest known distillation attack against Anthropic: more than 28.8 million exchanges through almost 25,000 fraudulent accounts between April 22 and June 5, 2026, targeting agentic reasoning, software engineering and long-horizon tasks Evidence: Confidential evidence provided to the Senate Banking Committee ahead of its June 11 hearing; described as following the same patterns as the February disclosure Outcome: Alibaba denied using proprietary model outputs and denied government involvement; it barred employees from using Anthropic products (announced 6 July 2026, effective 10 July 2026). Anthropic asked Congress for antitrust clarity on lab-to-lab sharing, chip-loophole closure, and penalties for responsible labs Source: https://d1e00ek4ebabms.cloudfront.net/production/uploaded-files/Anthropic%20letter%20Alibaba%E3%83%BBJun%2010%202026%20-ba307cab-ccb4-43f1-88c0-515f29694678.pdf - 2026-07-21 — US Treasury Secretary Scott Bessent vs Chinese AI model developers generally Claim: Watermarks of American large language models are appearing inside Chinese AI systems; the US has the ability to sanction overseas models that steal from US companies Evidence: Asserted on Fox Business; the watermark evidence was not published Outcome: Threat only. No Chinese AI model had been sanctioned as of early September 2026; China's Ministry of Commerce promised 'all necessary measures' in response Source: https://www.cnbc.com/2026/07/21/bessent-china-ai-sanctions.html - 2026-07-22 — White House OSTP (Michael Kratsios) vs Moonshot AI Claim: Covertly distilled Anthropic's Fable to build Kimi K3, using an internal platform that switched rapidly between access methods to avoid detection, and obtained export-controlled Nvidia servers Evidence: None published in the week following the accusation Outcome: Researchers including Nathan Lambert (Allen Institute for AI) and Braden Hancock (Laude Institute) publicly doubted distillation alone could explain K3's capabilities; Moonshot did not respond and Anthropic did not comment on the specific allegation Source: https://cyberscoop.com/white-house-accuses-moonshot-ai-anthropic-model-distillation/ - 2026-07-27 — China Ministry of Commerce vs Many American AI enterprises Claim: US firms have themselves distilled Chinese models; US allegations lack factual basis and legal support and constitute double standards and 'AI hegemony' Evidence: None supplied; no companies named Outcome: Paired with a promise of countermeasures if Chinese AI firms are sanctioned, and with a reported MOFCOM consultation on export controls covering model weights and training data Source: https://www.implicator.ai/china-says-us-firms-distilled-chinese-models/ - 2026-02 — Unnamed senior Trump administration officials vs DeepSeek Claim: DeepSeek trained V4 on banned Nvidia Blackwell GPUs smuggled via Southeast Asian shell companies and housed in an Inner Mongolia data centre Evidence: Official statements to press; no documentation published Outcome: Nvidia called the claim farfetched; the V4 preview released April 24, 2026 was notable for being optimised for Huawei Ascend silicon Source: https://www.technologyreview.com/2026/04/24/1136422/why-deepseeks-v4-matters/ TIMELINE OF THIS PERSPECTIVE (47 EVENTS) ---------------------------------------- 2022-10-07 | policy | BIS imposes advanced computing export controls on China 2023-10-17 | policy | BIS closes the A800/H800 workaround 2024-07-09 | product | OpenAI blocks API traffic from unsupported regions including mainland China 2024-12-27 | research | DeepSeek V3 is reported to self-identify as ChatGPT 2025-01-15 | policy | AI Diffusion Rule published 2025-01-28 | policy | David Sacks says there is 'substantial evidence' DeepSeek distilled OpenAI models 2025-01-29 | legal | Microsoft and OpenAI confirm they are investigating a DeepSeek-linked group 2025-01-29 | policy | Dario Amodei publishes 'On DeepSeek and Export Controls' 2025-01-29 | policy | Senator Hawley introduces S. 321, the Decoupling America's AI Capabilities from China Act 2025-01-30 | legal | Italy's Garante orders DeepSeek's chatbot blocked 2025-01-31 | policy | Texas becomes the first US state to ban DeepSeek on state devices 2025-02-04 | policy | Australia bans DeepSeek on government devices 2025-03-13 | policy | OpenAI's OSTP filing calls DeepSeek 'state-subsidized' and 'state-controlled' 2025-04-14 | product | OpenAI introduces Verified Organization ID checks for frontier API access 2025-04-16 | policy | House Select Committee on the CCP publishes 'DeepSeek Unmasked' 2025-05-13 | policy | BIS announces rescission of the AI Diffusion Rule 2025-06-25 | policy | No Adversarial AI Act introduced 2025-07-10 | policy | European Commission publishes the final GPAI Code of Practice 2025-07-23 | policy | White House releases 'Winning the Race: America's AI Action Plan' 2025-07-26 | policy | China unveils a Global AI Governance Action Plan and proposes a world AI cooperation body 2025-08-02 | policy | EU AI Act obligations for general-purpose AI providers begin to apply 2025-08-11 | policy | Nvidia and AMD reported to agree a 15% China revenue share for export licences 2025-09-05 | product | Anthropic bars entities majority-controlled from China, Russia, Iran and North Korea 2025-10-09 | policy | Senate passes its NDAA including the GAIN AI Act 2025-12-18 | policy | FY2026 NDAA signed without the GAIN AI Act 2026-01-15 | policy | BIS moves H200 and MI325X exports to case-by-case review; 25% tariff applies 2026-01-22 | policy | South Korea's AI Basic Act takes effect 2026-02-12 | legal | OpenAI memo to the House Select Committee and Google GTIG report land the same day 2026-02-23 | legal | Anthropic publicly discloses industrial-scale distillation attacks 2026-03-26 | policy | Chip Security Act reported out of House Foreign Affairs 42-0 2026-04-07 | market | OpenAI, Anthropic and Google agree to share distillation threat intelligence 2026-04-15 | policy | H.R. 8283, the Deterring American AI Model Theft Act of 2026, is introduced 2026-04-16 | policy | House Select Committee hearing: 'China's Illicit Campaign to Steal and Subvert American AI Technology' 2026-04-22 | policy | H.R. 8283 ordered reported 43-0 2026-04-23 | policy | OSTP issues NSTM-4, 'Adversarial Distillation of American AI Models' 2026-04-24 | policy | State Department cables posts worldwide to raise distillation with foreign counterparts 2026-04-24 | product | DeepSeek releases a preview of V4 2026-06-05 | policy | President signs NSPM-11 on AI in the national security enterprise 2026-06-08 | policy | Pentagon adds Alibaba, Baidu, BYD and Unitree to the 1260H list 2026-06-10 | legal | Anthropic tells Senate Banking that Alibaba ran the largest known distillation attack against it 2026-06-17 | policy | US holds off blacklisting DeepSeek and 100-plus other flagged firms 2026-07-06 | market | Alibaba bars employees from using Anthropic products 2026-07-21 | policy | Treasury Secretary Bessent threatens sanctions over AI model theft 2026-07-22 | legal | Kratsios accuses Moonshot AI of distilling Anthropic's Fable to build Kimi K3 2026-07-24 | policy | Open-weights letter launches with ~25 signatories, later exceeding 270 2026-07-27 | policy | China's Ministry of Commerce counter-accuses American AI firms of distilling Chinese models 2026-08-02 | policy | EU AI Act enforcement powers go live SOURCES CITED BY THIS SECTION (100) ----------------------------------- [1] Detecting and preventing distillation attacks — Anthropic, 2026-02-23 (blog) https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks [2] Letter to Senate Banking Committee on illicit access to American AI models by Alibaba-affiliated operators — Anthropic, 2026-06-10 (filing) https://d1e00ek4ebabms.cloudfront.net/production/uploaded-files/Anthropic%20letter%20Alibaba%E3%83%BBJun%2010%202026%20-ba307cab-ccb4-43f1-88c0-515f29694678.pdf [3] H.R. 8283, Deterring American AI Model Theft Act of 2026 (introduced text) — US Government Publishing Office, 2026-04-15 (law) https://www.govinfo.gov/content/pkg/BILLS-119hr8283ih/pdf/BILLS-119hr8283ih.pdf [4] H.R. 8283 bill page and status — Congress.gov, 2026-04-22 (law) https://www.congress.gov/bill/119th-congress/house-bill/8283/text [5] National Security Presidential Memorandum NSPM-11 — The White House, 2026-06-05 (law) https://www.whitehouse.gov/presidential-actions/2026/06/national-security-presidential-memorandum-nspm-11/ [6] Winning the Race: America's AI Action Plan — The White House, 2025-07-23 (law) https://www.whitehouse.gov/wp-content/uploads/2025/07/Americas-AI-Action-Plan.pdf [7] China's Illicit Campaign to Steal and Subvert American AI Technology (testimony of Yusuf Mahmood) — US House Select Committee on the CCP, 2026-04-16 (filing) https://docs.house.gov/meetings/ZS/ZS00/20260416/119165/HHRG-119-ZS00-Wstate-MahmoodY-20260416.pdf [8] OpenAI response to the OSTP/NSF RFI on the AI Action Plan — OpenAI, 2025-03-13 (filing) https://cdn.openai.com/global-affairs/ostp-rfi/ec680b75-d539-4653-b297-8bcf6e5f7686/openai-response-ostp-nsf-rfi-notice-request-for-information-on-the-development-of-an-artificial-intelligence-ai-action-plan.pdf [9] OpenAI calls DeepSeek 'state-controlled,' calls for bans on 'PRC-produced' models — TechCrunch, 2025-03-13 (news) https://techcrunch.com/2025/03/13/openai-calls-deepseek-state-controlled-calls-for-bans-on-prc-produced-models/ [10] AI Czar Sacks Says 'Evidence' DeepSeek Leaned On OpenAI's Models — Bloomberg, 2025-01-28 (news) https://www.bloomberg.com/news/articles/2025-01-28/ai-czar-sacks-says-evidence-deepseek-leaned-on-openai-s-models [11] Microsoft Probing If DeepSeek-Linked Group Improperly Obtained OpenAI Data — Bloomberg, 2025-01-29 (news) https://www.bloomberg.com/news/articles/2025-01-29/microsoft-probing-if-deepseek-linked-group-improperly-obtained-openai-data [12] OpenAI says DeepSeek may have 'inappropriately' used its models' output — Axios, 2025-01-29 (news) https://www.axios.com/2025/01/29/openai-deepseek-ai-models-data-training [13] On DeepSeek and Export Controls — Dario Amodei, 2025-01-29 (blog) https://darioamodei.com/post/on-deepseek-and-export-controls [14] Moolenaar, Krishnamoorthi unveil report on DeepSeek — US House Select Committee on the CCP, 2025-04-16 (filing) https://chinaselectcommittee.house.gov/media/press-releases/moolenaar-krishnamoorthi-unveil-explosive-report-on-chinese-ai-firm-deepseek-demand-answers-from-nvidia-over-chip-use [15] S. 321 Decoupling America's Artificial Intelligence Capabilities from China Act — Congress.gov, 2025-01-29 (law) https://www.congress.gov/bill/119th-congress/senate-bill/321 [16] H.R. 3447 Chip Security Act — Congress.gov, 2025-05-15 (law) https://www.congress.gov/bill/119th-congress/house-bill/3447 [17] S. 1705 Chip Security Act — Congress.gov, 2025-05-08 (law) https://www.congress.gov/bill/119th-congress/senate-bill/1705 [18] S. 3150 GAIN AI Act of 2025 — Congress.gov, 2025-11-06 (law) https://www.congress.gov/bill/119th-congress/senate-bill/3150 [19] S. 2177 No Adversarial AI Act (text) — Congress.gov, 2025-06-25 (law) https://www.congress.gov/bill/119th-congress/senate-bill/2177/text [20] H.R. 1121 No DeepSeek on Government Devices Act — Congress.gov, 2025-02-07 (law) https://www.congress.gov/bill/119th-congress/house-bill/1121/all-info [21] US Export Controls and China: Advanced Semiconductors (CRS R48642) — Congressional Research Service, 2025-08 (filing) https://www.congress.gov/crs-product/R48642 [22] Explainer: The Commerce Department's October 2023 Export Controls Update — CSET, Georgetown, 2023-10 (blog) https://cset.georgetown.edu/article/bis-2023-update-explainer/ [23] BIS Rescinds AI Diffusion Rule and Issues Guidance — Wiley Rein, 2025-05-14 (blog) https://www.wiley.law/alert-BIS-Rescinds-AI-Diffusion-Rule [24] Nvidia says it will record $5.5 billion charge tied to H20 processors — CNBC, 2025-04-15 (news) https://www.cnbc.com/2025/04/15/nvidia-says-it-will-record-5point5-billion-quarterly-charge-tied-to-h20-processors-exported-to-china.html [25] Nvidia, AMD to pay 15% of China chip revenue to US government — Fortune, 2025-08-10 (news) https://fortune.com/2025/08/10/nvidia-amd-chips-h20-mi308-china-sales-revenue-trump-export-license/ [26] Trump administration clears way for Nvidia H200 chip sales to China — CNBC, 2026-01-14 (news) https://www.cnbc.com/2026/01/14/trump-nvidia-h200-china-ai-chips.html [27] AI export control bill passes Senate as NDAA amendment — Nextgov/FCW, 2025-10-09 (news) https://www.nextgov.com/artificial-intelligence/2025/10/ai-export-control-bill-passes-senate-ndaa-amendment/408762/ [28] Bill prioritizing American customers for AI chips not expected to make it into final NDAA — Nextgov/FCW, 2025-12 (news) https://www.nextgov.com/policy/2025/12/bill-prioritizing-american-customers-ai-chips-not-expected-make-it-final-ndaa-sources-say/409920/ [29] White House accuses China of deliberate, industrial-scale campaigns to steal US AI models — Nextgov/FCW, 2026-04-23 (news) https://www.nextgov.com/artificial-intelligence/2026/04/white-house-accuses-china-deliberate-industrial-scale-campaigns-steal-us-ai-models/413083/ [30] US State Department orders global warning about alleged China AI theft — CNBC, 2026-04-25 (news) https://www.cnbc.com/2026/04/25/us-global-warning-alleged-china-ai-theft.html [31] Anthropic accuses Alibaba of campaign to 'brazenly' and 'illicitly' extract Claude capabilities — CNBC, 2026-06-24 (news) https://www.cnbc.com/2026/06/24/anthropic-alibaba-distillation-campaign.html [32] China's Alibaba bans Anthropic AI for employees after distillation accusation — CNBC, 2026-07-06 (news) https://www.cnbc.com/2026/07/06/alibaba-anthropic-ai-ban-claude-china.html [33] Bessent says U.S. could sanction China over AI model 'theft' — CNBC, 2026-07-21 (news) https://www.cnbc.com/2026/07/21/bessent-china-ai-sanctions.html [34] US holds off blacklisting China's DeepSeek and 100+ firms — CNBC, 2026-06-17 (news) https://www.cnbc.com/2026/06/17/us-deepseek-blacklist-cxmt-national-security-risks-.html [35] Pentagon expands list of China military-linked firms to include Alibaba, Baidu, BYD — CNBC, 2026-06-09 (news) https://www.cnbc.com/2026/06/09/alibaba-baidu-byd-named-on-pentagons-china-military-list-.html [36] Anthropic, OpenAI among firms facing new scrutiny under EU AI Act enforcement — CNBC, 2026-08-03 (news) https://www.cnbc.com/2026/08/03/eu-ai-act-enforcement-powers.html [37] Commission starts enforcing AI Act rules and new transparency obligations — European Commission, 2026-08-02 (law) https://ec.europa.eu/commission/presscorner/detail/en/ip_26_1714 [38] Overview of the General-Purpose AI Code of Practice — EU AI Act explorer, 2025-07-10 (docs) https://artificialintelligenceact.eu/code-of-practice-overview/ [39] EU AI Act rules on GPAI models under DeepSeek review — Pinsent Masons, 2025 (blog) https://www.pinsentmasons.com/out-law/analysis/eu-ai-act-gpai-deepseek-review [40] White House accuses Moonshot AI of distilling Anthropic's model — CyberScoop, 2026-07-22 (news) https://cyberscoop.com/white-house-accuses-moonshot-ai-anthropic-model-distillation/ [41] Experts say exploiting Anthropic's Fable isn't how Kimi K3 got so good — TechCrunch, 2026-07-23 (news) https://techcrunch.com/2026/07/23/experts-say-exploiting-anthropics-fable-isnt-how-kimi-k3-got-so-good/ [42] Three reasons why DeepSeek's new model matters — MIT Technology Review, 2026-04-24 (news) https://www.technologyreview.com/2026/04/24/1136422/why-deepseeks-v4-matters/ [43] The Case for Imposing Costs on China's AI Distillation Campaigns — Just Security, 2026 (blog) https://www.justsecurity.org/134124/costs-china-ai-distillation/ [44] From Diagnosis to Deterrence: The Emerging U.S. Response to Distillation — Just Security, 2026 (blog) https://www.justsecurity.org/137498/diagnosis-deterrence-us-response-distillation/ [45] Responding to AI Distillation Without Panic — Institute for Law & AI, 2026 (blog) https://law-ai.org/responding-to-ai-distillation-without-panic/ [46] How to Fix the AI Model Theft Bill Before It Becomes Law — ITIF, 2026-07-28 (blog) https://itif.org/publications/2026/07/28/how-to-fix-the-ai-model-theft-bill-before-it-becomes-law/ [47] Adversarial Distillation: China's Campaign to Extract American AI Capabilities — Center for a New American Security, 2026-06-02 (paper) https://www.cnas.org/publications/reports/adversarial-distillation [48] AI Distillation Attacks: Executive and Congressional Action Can Go Further — Institute for AI Policy and Strategy, 2026 (paper) https://www.iaps.ai/research/ai-distillation-attacks-executive-and-congressional-action-can-go-further [49] Trade Secrecy Meets Generative AI — Camilla Alexandra Hrdy, SSRN, 2025 (paper) https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5121745 [50] Dispute over AI model distillation tech in OpenAI-DeepSeek case — Law.asia, 2025 (blog) https://law.asia/openai-deepseek-ai-distillation/ [51] Copyright Office report on copyrightability of AI-generated material — US Copyright Office, 2025-01 (law) https://www.copyright.gov/newsnet/2025/1060.html [52] 'Ironic, hypocritical' of big tech to call out DeepSeek — Cornell University, 2025-01 (news) https://news.cornell.edu/media-relations/tip-sheets/ironic-hypocritical-big-tech-call-out-deepseek [53] OpenAI hit with mockery over DeepSeek complaint — Futurism, 2025-01 (news) https://futurism.com/openai-mockery-stole-work-deepseek [54] Why DeepSeek's new AI model thinks it's ChatGPT — TechCrunch, 2024-12-27 (news) https://techcrunch.com/2024/12/27/why-deepseeks-new-ai-model-thinks-its-chatgpt/ [55] Anthropic blocks sales of AI to Chinese firms — Semafor, 2025-09-05 (news) https://www.semafor.com/article/09/05/2025/anthropic-blocks-ai-sales-in-china [56] US AI giant Anthropic bars Chinese-owned entities — RTE, 2025-09-05 (news) https://www.rte.ie/news/business/2025/0905/1531954-us-ai-giant-anthropic-bars-chinese-owned-entities/ [57] OpenAI Terms of Use — OpenAI, 2025 (docs) https://openai.com/policies/row-terms-of-use/ [58] Anthropic Consumer Terms of Service — Anthropic, 2025 (docs) https://www.anthropic.com/legal/consumer-terms [59] Meta Llama 3 Community License — Meta, 2024 (docs) https://www.llama.com/llama3/license/ [60] Llama 3.3 70B Instruct model card — Meta / Hugging Face, 2024-12 (docs) https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct [61] API Organization Verification — OpenAI, 2025-04 (docs) https://help.openai.com/en/articles/10910291-api-organization-verification [62] OpenAI to cut off API access in China on July 9 — Rest of World, 2024-06 (news) https://restofworld.org/2024/exporter-openai-china-api-access/ [63] OpenAI accuses DeepSeek of malpractice ahead of AI launch — Rest of World, 2026 (news) https://restofworld.org/2026/openai-deepseek-distillation-dispute-us-china/ [64] These States Have Banned DeepSeek — StateTech Magazine, 2025-04 (news) https://statetechmagazine.com/article/2025/04/these-states-have-banned-deepseek [65] New York state bans DeepSeek from government devices — NBC News, 2025-02-10 (news) https://www.nbcnews.com/tech/new-york-state-bans-deepseek-government-devices-rcna191510 [66] Italy and South Korea ban DeepSeek and start investigation — Privacy Laws & Business, 2025-02 (news) https://www.privacylaws.com/news/italy-and-south-korea-ban-deepseek-and-start-investigation/ [67] Governor Stitt bans DeepSeek on all state-owned devices — State of Oklahoma, 2025-03 (law) https://oklahoma.gov/governor/newsroom/newsroom/2025/-governor-stitt-bans-deepseek-on-all-state-owned-devices-due-to-.html [68] South Korea's AI Basic Act: Overview and Key Takeaways — Cooley, 2026-01-27 (blog) https://www.cooley.com/news/insight/2026/2026-01-27-south-koreas-ai-basic-act-overview-and-key-takeaways [69] AI Watch: Global regulatory tracker - Japan — White & Case, 2025 (blog) https://www.whitecase.com/insight-our-thinking/ai-watch-global-regulatory-tracker-japan [70] UK Sovereign AI Unit — OECD.AI, 2025 (docs) https://oecd.ai/en/dashboards/policy-initiatives/uk-sovereign-ai-unit [71] China proposes new global AI cooperation organization, headquarters planned in Shanghai — TechNode, 2025-07-29 (news) https://technode.com/2025/07/29/china-proposes-new-global-ai-cooperation-organization-headquarter-planned-in-shanghai/ [72] China launches Shanghai-based AI governance body with 29 founding nations — Caixin Global, 2026-07-17 (news) https://www.caixinglobal.com/2026-07-17/china-launches-shanghai-based-ai-governance-body-with-29-founding-nations-102465524.html [73] China Accuses US AI Firms of Distilling Chinese Models — Implicator.ai, 2026-07-27 (news) https://www.implicator.ai/china-says-us-firms-distilled-chinese-models/ [74] China Considers Export Controls on AI Models, Training Data and Chip Technology — TechRepublic, 2026-07 (news) https://www.techrepublic.com/article/news-apac-china-ai-model-export-controls/ [75] Google disrupts Gemini model extraction attempts — TechInformed, 2026-02-16 (news) https://techinformed.com/google-disrupts-gemini-model-extraction-attempts/ [76] OpenAI, Anthropic, Google join forces against China — Tech Brew, 2026-04-07 (news) https://www.techbrew.com/stories/openai-anthropic-google-distillation-collab [77] How to Buy Cheap Claude Tokens in China — ChinaTalk, 2026 (news) https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens-in [78] Inside the Gray Market for LLM Access — DeepLearning.AI, The Batch, 2026 (news) https://www.deeplearning.ai/the-batch/inside-the-gray-market-for-llm-access [79] Open Weights and American AI Leadership (open letter) — Multi-company coalition, 2026-07-24 (filing) https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf [80] Gulf AI infrastructure and the limits of technological sovereignty — IISS, 2026-06 (paper) https://www.iiss.org/publications/strategic-comments/2026/06/gulf-ai-infrastructure-and-the-limits-of-technological-sovereignty/ [81] Which Countries Have Banned DeepSeek Already? — Tech.co, 2025 (news) https://tech.co/news/which-countries-have-banned-deepseek-already [82] DeepSeek and Chinese AI Models: GDPR Data Transfer Compliance — AI Policy Desk, 2026-06 (blog) https://www.aipolicydesk.com/blog/deepseek-chinese-ai-models-gdpr-compliance-2026 [83] Report: Commerce Department bans use of DeepSeek on government devices — PYMNTS, 2025-02 (news) https://www.pymnts.com/artificial-intelligence-2/2025/report-commerce-department-bans-use-of-deepseek-on-government-devices [84] Three States Ban DeepSeek Use on State Devices and Networks — National Law Review, 2025-02 (news) https://natlawreview.com/article/three-states-ban-deepseek-use-state-devices-and-networks [85] South Korea joins Italy, Australia in banning DeepSeek — The Cable, 2025-02 (news) https://www.thecable.ng/south-korea-joins-italy-australia-in-banning-deepseek-over-security-concerns/ [86] Nvidia expects to lose billions in revenue due to H20 chip licensing requirements — TechCrunch, 2025-05-28 (news) https://techcrunch.com/2025/05/28/nvidia-expects-to-lose-billions-in-revenue-due-to-h20-chip-licensing-requirements/ [87] Nvidia announces financial results for 1st quarter fiscal 2026 — HPCwire, 2025-05-28 (filing) https://www.hpcwire.com/off-the-wire/nvidia-announces-financial-results-for-1st-quarter-fiscal-2026/ [88] H.R. 9363, AI Security and Innovation Act — House Science, Space and Technology Committee, 2026-06-25 (law) https://science.house.gov/2026/6/h-r-9363-ai-security-and-innovation-act [89] CBO cost estimate, H.R. 9363 — Congressional Budget Office, 2026 (law) https://www.cbo.gov/publication/62730 [90] President Trump orders narrowly targeted 25% Section 232 tariff on certain advanced semiconductors — White & Case, 2026-01 (law) https://www.whitecase.com/insight-alert/president-trump-orders-narrowly-targeted-25-section-232-tariff-certain-advanced [91] BlnBDI press release: DeepSeek apps reported to Apple and Google — Berlin Commissioner for Data Protection and Freedom of Information, 2025-06-27 (law) https://www.datenschutz-berlin.de/fileadmin/user_upload/pdf/pressemitteilungen/2025/20250627-BlnBDI-Press-Release_DeepSeek.pdf [92] Nvidia and 24 other companies sign open-weights letter — Tom's Hardware, 2026-07-24 (news) https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidia-and-24-other-companies-sign-open-weights-letter-as-washington-weighs-chinese-ai-model-ban [93] Open weights, American AI leadership letter: OpenAI absent — The Next Web, 2026-07-25 (news) https://thenextweb.com/news/open-weights-american-ai-leadership-letter-huang-nvidia-openai-absent [94] Taiwan bans DeepSeek in the public sector — Taipei Times, 2025-02-02 (news) https://www.taipeitimes.com/News/front/archives/2025/02/02/2003831193 [95] H.R. 4142 No Adversarial AI Act (introduced) — GovInfo, 2025-06-25 (law) https://www.govinfo.gov/app/details/BILLS-119hr4142ih [96] House Foreign Affairs markup documents, H.R. 8283 (April 22, 2026) — Congress.gov, 2026-04-22 (law) https://www.congress.gov/119/meeting/house/119191/documents/HMKP-119-FA00-20260422-SD002.pdf [97] Distillation, experimentation and integration: adversarial use of AI — Google Threat Intelligence Group, 2026-02-12 (blog) https://cloud.google.com/blog/topics/threat-intelligence/distillation-experimentation-integration-ai-adversarial-use [98] Updated Stakes for American-Led, Democratic AI (memo to House Select Committee) — OpenAI, 2026-02-12 (filing) https://cdn.openai.com/pdf/045aa967-ee96-4a09-94ee-3098ddf6db2c/OpenAI-US-House-Select-Cmte-Update-%5B021226%5D.pdf [99] Anthropic Commercial Terms of Service — Anthropic, 2025 (docs) https://www.anthropic.com/legal/commercial-terms [100] Anthropic Usage Policy — Anthropic, 2025 (docs) https://www.anthropic.com/legal/usage-policy ================================================================================================ 6. COMPANY PERSPECTIVE — COMPANY VIEW OF AI DISTILLATION: WHO USES IT, WHO SELLS IT, WHO POLICES IT ================================================================================================ Page: https://global-distillation.com/company Data: https://global-distillation.com/data/company.json Updated: 2026-09-03 Content: 8 key figures, 7 tables, 7 charts, 40 dated events, 14 glossary terms, 75 sources SUMMARY -------- Every major AI lab now uses knowledge distillation to build its small and mid-tier models: Google states in the Gemini 2.5 report that all models 'Flash size and below' are distilled, Meta co-distilled Llama 4 Maverick from the 2-trillion-parameter Behemoth, Qwen3's small models are 'strong-to-weak' distilled from Qwen3-235B, and Apple retrains its 3B on-device model with a distillation loss from a 64-expert MoE teacher. Three hyperscalers have shipped distillation as a product - OpenAI Model Distillation (Oct 2024), Azure OpenAI stored completions (2024) and Amazon Bedrock Model Distillation (GA May 2025) - while Google's Vertex Gemini distillation is so far documented only as a pre-GA, allowlist-only service that prohibits production use. Open-source toolkits from Hugging Face (TRL GKD, Open R1) and Arcee (DistillKit) commoditised the technique, and open-weight licenses from DeepSeek, Qwen, Mistral, Moonshot, MiniMax, Zhipu, NVIDIA and Hugging Face expressly permit derivative distillation. The same companies police distillation of their own outputs through terms-of-service clauses - OpenAI, Anthropic, Google, xAI and Cohere all bar training competing models - and the dispute escalated from OpenAI's January 2025 claims against DeepSeek to Anthropic's February 2026 report of 24,000 fraudulent accounts, its June 2026 letter to the Senate Banking Committee alleging a 28.8-million-interaction campaign by Alibaba, and White House memorandum NSTM-4 in April 2026. The split is not purely US-versus-China: in April 2026 Elon Musk conceded under oath that xAI had 'partly' used OpenAI's technology to train its own models. KEY FIGURES ----------- - Fraudulent accounts in Anthropic's Feb 2026 report: 24,000 accounts (16M+ exchanges) DeepSeek, Moonshot AI and MiniMax combined; MiniMax alone drove 13M+ exchanges Source: https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks - Largest single distillation attack reported (Alibaba/Qwen on Claude): 28,800,000 interactions (25,000 accounts in ~6 weeks) Per White House OSTP director Kratsios and Anthropic's June 2026 letter to the Senate; Alibaba banned Claude Code internally two weeks later Source: https://cyberscoop.com/white-house-accuses-moonshot-ai-anthropic-model-distillation/ - Companies profiled that publicly document a distilled model: 11 of 18 (7 undisclosed) Derived from the company-matrix table in this file: OpenAI, Anthropic, Amazon, xAI, Moonshot AI, MiniMax and Cohere do not publicly document a distilled model. Separately, OpenAI, Microsoft/Azure, Amazon, Google and NVIDIA sell distillation tooling. - Companies with explicit anti-distillation / anti-competing-model ToS clauses: 5 companies (OpenAI, Anthropic, Google, xAI, Cohere) Versus 9 permissive open-weight licenses in the tos-clauses table of this file that allow derivative distillation (Meta, DeepSeek, Alibaba, Moonshot, MiniMax, Mistral, NVIDIA, Microsoft/Phi, Hugging Face); Cohere Labs' Command R7B is CC-BY-NC and permits non-commercial derivatives only, so it is not counted. - GPU-hour savings of on-policy distillation vs RL (Qwen3-8B): 10 x fewer (1,800 vs 17,920 GPU hours) Qwen3 technical report Table 21; distilled model also scored higher on AIME'24 (74.4 vs 67.6) Source: https://arxiv.org/html/2505.09388v1 - Training tokens saved by prune-and-distill (NVIDIA Minitron): 160 x fewer (94B tokens vs 15T for the Llama 3.1 8B teacher) Llama-3.1-Minitron 4B produced from Llama 3.1 8B. NVIDIA separately claims up to 40x fewer training tokens per additional model when producing a family from one trained parent, and a 1.8x total compute saving. Source: https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model/ - Output price gap, flagship vs nano tier (OpenAI GPT-6 Astra vs GPT-5-nano): 125 x cheaper ($50.00 vs $0.40 per 1M output tokens) Current flagship gpt-6-astra against the cheapest listed nano tier; the cheapest current-generation small tier is gpt-5.6-luna at $1.20 per 1M output. OpenAI does not disclose training method for mini/nano tiers. Source: https://developers.openai.com/api/docs/pricing - Distillation-as-a-service products from hyperscalers: 3 shipped (+1 pre-GA) (Oct 2024 to 2025) Shipped: OpenAI Model Distillation (Oct 2024), Azure OpenAI stored completions + distillation (2024), Amazon Bedrock Model Distillation (preview Dec 2024, GA May 2025). Google Vertex Gemini distillation is documented only as pre-GA, allowlist-only with production use prohibited, so it is not counted as shipped. Source: https://aws.amazon.com/about-aws/whats-new/2025/05/amazon-bedrock-model-distillation-generally-available KEY FINDINGS ------------ 1. Distillation is now the default way every lab builds its small models Google's Gemini 2.5 report states plainly that 'the smaller models in the Gemini 2.5 series — Flash size and below — use distillation', and Gemma 2 and Gemma 3 are trained with knowledge distillation rather than plain next-token prediction. Meta pruned Llama 3.1 8B and distilled logits from 8B/70B to make Llama 3.2 1B/3B, then co-distilled Llama 4 Maverick from Behemoth. Qwen3, Ministral 3, Apple's on-device 3B, NVIDIA's Minitron/Nemotron Nano and DeepSeek's R1-Distill series all document the same pattern in their technical reports. Sources: https://arxiv.org/html/2507.06261v1/ https://ai.meta.com/blog/llama-4-multimodal-intelligence/ https://arxiv.org/html/2505.09388v1 https://arxiv.org/abs/2507.13575 2. Labs report distillation beats RL on cost and often on quality Qwen3's ablation shows on-policy distillation lifting Qwen3-8B to 74.4 on AIME'24 versus 67.6 for RL, using 1,800 rather than 17,920 GPU hours. DeepSeek reports that R1-Distill-Qwen-32B (72.6 AIME'24) beats OpenAI o1-mini (63.6) with SFT-only distillation and no RL stage. NVIDIA reports up to 40x fewer training tokens and Mistral reports Ministral 3 trained on 1-3T tokens versus 15-36T for comparable Qwen 3 / Llama 3 models. Sources: https://arxiv.org/html/2505.09388v1 https://huggingface.co/deepseek-ai/DeepSeek-R1 https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model/ https://www.deeplearning.ai/the-batch/mistral-uses-cascade-distillation-on-mistral-3-to-build-ministral-family/ 3. Closed labs sell distillation, but only inside their own model family OpenAI's Model Distillation (stored completions + evals + fine-tuning, Oct 2024) lets customers distill GPT-4o/o1-preview into GPT-4o mini; Azure mirrors it; Amazon Bedrock requires teacher and student to be from the same model family, so Nova Premier distills into Nova Pro/Lite/Micro, Claude 3.5 Sonnet v2 into Claude 3 Haiku, and Llama 3.3 70B / Llama 3.1 405B into Llama 3.2 1B/3B and Llama 3.1 70B/8B; Google's Vertex early-access service distills Gemini 3.1 Pro into Gemini 2.5 Flash. In every case the student must be a model the vendor hosts, so distillation revenue stays on-platform. Notably OpenAI's docs now say it is 'winding down the fine-tuning platform' for new users, and Azure retires stored completions on 2026-10-15. Sources: https://www.infoworld.com/article/3544913/openai-updates-api-with-model-distillation-prompt-caching-abilities.html https://aws.amazon.com/about-aws/whats-new/2025/05/amazon-bedrock-model-distillation-generally-available https://aws.amazon.com/blogs/aws/build-faster-more-cost-efficient-highly-accurate-models-with-amazon-bedrock-model-distillation-preview/ https://developers.openai.com/api/docs/guides/supervised-fine-tuning https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/stored-completions 4. Terms of service, not copyright, are the main legal lever against cross-lab distillation OpenAI forbids using 'Output to develop models that compete with OpenAI'; Anthropic's commercial terms bar access 'to build a competing product or service, including to train competing AI models'; Google's Gemini API terms say 'You may not use the Services to develop models that compete with the Services'; xAI's terms list 'distilling' among prohibited acts; Cohere bars use 'for the purpose of building a similar or competitive product or service'. Meta's Llama 4 license takes the opposite approach: derivative models are allowed but must carry 'Llama' at the start of their name. Sources: https://openai.com/policies/row-terms-of-use/ https://www.anthropic.com/legal/commercial-terms https://ai.google.dev/gemini-api/terms https://developer.meta.com/ai/llama4/license/ 5. Open-weight labs explicitly invite distillation in their licenses DeepSeek-R1's model card states the series 'allow for any modifications and derivative works, including, but not limited to, distillation for training other LLMs' under MIT. Qwen3, Mistral 3/Ministral 3 and SmolLM3 ship under Apache 2.0; Kimi K2 and MiniMax M2 use a modified MIT that only adds an attribution requirement above 100M MAU or $20M monthly revenue; NVIDIA releases Nemotron under its Open Model License and even lists the teacher models (DeepSeek-R1, GPT-OSS-120B, Qwen) used to synthesise 3.5T of its 10.6T pre-training tokens. Sources: https://huggingface.co/deepseek-ai/DeepSeek-R1 https://huggingface.co/moonshotai/Kimi-K2-Instruct/blob/main/LICENSE https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 https://mistral.ai/news/mistral-3/ 6. The accusation cycle escalated from one lab to a US government policy in 15 months OpenAI first alleged DeepSeek distillation in January 2025; the House Select Committee on the CCP's April 2025 report called it 'highly likely'. In February 2026 OpenAI told the committee DeepSeek used 'obfuscated third-party routers' and Anthropic published per-lab exchange counts (DeepSeek 150K, Moonshot 3.4M, MiniMax 13M). The framing is not purely US-versus-China: on 30 April 2026 Elon Musk conceded under oath in Musk v. OpenAI that xAI had 'partly' used OpenAI's technology to train its models. By April 2026 OpenAI, Anthropic and Google were also sharing threat intelligence through the Frontier Model Forum and the White House issued NSTM-4 calling the campaigns 'deliberate, industrial-scale'. In July 2026 OSTP director Kratsios accused Moonshot of distilling Anthropic's Fable model to build Kimi K3. Sources: https://assets.bwbx.io/documents/users/iqjWHBFdfxIU/rRmql_jJcxb4/v0 https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://www.nextgov.com/artificial-intelligence/2026/04/white-house-accuses-china-deliberate-industrial-scale-campaigns-steal-us-ai-models/413083/ https://cyberscoop.com/white-house-accuses-moonshot-ai-anthropic-model-distillation/ https://www.forbesafrica.com/current-affairs/2026/05/01/musk-admits-distilling-openai-data-for-his-xai-heres-why-thats-controversial 7. Defensive measures now shape products: ownership bans, hidden fingerprints, CoT hiding Anthropic barred entities more than 50% owned by companies in unsupported regions in September 2025, citing that they 'could also potentially use our models to advance their own AI development through techniques like distillation'. OpenAI's memo describes classifiers for 'reinforcement learning-style grading behavior', models 'trained not to reveal reasoning traces', and account bans. Anthropic admitted a March 2026 Claude Code 'experiment' that embedded identifying markers to protect against distillation; Alibaba responded by banning Claude Code for staff from 10 July 2026. Sources: https://www.anthropic.com/news/updating-restrictions-of-sales-to-unsupported-regions https://assets.bwbx.io/documents/users/iqjWHBFdfxIU/rRmql_jJcxb4/v0 https://techcrunch.com/2026/07/04/alibaba-reportedly-bans-employees-from-using-claude-code/ 8. Accused labs have not answered, and an illicit reseller market has emerged As of the Feb 2026 reports, DeepSeek, Moonshot and MiniMax had not responded to Anthropic's allegations; Moonshot did not respond to the July 2026 K3 claim. OpenAI's memo says Chinese companies 'rely on networks of unauthorized resellers of OpenAI's services to evade our platform's controls', and on 3 September 2026 Anthropic's head of threat intelligence Jacob Klein described 'an entire illicit ecosystem' on the dark web spinning up accounts at scale. Meanwhile Qwen and Zhipu were conspicuously absent from Anthropic's February list, before Alibaba was named in June. Sources: https://www.latent.space/p/ainews-anthropic-accuses-deepseek https://www.cnbc.com/2026/09/03/anthropic-distillation-battle-turns-to-dark-web-china-concerns-swell.html https://assets.bwbx.io/documents/users/iqjWHBFdfxIU/rRmql_jJcxb4/v0 TABLES -------- TABLE: Company distillation matrix: uses, sells, bans, accused, accuser [id: company-matrix, 18 rows] One row per company profiled. 'Bans' means the company's ToS or license restricts using its outputs to train competing models. 'Accused' and 'Accuser' refer to public distillation disputes as of 2026-09-03. Company | HQ | Uses distillation (documented) | Sells distillation tooling | Bans distillation of its outputs | Publicly accused | Public accuser | Open weights | Stance | Source URL ------- | --- | ------------------------ | ------------------------ | ------------------------ | ---------------- | -------------- | ------------ | ------ | ---------- OpenAI | US | Undisclosed (mini/nano tiers presumed; sells GPT-4o->4o-mini distillation) | Yes (Model Distillation, Oct 2024) | Yes | No | Yes (DeepSeek, Jan 2025 and Feb 2026) | gpt-oss only | restrictive | https://assets.bwbx.io/documents/users/iqjWHBFdfxIU/rRmql_jJcxb4/v0 Anthropic | US | Undisclosed (Haiku lineage) | Via Amazon Bedrock (Claude 3.5 Sonnet v2 as teacher) | Yes | No | Yes (DeepSeek, Moonshot, MiniMax, Alibaba) | No | restrictive | https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks Google DeepMind | US | Yes (Gemini Flash/Flash-Lite, Gemma 2/3) | Yes (Vertex distillation, early access 2026) | Yes | No | Partial (reported 'distillation attacks' on Gemini; joined FMF intel sharing) | Gemma | restrictive | https://arxiv.org/html/2507.06261v1/ Meta | US | Yes (Llama 3.2 1B/3B, Llama 4 Maverick/Scout) | No (teacher on Bedrock) | No, but derivative must be named 'Llama...' | No | No | Yes (Llama license) | mixed | https://developer.meta.com/ai/llama4/license/ DeepSeek | China | Yes (R1-Distill-Qwen/Llama, six models) | No | No (MIT, distillation expressly allowed) | Yes (OpenAI 2025/2026, House report 2025, Anthropic 2026, Gemini-similarity claims 2025) | No | Yes (MIT) | permissive | https://huggingface.co/deepseek-ai/DeepSeek-R1 Alibaba (Qwen) | China | Yes (Qwen3 strong-to-weak distillation) | No | No (Apache 2.0) | Yes (Anthropic, June 2026: 'largest known distillation attack') | No (banned Claude Code internally July 2026) | Yes (Apache 2.0) | permissive | https://arxiv.org/html/2505.09388v1 Microsoft | US | Yes (Phi family distilled from GPT-4; Phi-4-reasoning from o3-mini traces) | Yes (Azure OpenAI stored completions + distillation) | Azure OpenAI inherits OpenAI-style restrictions; Phi is MIT | No | No (FMF founding member) | Phi (MIT) | mixed | https://arxiv.org/abs/2412.08905 NVIDIA | US | Yes (Minitron, Nemotron Nano 2, Nemotron 3 Nano) | Yes (NeMo pruning/distillation recipes, TensorRT-LLM) | No (Nemotron Open Model License) | No | No | Yes | permissive | https://arxiv.org/abs/2508.14444 Amazon (AWS) | US | Undisclosed for Nova tiers | Yes (Bedrock Model Distillation, GA May 2025) | Undisclosed (AWS service terms) | No | No | No | mixed | https://aws.amazon.com/about-aws/whats-new/2025/05/amazon-bedrock-model-distillation-generally-available Mistral AI | France | Yes (Ministral 3 via cascade distillation from Mistral Small 3.1) | No | No (Apache 2.0 for Mistral 3 family) | No | No | Yes | permissive | https://arxiv.org/abs/2601.08584 Hugging Face | US/France | Yes (SmolLM3 synthetic traces from Qwen3-32B; OpenR1-Distill-7B from DeepSeek-R1) | Open-source tooling (TRL GKDTrainer, open-r1) | No (Apache 2.0) | No | No | Yes | permissive | https://github.com/huggingface/open-r1 Arcee AI | US | Yes (1.5B student from 7B Arcee-Agent; AFM-4.5B) | Open-source DistillKit (Apache 2.0) | No | No | No | Yes | permissive | https://github.com/arcee-ai/DistillKit xAI | US | Undisclosed (Grok mini/fast tiers; Grok 4 Fast described as RL, not distillation) | No | Yes ('distilling' listed as prohibited) | Yes (admitted under oath, 2026-04-30) | No | Older Grok-1 only | restrictive | https://x.ai/legal/terms-of-service Moonshot AI (Kimi) | China | Undisclosed | No | No (modified MIT) | Yes (Anthropic Feb 2026: 3.4M exchanges; White House July 2026: K3 distilled from Anthropic) | No | Yes | mixed | https://cyberscoop.com/white-house-accuses-moonshot-ai-anthropic-model-distillation/ MiniMax | China | Undisclosed | No | No (modified MIT) | Yes (Anthropic Feb 2026: 13M+ exchanges, largest of the three) | No | Yes | mixed | https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks Zhipu AI (Z.ai) | China | Partial (GLM-4.5 'expert model iteration' post-training) | No | No (MIT) | No (explicitly not named by Anthropic) | No | Yes | permissive | https://www.latent.space/p/ainews-anthropic-accuses-deepseek Apple | US | Yes (3B on-device model distilled from 64-expert MoE teacher) | No | N/A (no public API for its foundation models) | No | No | No | mixed | https://arxiv.org/abs/2507.13575 Cohere | Canada | Not publicly described (Command A uses self-refinement and model merging) | No | Yes ('building a similar or competitive product or service') | No | No | Command R7B (CC-BY-NC) | restrictive | https://cohere.com/terms-of-use Notes: Stance: restrictive = closed weights plus ToS ban on training competing models; permissive = open weights with license expressly allowing derivatives; mixed = open weights with conditions, or closed weights without a public ban, or an accused open-weight lab. Sources: https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://assets.bwbx.io/documents/users/iqjWHBFdfxIU/rRmql_jJcxb4/v0 https://www.cnbc.com/2026/06/24/anthropic-alibaba-distillation-campaign.html TABLE: Distilled model lineage: teacher to student, as documented by the companies [id: distilled-lineage, 22 rows] Only pairs that the releasing company (or its technical report) describes as distillation. Alleged cross-lab distillation is in the disputes table instead. Company | Teacher | Student | Student params (B) (B) | Method | Release | Source URL ------- | ------- | ------- | ---------------------- | ------ | ------- | ---------- Google | Gemini 1.5 Pro | Gemini 1.5 Flash | n/a | Distillation ('trained by 1.5 Pro through a process called distillation') | 2024-05 | https://blog.google/technology/ai/google-gemini-update-flash-ai-assistant-io-2024/ Google | Larger Gemma 2 / undisclosed | Gemma 2 2B, 9B | 9 | KD instead of next-token prediction (Hinton et al.) | 2024-07 | https://arxiv.org/abs/2408.00118 Arcee AI | Arcee-Agent 7B | 1.5B-Distilled | 1.5 | Logit + hidden-state distillation (DistillKit v0.1) | 2024-08 | https://arcee.ai/blog/announcing-distillkit/ NVIDIA | Llama 3.1 8B | Llama-3.1-Minitron 4B (width / depth) | 4 | Structured pruning + logit KD on 94B tokens | 2024-08 | https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model/ NVIDIA | Mistral NeMo 12B | Mistral-NeMo-Minitron 8B | 8 | Pruning + KD (Minitron) | 2024-08 | https://arxiv.org/abs/2408.11796 Meta | Llama 3.1 8B and 70B (logits) | Llama 3.2 1B, 3B | 3 | Single-shot structured pruning from 8B + logit KD in pre-training | 2024-09 | https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ OpenAI (customer-run) | GPT-4o, o1-preview | GPT-4o mini (fine-tuned) | n/a | Stored completions -> evals -> SFT (Model Distillation API) | 2024-10 | https://www.infoworld.com/article/3544913/openai-updates-api-with-model-distillation-prompt-caching-abilities.html Microsoft | GPT-4 | Phi-1 / Phi-2 / Phi-3 family | 14 | Synthetic 'textbook' data ('largely distill the capabilities of a teacher model (specifically GPT-4)') | 2023-2024 | https://arxiv.org/abs/2412.08905 DeepSeek | DeepSeek-R1 (671B MoE) | R1-Distill-Qwen 1.5B/7B/14B/32B; R1-Distill-Llama 8B/70B | 70 | SFT on ~800K R1 reasoning samples, 2-3 epochs, no RL | 2025-01 | https://arxiv.org/abs/2501.12948 Google | Undisclosed large teacher; 'large IT teacher' for post-training | Gemma 3 1B/4B/12B/27B | 27 | KD sampling 256 logits per token; 2T/4T/12T/14T tokens | 2025-03 | https://arxiv.org/html/2503.19786v1 Meta | Llama 4 Behemoth (~2T total, 288B active) | Llama 4 Maverick (400B total, 17B active); Scout | 400 | Codistillation with dynamically weighted soft/hard targets | 2025-04 | https://ai.meta.com/blog/llama-4-multimodal-intelligence/ Microsoft | OpenAI o3-mini (reasoning traces) | Phi-4-reasoning 14B | 14 | SFT on o3-mini demonstrations + RL | 2025-04 | https://arxiv.org/abs/2504.21318 Hugging Face | DeepSeek-R1 | OpenR1-Distill-7B | 7 | SFT on Mixture-of-Thoughts (350K traces) | 2025-05 | https://github.com/huggingface/open-r1 Alibaba | Qwen3-235B-A22B and Qwen3-32B | Qwen3 0.6B/1.7B/4B/8B/14B, 30B-A3B | 30 | Strong-to-weak: off-policy response distillation, then on-policy logit KL | 2025-05 | https://arxiv.org/html/2505.09388v1 Google | Larger Gemini 2.5 models | Gemini 2.5 Flash, Flash-Lite | n/a | Distillation with k-sparse teacher distribution | 2025-06 | https://arxiv.org/html/2507.06261v1/ Hugging Face | Qwen3-32B (synthetic reasoning traces) | SmolLM3 3B | 3 | Synthetic data generation + SFT + APO | 2025-07 | https://huggingface.co/blog/smollm3 Apple | 64-expert sparse-upcycled MoE (from 14T-token dense model) | On-device ~3B model | 3 | Distillation loss for last 10% (~1.4T) of tokens; teacher cost cut 90% | 2025-07 | https://arxiv.org/abs/2507.13575 NVIDIA | Nemotron-Nano-12B-v2-Base (20T tokens) | Nemotron-Nano-9B-v2 | 9 | Minitron pruning + distillation | 2025-08 | https://arxiv.org/abs/2508.14444 Mistral AI | Mistral Small 3.1 (24B) | Ministral 3 3B/8B/14B | 14 | Cascade distillation: iterative pruning + continued training with distillation (1-3T tokens) | 2025-12 | https://arxiv.org/abs/2601.08584 NVIDIA | DeepSeek-R1, GPT-OSS-120B, Qwen models (synthetic) | Nemotron 3 Nano 30B-A3B | 30 | ~3.5T of 10.6T pre-training tokens synthesised from teachers | 2025-12 | https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Amazon (customer-run) | Nova Premier; Claude 3.5 Sonnet v2; Llama 3.3 70B / Llama 3.1 405B | Nova Pro/Lite/Micro; Claude 3 Haiku; Llama 3.2 1B/3B, Llama 3.1 70B/8B | 3 | Bedrock synthetic data generation + fine-tuning (teacher and student must be from the same model family) | 2025-05 | https://aws.amazon.com/about-aws/whats-new/2025/05/amazon-bedrock-model-distillation-generally-available Google (customer-run) | Gemini 3.1 Pro | Gemini 2.5 Flash (custom) | n/a | Vertex distillation service (pre-GA, allowlist) | 2026-07 | https://runtimewire.com/article/google-cloud-page-describes-gemini-distillation-service-but-its-release-status-i Notes: Parameter counts for Gemini and GPT tiers are undisclosed. Phi row uses Phi-4's 14B as representative size; the Phi-4 report says phi-4 itself 'substantially surpasses its teacher model'. Amazon Bedrock Model Distillation only permits same-family teacher/student pairs (AWS: 'The teacher and the student model must be from the same family'), so each Amazon teacher is paired with a student of its own family. Sources: https://arxiv.org/html/2507.06261v1/ https://arxiv.org/abs/2501.12948 https://arxiv.org/html/2505.09388v1 TABLE: Terms-of-service and license clauses governing distillation [id: tos-clauses, 18 rows] Brief verbatim quotes from each company's governing document, with effective/last-updated date where the page states it. Company | Document | Clause (quoted) | Effect on distillation | Date | Source URL ------- | -------- | --------------- | ---------------------- | ---- | ---------- OpenAI | Terms of Use (ROW) | "use Output to develop models that compete with OpenAI" (listed under what you cannot do) | Prohibited for competing models; OpenAI's own Distillation API is the sanctioned path | current | https://openai.com/policies/row-terms-of-use/ Anthropic | Commercial Terms of Service, D.4 | "access the Services to build a competing product or service, including to train competing AI models or resell the Services" | Prohibited | 2025-06-17 | https://www.anthropic.com/legal/commercial-terms Anthropic | Consumer Terms of Service, s.3 | "To develop any products or services that compete with our Services, including to develop or train any artificial intelligence or machine learning algorithms or models" | Prohibited | 2025-10-08 | https://www.anthropic.com/legal/consumer-terms Google | Gemini API Additional Terms, Use Restrictions | "You may not use the Services to develop models that compete with the Services (e.g., Gemini API or Google AI Studio)" | Prohibited; also bars extracting 'parameter weights' | 2026-04-28 | https://ai.google.dev/gemini-api/terms xAI | Terms of Service - Consumer; Acceptable Use Policy | Prohibits "distilling" the Service and using "the Service or Output to develop models or services that compete with xAI" | Prohibited (distillation named explicitly) | undisclosed (page not fetchable) | https://x.ai/legal/acceptable-use-policy Cohere | Terms of Use, s.14(12) | "for the purpose of building a similar or competitive product or service" | Prohibited | 2022-09-07 | https://cohere.com/terms-of-use Meta | Llama 4 Community License | "If you use the Llama Materials or any outputs ... to create, train, fine tune, or otherwise improve an AI model ... you shall also include 'Llama' at the beginning of any such AI model name" | Allowed with naming + 'Built with Llama' attribution; >700M MAU needs a license | 2025-04-05 | https://developer.meta.com/ai/llama4/license/ DeepSeek | DeepSeek-R1 model card (MIT) | "allow for any modifications and derivative works, including, but not limited to, distillation for training other LLMs" | Expressly allowed | 2025-01 | https://huggingface.co/deepseek-ai/DeepSeek-R1 Alibaba (Qwen) | Qwen3 model cards | Apache 2.0 | Allowed | 2025-05 | https://huggingface.co/Qwen/Qwen3-235B-A22B Moonshot AI | Kimi K2 Modified MIT License | "more than 100 million monthly active users, or more than 20 million US dollars ... in monthly revenue, you shall prominently display 'Kimi K2'" | Allowed with attribution above thresholds | 2025-07 | https://huggingface.co/moonshotai/Kimi-K2-Instruct/blob/main/LICENSE MiniMax | MiniMax-M2 model card | License: modified-mit | Allowed | 2025 | https://huggingface.co/MiniMaxAI/MiniMax-M2 Mistral AI | Mistral 3 release | "All models are released under the Apache 2.0 license" | Allowed | 2025-12-02 | https://mistral.ai/news/mistral-3/ NVIDIA | NVIDIA Nemotron Open Model License | Open model license; model card lists teacher models used for synthetic data | Allowed | 2025-12-15 | https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Microsoft | Phi-4-mini model card | MIT license | Allowed for Phi weights; Azure OpenAI service outputs governed separately | 2025-02 | https://huggingface.co/microsoft/Phi-4-mini-instruct Hugging Face | SmolLM3 / TRL | Apache 2.0 | Allowed; TRL ships GKDTrainer for on-policy distillation | 2025-07-08 | https://huggingface.co/blog/smollm3 Zhipu AI (Z.ai) | GLM-4.5 model card | MIT license | Allowed | 2025-07 | https://huggingface.co/zai-org/GLM-4.5 Cohere Labs | Command R7B model card | CC-BY-NC plus Acceptable Use Policy | Non-commercial derivatives only | 2024-12 | https://huggingface.co/CohereLabs/c4ai-command-r7b-12-2024 Amazon | AWS Service Terms (Bedrock) | undisclosed | Distillation sold in-platform; cross-platform terms not verified | undisclosed | https://aws.amazon.com/about-aws/whats-new/2025/05/amazon-bedrock-model-distillation-generally-available Notes: OpenAI's and xAI's ToS pages returned HTTP 403 to automated fetching; OpenAI's clause is corroborated by third-party legal commentary and xAI's by search-index text. The effective date of xAI's terms could not be verified and is recorded as undisclosed; the 'distilling' prohibition also appears in xAI's separate Acceptable Use Policy. Quotes are kept under 25 words. Sources: https://ospo.co/blog/be-careful-with-openais-terms-of-use/ https://www.anthropic.com/legal/commercial-terms https://ai.google.dev/gemini-api/terms https://x.ai/legal/acceptable-use-policy TABLE: Distillation products and toolkits offered by companies [id: distillation-products, 9 rows] Commercial services and open-source tools that let third parties run distillation. Product | Vendor | Launched | Teachers | Students | Status (2026-09) | Source URL ------- | ------ | -------- | -------- | -------- | ---------------- | ---------- Model Distillation (Stored Completions + Evals + Fine-tuning) | OpenAI | 2024-10-01 | GPT-4o, o1-preview, later gpt-4.1 | GPT-4o mini, gpt-4.1-mini | Docs state fine-tuning platform is 'winding down' for new users | https://developers.openai.com/api/docs/guides/supervised-fine-tuning Stored completions & distillation (Azure OpenAI / Foundry classic) | Microsoft | 2024 | Any Azure OpenAI chat model (e.g. gpt-4o) | Azure OpenAI fine-tunable models | Stored completions retire 2026-10-15 | https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/stored-completions Amazon Bedrock Model Distillation (preview) | Amazon | 2024-12-03 | Nova Premier; Claude 3.5 Sonnet v2; Llama 3.1 405B / 70B | Nova Lite/Micro; Claude 3 Haiku; Llama 3.1 70B/8B, Llama 3.2 1B/3B | Superseded by GA; preview promised 'up to 500% faster and 75% less expensive' | https://aws.amazon.com/blogs/aws/build-faster-more-cost-efficient-highly-accurate-models-with-amazon-bedrock-model-distillation-preview/ Amazon Bedrock Model Distillation (GA) | Amazon | 2025-05-01 | Nova Premier, Claude 3.5 Sonnet v2, Llama 3.3 70B | Nova Pro, Llama 3.2 1B/3B | GA; 'up to 500% faster and 75% less expensive ... less than 2% accuracy loss' | https://aws.amazon.com/about-aws/whats-new/2025/05/amazon-bedrock-model-distillation-generally-available Gemini distillation (Vertex AI / Gemini Enterprise Agent Platform) | Google | 2026-07 (docs) | Gemini 3.1 Pro | Gemini 2.5 Flash | Pre-GA, allowlist only, no production use | https://runtimewire.com/article/google-cloud-page-describes-gemini-distillation-service-but-its-release-status-i NeMo pruning + distillation (Minitron recipes), TensorRT-LLM | NVIDIA | 2024-08-14 | Llama 3.1 8B, Mistral NeMo 12B, Nemotron 12B | 4B-9B pruned students | Open recipes; Nemotron Nano 2/3 built with them | https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model/ TRL GKDTrainer (Generalized Knowledge Distillation) | Hugging Face | 2024 | Any HF causal LM | Any HF causal LM | Experimental module in TRL v1.12; lmbda/beta/seq_kd controls | https://huggingface.co/docs/trl/gkd_trainer Open R1 (open reproduction of DeepSeek-R1 distillation) | Hugging Face | 2025-01 | DeepSeek-R1 | Qwen2.5-based 1.5B-7B | OpenR1-Math-220k (Feb 2025), Mixture-of-Thoughts 350K (May 2025) | https://github.com/huggingface/open-r1 DistillKit | Arcee AI | 2024-08-01 | Any (online or offline logits) | Any (cross-architecture via hidden-state loss) | Apache 2.0; logit compression via polynomial approximation + quantization | https://github.com/arcee-ai/DistillKit Notes: Every commercial service restricts the student to a model hosted on the same platform. Amazon Bedrock additionally requires the teacher and student to be from the same model family, and its row is split because the preview and GA announcements list different teacher/student sets. Sources: https://www.infoworld.com/article/3544913/openai-updates-api-with-model-distillation-prompt-caching-abilities.html https://aws.amazon.com/blogs/aws/build-faster-more-cost-efficient-highly-accurate-models-with-amazon-bedrock-model-distillation-preview/ TABLE: Public distillation accusations between companies [id: disputes, 11 rows] Who accused whom, the evidence disclosed, and the outcome so far. Date | Accuser | Accused | Exchanges / queries claimed | Accounts claimed | Claim | Outcome | Source URL ---- | ------- | ------- | ------------------------ | ---------------- | ----- | ------- | ---------- 2025-01 | OpenAI | DeepSeek | n/a | n/a | DeepSeek distilled OpenAI outputs in violation of ToS; Microsoft flagged API exfiltration | House Select Committee report (April 2025) found it 'highly likely'; no lawsuit | https://sites.law.berkeley.edu/thenetwork/2025/03/30/the-innovation-dilemma-ai-distillation-in-openai-v-deepseek/ 2025-06-03 | Independent researchers (EQ-Bench, SpeechMap) | DeepSeek (R1-0528) | n/a | n/a | Outputs stylistically resemble Gemini 2.5 Pro | Suggestive only; Google did not comment | https://winbuzzer.com/2025/06/03/is-deepseek-training-its-ai-with-data-from-google-gemini-new-distillation-claims-emerge-xcxwbn/ 2026-02-12 | OpenAI (memo to House Select Committee) | DeepSeek | n/a | n/a | 'obfuscated third-party routers', programmatic extraction code, unauthorized reseller networks | Closed-door briefing offered; DeepSeek V4 shipped April-August 2026 regardless | https://assets.bwbx.io/documents/users/iqjWHBFdfxIU/rRmql_jJcxb4/v0 2026-02-23 | Anthropic | DeepSeek | 150000 | n/a | Targeted agentic reasoning, reward modeling and censorship-safe alternatives | No response from DeepSeek | https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks 2026-02-23 | Anthropic | Moonshot AI | 3400000 | n/a | Targeted computer-use agents and vision | No response; later named by White House over Kimi K3 | https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks 2026-02-23 | Anthropic | MiniMax | 13000000 | n/a | Agentic coding and tool orchestration; redirected nearly half of traffic to a new Claude model within 24h | No response | https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks 2026-02-23 | Anthropic (aggregate) | DeepSeek + Moonshot + MiniMax | 16000000 | 24000 | 'industrial-scale distillation attacks' | Intel shared with industry and authorities | https://cyberscoop.com/anthropic-accuses-chinese-labs-ai-distillation-cyber-risk/ 2026-02 | Google (Threat Intelligence Group) | Suspected state-aligned actors (China, Russia, North Korea) plus commercial firms and researchers | 100000 | n/a | Gemini hit with 100,000+ structured prompts in an apparent cloning attempt; Google classifies model extraction as IP theft | Google joined FMF intel sharing (April 2026) | https://www.nbcnews.com/tech/security/google-gemini-hit-100000-prompts-cloning-attempt-rcna258657 2026-04-30 | OpenAI litigation / sworn testimony (Musk v. OpenAI) | xAI | n/a | n/a | Musk conceded under cross-examination that xAI had 'partly' used OpenAI's technology: 'Generally A.I. companies distill other A.I. companies.' | On-the-record admission by a company whose own terms prohibit distilling; litigation ongoing | https://www.forbesafrica.com/current-affairs/2026/05/01/musk-admits-distilling-openai-data-for-his-xai-heres-why-thats-controversial 2026-06-10 | Anthropic (letter to US Senate Banking Committee: Chair Tim Scott, Ranking Member Elizabeth Warren) | Alibaba (Qwen) | 28800000 | 25000 | 'the largest known distillation attack' on Anthropic; ~6 weeks (22 April - 5 June) | Reported by CNBC 2026-06-24; Alibaba banned Claude Code for staff from 2026-07-10 | https://www.cnbc.com/2026/06/24/anthropic-alibaba-distillation-campaign.html 2026-07-22 | White House OSTP (Kratsios) | Moonshot AI | n/a | n/a | Distilled Anthropic's Fable model to build Kimi K3 ('first open 2.8 trillion parameter model') using an internal platform and GB300 servers | Moonshot did not respond | https://cyberscoop.com/white-house-accuses-moonshot-ai-anthropic-model-distillation/ Notes: Exchange counts are the accusers' figures and have not been independently verified. The Google figure is a count of structured prompts reported by Google Threat Intelligence Group (via NBC News), not a per-lab exchange total comparable to Anthropic's. Sources: https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://assets.bwbx.io/documents/users/iqjWHBFdfxIU/rRmql_jJcxb4/v0 https://cyberscoop.com/white-house-accuses-moonshot-ai-anthropic-model-distillation/ TABLE: Distillation efficiency claims made in company reports [id: efficiency-claims, 17 rows] Quantified benefits each company attributes to distillation, taken from its own technical report or announcement. Company | Model | Metric | Distilled | Baseline | Unit | Source URL ------- | ----- | ------ | --------- | -------- | ---- | ---------- Alibaba | Qwen3-8B | GPU hours (on-policy distillation vs RL) | 1800 | 17920 | GPU hours | https://arxiv.org/html/2505.09388v1 Alibaba | Qwen3-8B | AIME'24 (on-policy distillation vs RL) | 74.4 | 67.6 | % | https://arxiv.org/html/2505.09388v1 Alibaba | Qwen3-8B | LiveCodeBench (on-policy distillation vs RL) | 60.3 | 52.9 | % | https://arxiv.org/html/2505.09388v1 NVIDIA | Llama-3.1-Minitron 4B | Training tokens (student vs Llama 3.1 8B teacher) | 94 | 15000 | B tokens | https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model/ NVIDIA | Llama-3.1-Minitron 4B (width) | MMLU (student vs 8B teacher) | 60.5 | 65.3 | % | https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model/ NVIDIA | Llama-3.1-Minitron 4B (depth) | Throughput vs Llama 3.1 8B (TensorRT-LLM, H100) | 2.7 | 1 | x | https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model/ Mistral AI | Ministral 3 (3B-14B) | Training tokens vs Qwen 3 / Llama 3 peers (low end) | 1000 | 15000 | B tokens | https://www.deeplearning.ai/the-batch/mistral-uses-cascade-distillation-on-mistral-3-to-build-ministral-family/ Mistral AI | Ministral 3 14B reasoning | AIME 2025 vs Qwen 3 14B Thinking | 85 | 73.7 | % | https://www.deeplearning.ai/the-batch/mistral-uses-cascade-distillation-on-mistral-3-to-build-ministral-family/ Apple | On-device ~3B | Teacher training cost reduction from new pipeline | 90 | 0 | % saved | https://arxiv.org/abs/2507.13575 Amazon | Bedrock distilled students | Latency (AWS: 'up to 500% faster' than the original model) | 600 | 100 | % of baseline speed | https://aws.amazon.com/about-aws/whats-new/2025/05/amazon-bedrock-model-distillation-generally-available Amazon | Bedrock distilled students | Cost reduction (up to), accuracy loss <2% for RAG | 75 | 0 | % cheaper | https://aws.amazon.com/about-aws/whats-new/2025/05/amazon-bedrock-model-distillation-generally-available DeepSeek | R1-Distill-Qwen-32B vs OpenAI o1-mini | AIME 2024 | 72.6 | 63.6 | % | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek | R1-Distill-Llama-70B vs OpenAI o1-mini | MATH-500 | 94.5 | 90 | % | https://huggingface.co/deepseek-ai/DeepSeek-R1 Hugging Face | OpenR1-Distill-7B vs DeepSeek-R1-Distill-Qwen-7B (reproduction) | AIME 2024 | 52.7 | 51.3 | % | https://github.com/huggingface/open-r1 Hugging Face | OpenR1-Distill-7B vs DeepSeek-R1-Distill-Qwen-7B (reproduction) | MATH-500 | 89 | 93.5 | % | https://github.com/huggingface/open-r1 Google | Gemma 3 4B-IT vs Gemma 2 27B-IT | Competitive per report (distillation-trained 4B matches prior 27B) | 4 | 27 | B params for similar IT quality | https://arxiv.org/html/2503.19786v1 NVIDIA | Nemotron-Nano-9B-v2 vs Qwen3-8B | Inference throughput in reasoning settings (up to) | 6 | 1 | x | https://arxiv.org/abs/2508.14444 Notes: Open R1 rows use Open R1's own like-for-like re-evaluation of both models under one harness (AIME'24 52.7 vs 51.3; MATH-500 89.0 vs 93.5); DeepSeek's self-reported AIME'24 figure for DeepSeek-R1-Distill-Qwen-7B is 55.5, measured under a different setup. Gemma 3 row is qualitative from the report's claim that Gemma3-4B-IT is competitive with Gemma2-27B-IT. Sources: https://arxiv.org/html/2505.09388v1 https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model/ https://huggingface.co/deepseek-ai/DeepSeek-R1 TABLE: Flagship vs small-tier API pricing at companies that build small tiers by distillation (or undisclosed) [id: small-tier-pricing, 16 rows] Per-million-token prices as listed on vendor pricing pages, 2026-09-03. Distillation status per vendor disclosure. Vendor | Model | Tier | Input $/1M (USD) | Output $/1M (USD) | Distilled? | Source URL ------ | ----- | ---- | ---------------- | ----------------- | ---------- | ---------- OpenAI | GPT-6 Astra (gpt-6-astra) | flagship | 10 | 50 | n/a | https://developers.openai.com/api/docs/pricing OpenAI | GPT-5.6 Sol (gpt-5.6-sol) | mid | 4 | 20 | undisclosed | https://developers.openai.com/api/docs/pricing OpenAI | GPT-5.6 Terra (gpt-5.6-terra) | mid | 2 | 12 | undisclosed | https://developers.openai.com/api/docs/pricing OpenAI | GPT-5.6 Luna (gpt-5.6-luna) | small | 0.2 | 1.2 | undisclosed | https://developers.openai.com/api/docs/pricing OpenAI | GPT-5.5 | prior flagship | 5 | 30 | n/a | https://developers.openai.com/api/docs/pricing OpenAI | GPT-5 | prior flagship | 1.25 | 10 | n/a | https://developers.openai.com/api/docs/pricing OpenAI | GPT-5 mini | small | 0.25 | 2 | undisclosed | https://developers.openai.com/api/docs/pricing OpenAI | GPT-5 nano | small | 0.05 | 0.4 | undisclosed | https://developers.openai.com/api/docs/pricing OpenAI | GPT-5.4 mini | small | 0.75 | 4.5 | undisclosed | https://developers.openai.com/api/docs/pricing OpenAI | GPT-5.4 nano | small | 0.2 | 1.25 | undisclosed | https://developers.openai.com/api/docs/pricing OpenAI | GPT-4o | flagship (2024) | 2.5 | 10 | n/a (teacher in Distillation API) | https://developers.openai.com/api/docs/pricing OpenAI | GPT-4o mini | small | 0.15 | 0.6 | undisclosed (student in Distillation API) | https://developers.openai.com/api/docs/pricing OpenAI | o4-mini | small reasoning | 1.1 | 4.4 | undisclosed | https://developers.openai.com/api/docs/pricing OpenAI | o3 | flagship reasoning | 2 | 8 | n/a | https://developers.openai.com/api/docs/pricing Anthropic | Claude Haiku 4.5 | small | 1 | 5 | undisclosed ('one-third the cost' of Sonnet 4) | https://www.anthropic.com/news/claude-haiku-4-5 xAI | Grok 4 Fast (<128k) | small | 0.2 | 0.5 | No per xAI (RL; '98% reduction in price' vs Grok 4) | https://x.ai/news/grok-4-fast Notes: gpt-6-astra is OpenAI's current flagship (listed as 'rolling out today' on the pricing page on 2026-09-03); GPT-5.5 and GPT-5 are prior flagships kept for comparison. Google Gemini Flash prices are covered in the customer/financial perspectives; Google does confirm Flash tiers are distilled. Sources: https://developers.openai.com/api/docs/pricing https://www.anthropic.com/news/claude-haiku-4-5 https://x.ai/news/grok-4-fast CHART DATA ---------- CHART: 18 companies by distillation stance [id: stance-donut, type: donut, unit: companies] Stance | Stance (companies) Restrictive (closed weights + ToS ban) | 5 Permissive (open weights, distillation allowed) | 7 Mixed | 6 Notes: Restrictive: OpenAI, Anthropic, Google, xAI, Cohere. Permissive: DeepSeek, Alibaba/Qwen, NVIDIA, Mistral, Hugging Face, Arcee, Zhipu. Mixed: Meta, Microsoft, Amazon, Moonshot, MiniMax, Apple. Classification from the company-matrix table. Sources: https://ai.google.dev/gemini-api/terms https://huggingface.co/deepseek-ai/DeepSeek-R1 https://developer.meta.com/ai/llama4/license/ CHART: Documented distilled model families released per company per year [id: distilled-releases-by-year, type: stackedBar, unit: releases] Year | Google (releases) 2024 | 2 2025 | 2 2026 | 0 Year | NVIDIA (releases) 2024 | 2 2025 | 2 2026 | 0 Year | Meta (releases) 2024 | 1 2025 | 1 2026 | 0 Year | Microsoft (releases) 2024 | 1 2025 | 1 2026 | 0 Year | Hugging Face (releases) 2024 | 0 2025 | 2 2026 | 0 Year | DeepSeek (releases) 2024 | 0 2025 | 1 2026 | 0 Year | Alibaba (Qwen) (releases) 2024 | 0 2025 | 1 2026 | 0 Year | Apple (releases) 2024 | 0 2025 | 1 2026 | 0 Year | Mistral (releases) 2024 | 0 2025 | 1 2026 | 0 Year | Amazon (service) (releases) 2024 | 0 2025 | 1 2026 | 0 Year | OpenAI (service) (releases) 2024 | 1 2025 | 0 2026 | 0 Year | Arcee (releases) 2024 | 1 2025 | 0 2026 | 0 Notes: Counts are rows in the distilled-lineage table (one per teacher->student family, dated by release). 2026 is year-to-date through 3 September. Google's Vertex distillation service (Gemini 3.1 Pro -> 2.5 Flash) is excluded because it is pre-GA, allowlist-only and not a released distilled model family. Only company-documented distillations are counted, so OpenAI/Anthropic/xAI small tiers (undisclosed method) are excluded. Sources: https://arxiv.org/html/2507.06261v1/ https://arxiv.org/abs/2501.12948 https://arxiv.org/abs/2601.08584 CHART: Distillation product and toolkit launches (scatter by date) [id: product-launch-timeline, type: scatter] Launch date | Launches 2024 | 1 2024 | 2 2024-08-01 | 3 2024-08-14 | 4 2024-10-01 | 5 2024-12-03 | 6 2025-01-25 | 7 2025-05-01 | 8 2026-07-23 | 9 Launch date | Retirements 2026-10-15 | 1 Notes: One point per row of the distillation-products table. Azure stored completions and TRL GKDTrainer are known only to year precision (month not disclosed) and are plotted at 2024. Open R1 dated to late January 2025 per the repository. The scheduled Azure retirement is a separate series because it is not a launch; its date is from Microsoft Learn. Sources: https://arcee.ai/blog/announcing-distillkit/ https://aws.amazon.com/blogs/aws/build-faster-more-cost-efficient-highly-accurate-models-with-amazon-bedrock-model-distillation-preview/ https://huggingface.co/docs/trl/gkd_trainer https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/stored-completions CHART: Exchanges attributed to each accused lab in Anthropic's reports [id: attack-exchanges, type: bar, unit: M exchanges] Accused lab | Exchanges with Claude (M exchanges) DeepSeek (Feb 2026) | 0.15 Moonshot AI (Feb 2026) | 3.4 MiniMax (Feb 2026) | 13 Alibaba / Qwen (Jun 2026) | 28.8 Notes: Feb 2026 figures from Anthropic's 'Detecting and preventing distillation attacks'; Alibaba figure from Anthropic's June 2026 Senate letter as reported by CNBC and CyberScoop (25,000 accounts over ~6 weeks). Sources: https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://cyberscoop.com/white-house-accuses-moonshot-ai-anthropic-model-distillation/ CHART: Qwen3-8B: on-policy distillation vs reinforcement learning (Qwen3 report Table 21) [id: qwen3-distill-vs-rl, type: bar, unit: %] Benchmark | Off-policy distillation only (%) AIME'24 | 55 AIME'25 | 42.8 MATH500 | 92.4 LiveCodeBench | 42 Benchmark | + RL (17,920 GPU h) (%) AIME'24 | 67.6 AIME'25 | 55.5 MATH500 | 94.8 LiveCodeBench | 52.9 Benchmark | + On-policy distillation (1,800 GPU h) (%) AIME'24 | 74.4 AIME'25 | 65.5 MATH500 | 97 LiveCodeBench | 60.3 Notes: Alibaba's own ablation; teacher is Qwen3-32B / Qwen3-235B-A22B. The report concludes distillation needs 'approximately only 1/10 of the GPU hours'. Sources: https://arxiv.org/html/2505.09388v1 CHART: Output price: flagship vs small tier at OpenAI and Anthropic [id: flagship-vs-small-output-price, type: bar, unit: USD] Model | Output $/1M (USD) GPT-6 Astra | 50 GPT-5.5 | 30 GPT-5.6 Sol | 20 GPT-5.6 Terra | 12 GPT-5.6 Luna | 1.2 GPT-5 | 10 GPT-5 mini | 2 GPT-5 nano | 0.4 GPT-4o | 10 GPT-4o mini | 0.6 o3 | 8 o4-mini | 4.4 Claude Haiku 4.5 | 5 Grok 4 Fast | 0.5 Notes: Prices from the OpenAI pricing page on 2026-09-03; gpt-6-astra is the current flagship. OpenAI and Anthropic do not disclose whether mini/nano/Haiku tiers are distilled; OpenAI's own Distillation API positions GPT-4o mini as the student of GPT-4o. xAI says Grok 4 Fast was built with RL rather than distillation. Sources: https://developers.openai.com/api/docs/pricing https://www.anthropic.com/news/claude-haiku-4-5 https://x.ai/news/grok-4-fast CHART: Pre-training tokens: distilled student vs from-scratch teacher or peer [id: training-tokens-student-vs-teacher, type: bar, unit: T tokens] Model | Distilled student (T tokens) Llama-3.1-Minitron 4B | 0.094 Ministral 3 (low end) | 1 Ministral 3 (high end) | 3 Gemma 3 1B | 2 Gemma 3 4B | 4 Apple on-device 3B (distill phase) | 1.4 Model | Teacher / from-scratch peer (T tokens) Llama-3.1-Minitron 4B | 15 Ministral 3 (low end) | 15 Ministral 3 (high end) | 36 Gemma 3 1B | 14 Gemma 3 4B | 14 Apple on-device 3B (distill phase) | 14 Notes: Gemma 3 comparison uses the 27B sibling's 14T tokens. Apple: dense model trained ~14T, then last 10% (~1.4T) retrained with distillation loss. Ministral peers are Qwen 3 / Llama 3 models of similar size per Mistral (15-36T). Sources: https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model/ https://www.deeplearning.ai/the-batch/mistral-uses-cascade-distillation-on-mistral-3-to-build-ministral-family/ https://arxiv.org/html/2503.19786v1 https://arxiv.org/abs/2507.13575 DETAIL RECORDS -------------- COMPANY PROFILES (18) - OpenAI (San Francisco, US) — stance: restrictive Distillation products: Model Distillation in the API (Stored Completions + Evals + Fine-tuning), launched 2024-10-01; docs now state fine-tuning platform is winding down for new users Distilled models: GPT-4o mini (student in OpenAI's own Distillation API; training method undisclosed); GPT-5 mini / GPT-5 nano / GPT-5.4 mini / nano (undisclosed); o1-mini / o3-mini / o4-mini (undisclosed) Terms-of-service clause: "use Output to develop models that compete with OpenAI" — Terms of Use, https://openai.com/policies/row-terms-of-use/ Notable events: 2024-10-01: Launches Model Distillation with free training tokens promotion | 2025-01: Alleges DeepSeek distilled its models; Microsoft probes API exfiltration | 2025-03: Provides assessment of DeepSeek distillation to House Select Committee | 2026-02-12: Memo to House Select Committee describing 'obfuscated third-party routers', reseller networks, RL-style grading classifiers and CoT-hiding defenses; states 'we do not allow our outputs to be used to create imitation frontier AI models' | 2026-04-07: Joins Anthropic and Google in sharing distillation intel via Frontier Model Forum Source: https://assets.bwbx.io/documents/users/iqjWHBFdfxIU/rRmql_jJcxb4/v0 - Anthropic (San Francisco, US) — stance: restrictive Distillation products: None first-party; Claude 3.5 Sonnet v2 is offered as a teacher inside Amazon Bedrock Model Distillation Distilled models: Claude Haiku lineage (Haiku 3 / 3.5 / 4.5) — training method undisclosed; Haiku 4.5 priced $1/$5 at 'one-third the cost' of Sonnet 4 Terms-of-service clause: "access the Services to build a competing product or service, including to train competing AI models" — Commercial Terms D.4, https://www.anthropic.com/legal/commercial-terms Notable events: 2025-09-04: Bars entities >50% owned from unsupported regions, citing distillation risk | 2026-02-23: Reports 24,000 fraudulent accounts and 16M+ exchanges by DeepSeek (150K), Moonshot (3.4M), MiniMax (13M) | 2026-03: Runs Claude Code fingerprinting 'experiment' against resellers/distillation (disclosed July 2026) | 2026-06-24: Letter to US Senate calls Alibaba's ~25,000-account, 28.8M-interaction campaign 'the largest known distillation attack' | 2026-09-03: Threat-intel head Jacob Klein describes dark-web reseller ecosystem Source: https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks - Google DeepMind (Mountain View / London) — stance: restrictive Distillation products: Vertex AI / Gemini Enterprise Agent Platform distillation: Gemini 3.1 Pro teacher -> Gemini 2.5 Flash student (pre-GA, allowlist, docs updated 2026-07-23) Distilled models: Gemini 1.5 Flash (distilled from 1.5 Pro); Gemini 2.5 Flash and Flash-Lite ('Flash size and below — use distillation'); Gemma 2 2B/9B (KD instead of next-token prediction); Gemma 3 1B/4B/12B/27B (256 sampled logits per token; post-training from 'a large IT teacher') Terms-of-service clause: "You may not use the Services to develop models that compete with the Services (e.g., Gemini API or Google AI Studio)" — https://ai.google.dev/gemini-api/terms (updated 2026-04-28) Notable events: 2024-05-14: First lab to publicly call a production tier (1.5 Flash) distilled | 2025-06-03: Third-party claims DeepSeek R1-0528 resembles Gemini 2.5 Pro; Google begins summarising reasoning traces | 2026-02: Reports Gemini targeted by distillation/model-extraction campaigns (100,000+ queries) and classifies distillation as IP theft (secondary reporting) | 2026-04-07: Joins FMF distillation intel sharing Source: https://arxiv.org/html/2507.06261v1/ - Meta (Menlo Park, US) — stance: mixed Distillation products: None first-party; Llama 3.1 405B / 3.3 70B offered as teachers on Amazon Bedrock Distilled models: Llama 3.2 1B/3B (pruned from Llama 3.1 8B; logits from 8B and 70B as token-level targets); Llama 4 Maverick 400B/17B-active and Scout (codistilled from Behemoth ~2T/288B-active) Terms-of-service clause: "you shall also include 'Llama' at the beginning of any such AI model name" — Llama 4 Community License, https://developer.meta.com/ai/llama4/license/ (700M MAU threshold requires separate license) Notable events: 2024-09-25: Llama 3.2 edge models via pruning + distillation | 2025-04-05: Llama 4 codistillation with 'novel distillation loss function that dynamically weights the soft and hard targets'; Behemoth never released Source: https://ai.meta.com/blog/llama-4-multimodal-intelligence/ - DeepSeek (Hangzhou, China) — stance: permissive Distillation products: None; releases distillation datasets implicitly via open R1 outputs used by Open R1 and others Distilled models: DeepSeek-R1-Distill-Qwen-1.5B/7B/14B/32B; DeepSeek-R1-Distill-Llama-8B/70B (SFT on ~800K R1 samples; 32B scores 72.6 AIME'24 vs o1-mini 63.6) Terms-of-service clause: "allow for any modifications and derivative works, including, but not limited to, distillation for training other LLMs" (MIT) — https://huggingface.co/deepseek-ai/DeepSeek-R1 Notable events: 2025-01-22: R1 paper (later Nature 645:633-638) and six distills | 2025-01: Accused by OpenAI; 2025-04 House Select Committee report calls distillation 'highly likely' | 2025-06-03: R1-0528 Gemini-similarity claims | 2025-12-01: V3.2 released | 2026-02: OpenAI memo and Anthropic report (150K+ exchanges); no public response | 2026-04-24 to 2026-08-13: V4 preview, V4-Flash, V4-Pro released Source: https://arxiv.org/abs/2501.12948 - Alibaba Cloud (Qwen) (Hangzhou, China) — stance: permissive Distillation products: None public; Qwen models are the most common open students/teachers in third-party distillation (DeepSeek R1-Distill, SmolLM3, Open R1) Distilled models: Qwen3-0.6B/1.7B/4B/8B/14B and Qwen3-30B-A3B via 'strong-to-weak distillation' from Qwen3-235B-A22B and Qwen3-32B Terms-of-service clause: Apache 2.0 (Qwen3) — https://huggingface.co/Qwen/Qwen3-235B-A22B Notable events: 2025-05-14: Qwen3 report shows distillation beats RL at 1/10 GPU hours (1,800 vs 17,920) | 2026-02: Not named in Anthropic's February report | 2026-06-10: Anthropic's letter to the Senate Banking Committee (Chair Tim Scott, Ranking Member Elizabeth Warren) calls Alibaba's campaign the 'largest known distillation attack' (25,000 accounts, 28.8M interactions, 22 Apr-5 Jun); reported 2026-06-24 | 2026-07-10: Alibaba bans Claude Code for employees, mandates Qoder Source: https://arxiv.org/html/2505.09388v1 - Microsoft (Redmond, US) — stance: mixed Distillation products: Azure OpenAI / Foundry stored completions + Distill workflow (min 10 stored completions; retiring 2026-10-15) Distilled models: Phi-1/2/3 ('largely distill the capabilities of a teacher model (specifically GPT-4)'); Phi-4 14B (synthetic data; surpasses teacher on STEM QA); Phi-4-reasoning 14B (SFT on o3-mini reasoning traces); Phi-4-mini 3.8B (MIT, synthetic 'textbook-like' data); Orca lineage (explanation tuning from GPT-4) Terms-of-service clause: Phi weights: MIT; Azure OpenAI outputs governed by OpenAI-style competing-model restrictions (not independently quoted) Notable events: 2024: Azure OpenAI distillation mirrors OpenAI's stored completions | 2024-12-12: Phi-4 report | 2025-04-30: Phi-4-reasoning distils o3-mini traces | 2025-01: Microsoft security team probes DeepSeek-linked API exfiltration (reported) | 2026-07-06: Docs schedule stored-completions retirement for 2026-10-15 Source: https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/stored-completions - NVIDIA (Santa Clara, US) — stance: permissive Distillation products: NeMo pruning/distillation recipes (Minitron); TensorRT-LLM optimisation of distilled students; NeMo-Aligner SFT Distilled models: Llama-3.1-Minitron 4B width/depth (from Llama 3.1 8B, 94B tokens); Mistral-NeMo-Minitron 8B (from Mistral NeMo 12B); Nemotron-Nano-9B-v2 (from 12B via Minitron); Nemotron 3 Nano 30B-A3B (~3.5T of 10.6T tokens synthesised from DeepSeek-R1, GPT-OSS-120B, Qwen) Terms-of-service clause: NVIDIA Nemotron Open Model License; CC-BY-4.0 for Nemotron Nano 2 — https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Notable events: 2024-08-14: Minitron blog claims up to 40x fewer tokens and 1.8x compute saving | 2025-08-20: Nemotron Nano 2, up to 6x throughput vs Qwen3-8B | 2025-12-15: Nemotron 3 Nano openly lists other labs' models as teachers Source: https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model/ - Amazon (AWS) (Seattle, US) — stance: mixed Distillation products: Amazon Bedrock Model Distillation (preview 2024-12-03, GA 2025-05-01): synthetic data generation from teacher + fine-tuning of student; teachers Nova Premier, Claude 3.5 Sonnet v2, Llama 3.3 70B, Llama 3.1 405B; students Nova Pro/Lite/Micro, Llama 3.2 1B/3B, Llama 3.1 70B/8B Distilled models: Nova Micro / Lite / Pro tiers (training method not disclosed in fetched material) Terms-of-service clause: undisclosed — AWS Service Terms not verified for a distillation clause Notable events: 2024-12-03: Nova family + distillation preview at re:Invent | 2025-05-01: GA with 'up to 500% faster and 75% less expensive ... less than 2% accuracy loss' | 2025: Function-calling distillation for Agents use cases Source: https://aws.amazon.com/about-aws/whats-new/2025/05/amazon-bedrock-model-distillation-generally-available - Mistral AI (Paris, France) — stance: permissive Distillation products: None Distilled models: Ministral 3 3B/8B/14B (base, instruct, reasoning) via cascade distillation from Mistral Small 3.1 24B; 1-3T tokens; Mistral-NeMo-Minitron 8B (by NVIDIA from Mistral NeMo 12B) Terms-of-service clause: "All models are released under the Apache 2.0 license" — https://mistral.ai/news/mistral-3/ Notable events: 2025-12-02: Mistral 3 release (Large 3 675B/41B-active; Ministral 3) | 2026-01-13: Ministral 3 paper describes 'iterative pruning and continued training with distillation' | 2026-02-06: Ministral 3 14B reasoning reported at 85% AIME 2025 vs Qwen 3 14B Thinking 73.7% Source: https://arxiv.org/abs/2601.08584 - Hugging Face (New York, US / Paris, France) — stance: permissive Distillation products: TRL GKDTrainer (on-policy generalized KD; lmbda / beta / seq_kd); TRL Distillation and MiniLLM trainers; Open R1 training scripts and datasets (OpenR1-Math-220k, Mixture-of-Thoughts 350K) Distilled models: OpenR1-Distill-7B (52.7 AIME'24, 89.0 MATH-500); SmolLM3 3B (reasoning SFT data generated by Qwen3-32B; 11.2T tokens) Terms-of-service clause: Apache 2.0 (SmolLM3, TRL, Open R1) — https://huggingface.co/blog/smollm3 Notable events: 2025-01: Open R1 launched days after DeepSeek-R1 | 2025-02: OpenR1-Math-220k released | 2025-05: Mixture-of-Thoughts 350K verified traces | 2025-07-08: SmolLM3 Source: https://github.com/huggingface/open-r1 - Arcee AI (San Francisco, US) — stance: permissive Distillation products: DistillKit (Apache 2.0): logit-based and hidden-state distillation, online/offline modes, logit compression via polynomial approximation + error-diffusion quantization + bit packing Distilled models: 1.5B-Distilled from 7B Arcee-Agent (DistillKit v0.1); AFM-4.5B (Kimi Delta Attention distilled in, per Arcee blog); LLama-405B-Logits dataset published for offline distillation Terms-of-service clause: Apache 2.0 — https://github.com/arcee-ai/DistillKit Notable events: 2024-08-01: DistillKit v0.1 announced | 2025: DistillKit adds offline logit compression; Arcee publishes Llama-405B logits dataset Source: https://arcee.ai/blog/announcing-distillkit/ - xAI (Palo Alto, US) — stance: restrictive Distillation products: None Distilled models: Grok 3 mini / Grok 4 Fast tiers (xAI attributes Grok 4 Fast to 'large-scale reinforcement learning', not distillation; Grok 3 mini deprecated 2026-05-15) Terms-of-service clause: Prohibits "distilling" the Service and using "the Service or Output to develop models or services that compete with xAI" — https://x.ai/legal/terms-of-service and https://x.ai/legal/acceptable-use-policy (effective date undisclosed; pages not directly fetchable) Notable events: 2025-09-19: Grok 4 Fast at $0.20/$0.50 per 1M, '98% reduction in price' vs Grok 4 | 2026-04-30: Musk concedes under oath in Musk v. OpenAI that xAI 'partly' used OpenAI's technology to train its models | 2026-05-15: Grok 3 family aliases redirect to Grok 4.3 Source: https://x.ai/news/grok-4-fast - Moonshot AI (Kimi) (Beijing, China) — stance: mixed Distillation products: None Distilled models: Undisclosed; Kimi K2 (1T/32B-active, 15.5T tokens) model card does not mention distillation; Kimi K3 (alleged by White House to be distilled from Anthropic's Fable; '2.8 trillion parameter') Terms-of-service clause: Modified MIT: "more than 100 million monthly active users, or more than 20 million US dollars ... in monthly revenue, you shall prominently display 'Kimi K2'" — https://huggingface.co/moonshotai/Kimi-K2-Instruct/blob/main/LICENSE Notable events: 2025-07: Kimi K2 open-weight release | 2026-02-23: Anthropic attributes 3.4M+ exchanges targeting computer-use agents and vision | 2026-07-22: OSTP director Kratsios says K3 was built by distilling Anthropic's Fable on GB300 servers; Moonshot did not respond Source: https://cyberscoop.com/white-house-accuses-moonshot-ai-anthropic-model-distillation/ - MiniMax (Shanghai, China) — stance: mixed Distillation products: None Distilled models: Undisclosed; MiniMax-M2 (230B/10B-active) model card gives no training-data provenance Terms-of-service clause: License: modified-mit — https://huggingface.co/MiniMaxAI/MiniMax-M2 Notable events: 2026-02-23: Anthropic attributes 13M+ exchanges (largest of three labs) focused on agentic coding and tool orchestration; redirected nearly half its traffic to a new Claude model within 24 hours | 2026-04-23: Named in White House NSTM-4 coverage Source: https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks - Zhipu AI (Z.ai) (Beijing, China) — stance: permissive Distillation products: None Distilled models: GLM-4.5 355B/32B-active and GLM-4.5-Air 106B/12B-active built with 'expert model iteration and reinforcement learning' (report abstract); explicit self-distillation wording not verified Terms-of-service clause: MIT (GLM-4.5 family) — https://huggingface.co/zai-org/GLM-4.5 Notable events: 2025-08-08: GLM-4.5 report | 2026-02-23: Explicitly not among labs accused by Anthropic Source: https://www.latent.space/p/ainews-anthropic-accuses-deepseek - Apple (Cupertino, US) — stance: mixed Distillation products: None (Foundation Models framework exposes on-device model to developers; no cloud API) Distilled models: On-device ~3B model: dense model trained ~14T tokens, sparse-upcycled into a 64-expert every-2-layer MoE on 1T tokens, then dense model 'retrained ... for the last 10% of tokens (about 1.4T) using a distillation loss from the MoE teacher'; teacher cost cut 90% Terms-of-service clause: N/A — no public API for Apple foundation models; Apple states it does not train on user data Notable events: 2025-07: 2025 tech report documents distillation pipeline and PT-MoE server model (13.4T tokens on 8,192 TPU v5p) Source: https://arxiv.org/abs/2507.13575 - Cohere (Toronto, Canada) — stance: restrictive Distillation products: None Distilled models: None publicly described; Command A (111B) uses 'self-refinement algorithms and model merging techniques'; Command R7B released under CC-BY-NC Terms-of-service clause: "for the purpose of building a similar or competitive product or service" — https://cohere.com/terms-of-use (last updated 2022-09-07) Notable events: 2024-12: Command R7B open weights (CC-BY-NC) | 2025-04-01: Command A technical report Source: https://arxiv.org/abs/2504.00698 TIMELINE OF THIS PERSPECTIVE (40 EVENTS) ---------------------------------------- 2024-05-14 | product | Google says Gemini 1.5 Flash was distilled from 1.5 Pro 2024-07-31 | research | Gemma 2 2B/9B trained with knowledge distillation 2024-08-01 | product | Arcee AI open-sources DistillKit 2024-08-14 | research | NVIDIA publishes Minitron prune-and-distill recipe 2024-09-25 | product | Meta releases Llama 3.2 1B/3B via pruning + logit distillation 2024-10-01 | product | OpenAI launches Model Distillation in the API 2024-12-03 | product | Amazon announces Bedrock Model Distillation (preview) alongside Nova 2024-12-12 | research | Microsoft Phi-4 report: earlier Phi models 'largely distill' GPT-4 2025-01-20 | product | DeepSeek releases R1 plus six R1-Distill models under MIT 2025-01 | legal | OpenAI alleges DeepSeek distilled its models 2025-01 | research | Hugging Face launches Open R1 2025-03-12 | research | Gemma 3: all sizes trained with distillation 2025-04-05 | product | Meta codistills Llama 4 Maverick from Behemoth; license requires 'Llama' naming 2025-04 | policy | House Select Committee report: DeepSeek 'highly likely' used unlawful distillation 2025-04-30 | research | Phi-4-reasoning trained on o3-mini reasoning traces 2025-05-01 | product | Amazon Bedrock Model Distillation reaches GA 2025-05-14 | research | Qwen3 report formalises 'strong-to-weak distillation' 2025-06-03 | market | Researchers claim DeepSeek R1-0528 resembles Gemini 2.5 Pro 2025-06-17 | product | Gemini 2.5 Flash / Flash-Lite ship; report confirms distillation 2025-07-08 | product | Hugging Face SmolLM3 uses Qwen3-32B synthetic reasoning traces 2025-07 | research | Apple reports distilling its on-device 3B model from a 64-expert MoE 2025-08-20 | research | NVIDIA Nemotron Nano 2: 12B pruned and distilled to 9B 2025-09-04 | policy | Anthropic bars Chinese-controlled entities, citing distillation risk 2025-09-19 | product | xAI ships Grok 4 Fast, framed as RL not distillation 2025-10-15 | product | Anthropic releases Claude Haiku 4.5 at $1/$5 2025-12-02 | product | Mistral 3: Ministral 3 built by cascade distillation from Mistral Small 3.1 2025-12-15 | product | NVIDIA Nemotron 3 Nano lists DeepSeek-R1, GPT-OSS-120B and Qwen as synthetic-data teachers 2026-02-12 | policy | OpenAI memo to House Select Committee on DeepSeek's 'obfuscated' distillation 2026-02-23 | legal | Anthropic: 24,000 fraudulent accounts, 16M+ exchanges by DeepSeek, Moonshot, MiniMax 2026-04-07 | market | OpenAI, Anthropic and Google agree to share distillation threat intel via Frontier Model Forum 2026-04-23 | policy | White House memorandum NSTM-4 on 'industrial-scale' distillation 2026-04-24 | product | DeepSeek V4 preview released despite allegations 2026-04-30 | legal | Musk concedes under oath that xAI 'partly' used OpenAI technology 2026-06-10 | legal | Anthropic tells US Senate Alibaba ran 'the largest known distillation attack' 2026-07-04 | market | Anthropic admits Claude Code anti-distillation fingerprinting 'experiment' 2026-07-10 | market | Alibaba bans employees from Claude Code 2026-07-22 | policy | White House accuses Moonshot of distilling Anthropic's Fable into Kimi K3 2026-07-23 | product | Google Cloud documents Gemini 3.1 Pro -> 2.5 Flash distillation service (pre-GA) 2026-09-03 | market | Anthropic: distillation fight moves to the dark web 2026-10-15 | product | Azure OpenAI stored completions (distillation input) scheduled to retire SOURCES CITED BY THIS SECTION (75) ---------------------------------- [1] Model Distillation in the API — OpenAI, 2024-10-01 (blog) https://openai.com/index/api-model-distillation/ [2] OpenAI updates API with model distillation, prompt caching abilities — InfoWorld, 2024-10-03 (news) https://www.infoworld.com/article/3544913/openai-updates-api-with-model-distillation-prompt-caching-abilities.html [3] OpenAI Terms of Use (Rest of World) — OpenAI, 2025 (law) https://openai.com/policies/row-terms-of-use/ [4] Be Careful With OpenAI's Terms of Use — OSPOCO, 2024 (blog) https://ospo.co/blog/be-careful-with-openais-terms-of-use/ [5] OpenAI API pricing — OpenAI, 2026-09 (pricing) https://developers.openai.com/api/docs/pricing [6] Supervised fine-tuning: distilling from a larger model — OpenAI, 2026 (docs) https://developers.openai.com/api/docs/guides/supervised-fine-tuning [7] Memo to US House Select Committee: Updated Stakes for American-Led, Democratic AI — OpenAI (via Bloomberg), 2026-02-12 (filing) https://assets.bwbx.io/documents/users/iqjWHBFdfxIU/rRmql_jJcxb4/v0 [8] OpenAI Accuses China's DeepSeek of Distilling US AI Models to Gain an Edge — Bloomberg, 2026-02-12 (news) https://www.bloomberg.com/news/articles/2026-02-12/openai-accuses-deepseek-of-distilling-us-models-to-gain-an-edge [9] OpenAI accuses DeepSeek of malpractice ahead of AI launch — Rest of World, 2026-02-12 (news) https://restofworld.org/2026/openai-deepseek-distillation-dispute-us-china/ [10] The Innovation Dilemma: AI Distillation in OpenAI v. DeepSeek — Berkeley Law, 2025-03-30 (blog) https://sites.law.berkeley.edu/thenetwork/2025/03/30/the-innovation-dilemma-ai-distillation-in-openai-v-deepseek/ [11] OpenAI Alleges China's DeepSeek Stole its IP to Train its Own Models — FDD, 2026-02-13 (news) https://www.fdd.org/analysis/2026/02/13/openai-alleges-chinas-deepseek-stole-its-intellectual-property-to-train-its-own-models/ [12] Detecting and preventing distillation attacks — Anthropic, 2026-02-23 (blog) https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks [13] Anthropic accuses Chinese labs of trying to illicitly take Claude's capabilities — CyberScoop, 2026-02-23 (news) https://cyberscoop.com/anthropic-accuses-chinese-labs-ai-distillation-cyber-risk/ [14] Updating restrictions of sales to unsupported regions — Anthropic, 2025-09-04 (blog) https://www.anthropic.com/news/updating-restrictions-of-sales-to-unsupported-regions [15] Anthropic tightens AI access rules, targeting Chinese-controlled entities — CRN Asia, 2025-09-05 (news) https://www.crnasia.com/news/2025/artificial-intelligence/anthropic-tightens-ai-access-rules [16] Anthropic Commercial Terms of Service — Anthropic, 2025-06-17 (law) https://www.anthropic.com/legal/commercial-terms [17] Anthropic Consumer Terms of Service — Anthropic, 2025-10-08 (law) https://www.anthropic.com/legal/consumer-terms [18] Introducing Claude Haiku 4.5 — Anthropic, 2025-10-15 (blog) https://www.anthropic.com/news/claude-haiku-4-5 [19] Anthropic accuses Alibaba of campaign to 'brazenly' and 'illicitly' distill Claude — CNBC, 2026-06-24 (news) https://www.cnbc.com/2026/06/24/anthropic-alibaba-distillation-campaign.html [20] Alibaba reportedly bans employees from using Claude Code — TechCrunch, 2026-07-04 (news) https://techcrunch.com/2026/07/04/alibaba-reportedly-bans-employees-from-using-claude-code/ [21] White House official accuses Chinese startup of distilling Anthropic's model — CyberScoop, 2026-07-22 (news) https://cyberscoop.com/white-house-accuses-moonshot-ai-anthropic-model-distillation/ [22] Anthropic's distillation battle turns to the dark web as China concerns swell — CNBC, 2026-09-03 (news) https://www.cnbc.com/2026/09/03/anthropic-distillation-battle-turns-to-dark-web-china-concerns-swell.html [23] White House accuses China of 'deliberate, industrial-scale' campaigns to steal US AI models — Nextgov, 2026-04-23 (news) https://www.nextgov.com/artificial-intelligence/2026/04/white-house-accuses-china-deliberate-industrial-scale-campaigns-steal-us-ai-models/413083/ [24] White House accuses China of industrial-scale theft of US AI frontier models — Interesting Engineering, 2026-04-23 (news) https://interestingengineering.com/ai-robotics/ai-war-white-house-accuses-china-of-industrial-scale-theft-of-us-ai-frontier-models [25] OpenAI, Anthropic, Google join forces against China — Tech Brew (citing Bloomberg), 2026-04-07 (news) https://www.techbrew.com/stories/openai-anthropic-google-distillation-collab [26] Gemini 2.5: Pushing the Frontier with Advanced Reasoning — Google DeepMind, 2025-07 (paper) https://arxiv.org/html/2507.06261v1/ [27] Gemini 1.5 Flash announcement (Google I/O 2024) — Google, 2024-05-14 (blog) https://blog.google/technology/ai/google-gemini-update-flash-ai-assistant-io-2024/ [28] Gemma 2: Improving Open Language Models at a Practical Size — Google DeepMind, 2024-07-31 (paper) https://arxiv.org/abs/2408.00118 [29] Gemma 3 Technical Report — Google DeepMind, 2025-03-12 (paper) https://arxiv.org/html/2503.19786v1 [30] Gemini API Additional Terms of Service — Google, 2026-04-28 (law) https://ai.google.dev/gemini-api/terms [31] Google Cloud page describes Gemini distillation service, but its release status is unclear — RuntimeWire, 2026-07-28 (news) https://runtimewire.com/article/google-cloud-page-describes-gemini-distillation-service-but-its-release-status-i [32] Google says Gemini was hit with 100,000 prompts in apparent cloning attempt — NBC News, 2026-02 (news) https://www.nbcnews.com/tech/security/google-gemini-hit-100000-prompts-cloning-attempt-rcna258657 [33] Is DeepSeek Training its AI with Data from Google Gemini? — WinBuzzer, 2025-06-03 (news) https://winbuzzer.com/2025/06/03/is-deepseek-training-its-ai-with-data-from-google-gemini-new-distillation-claims-emerge-xcxwbn/ [34] Llama 3.2: Revolutionizing edge AI and vision with open, customizable models — Meta, 2024-09-25 (blog) https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ [35] The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation — Meta, 2025-04-05 (blog) https://ai.meta.com/blog/llama-4-multimodal-intelligence/ [36] Llama 4 Community License Agreement — Meta, 2025-04-05 (law) https://developer.meta.com/ai/llama4/license/ [37] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek, 2025-01-22 (paper) https://arxiv.org/abs/2501.12948 [38] DeepSeek-R1 model card and license — DeepSeek / Hugging Face, 2025-01 (docs) https://huggingface.co/deepseek-ai/DeepSeek-R1 [39] DeepSeek (Wikipedia) - V4 release dates — Wikipedia, 2026 (docs) https://en.wikipedia.org/wiki/DeepSeek [40] Qwen3 Technical Report — Alibaba Qwen Team, 2025-05-14 (paper) https://arxiv.org/html/2505.09388v1 [41] Qwen3-235B-A22B model card — Alibaba / Hugging Face, 2025-05 (docs) https://huggingface.co/Qwen/Qwen3-235B-A22B [42] Phi-4 Technical Report — Microsoft, 2024-12-12 (paper) https://arxiv.org/abs/2412.08905 [43] Phi-4-reasoning Technical Report — Microsoft, 2025-04-30 (paper) https://arxiv.org/abs/2504.21318 [44] Phi-4-mini-instruct model card — Microsoft / Hugging Face, 2025-02 (docs) https://huggingface.co/microsoft/Phi-4-mini-instruct [45] How to use Azure OpenAI stored completions & distillation — Microsoft Learn, 2026-07-06 (docs) https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/stored-completions [46] Introducing Model Distillation in Azure OpenAI Service — Microsoft, 2024 (blog) https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/introducing-model-distillation-in-azure-openai-service/4298627 [47] How to Prune and Distill Llama-3.1 8B to an NVIDIA Llama-3.1-Minitron 4B Model — NVIDIA, 2024-08-14 (blog) https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model/ [48] LLM Pruning and Distillation in Practice: The Minitron Approach — NVIDIA, 2024-08 (paper) https://arxiv.org/abs/2408.11796 [49] NVIDIA Nemotron Nano 2 — NVIDIA, 2025-08-20 (paper) https://arxiv.org/abs/2508.14444 [50] NVIDIA-Nemotron-3-Nano-30B-A3B model card — NVIDIA / Hugging Face, 2025-12-15 (docs) https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 [51] Amazon Bedrock Model Distillation (preview) — AWS, 2024-12-03 (blog) https://aws.amazon.com/blogs/aws/build-faster-more-cost-efficient-highly-accurate-models-with-amazon-bedrock-model-distillation-preview/ [52] Amazon Bedrock Model Distillation is now generally available — AWS, 2025-05-01 (docs) https://aws.amazon.com/about-aws/whats-new/2025/05/amazon-bedrock-model-distillation-generally-available [53] The Amazon Nova Family of Models: Technical Report and Model Card — Amazon, 2024-12-03 (paper) https://www.amazon.science/publications/the-amazon-nova-family-of-models-technical-report-and-model-card [54] Introducing Mistral 3 — Mistral AI, 2025-12-02 (blog) https://mistral.ai/news/mistral-3/ [55] Ministral 3 (paper) — Mistral AI, 2026-01-13 (paper) https://arxiv.org/abs/2601.08584 [56] Mistral Uses Cascade Distillation on Mistral 3 To Build Ministral Family — DeepLearning.AI The Batch, 2026-02-06 (news) https://www.deeplearning.ai/the-batch/mistral-uses-cascade-distillation-on-mistral-3-to-build-ministral-family/ [57] Generalized Knowledge Distillation Trainer — Hugging Face, 2026 (docs) https://huggingface.co/docs/trl/gkd_trainer [58] Open R1: fully open reproduction of DeepSeek-R1 — Hugging Face, 2025-01 (docs) https://github.com/huggingface/open-r1 [59] SmolLM3: smol, multilingual, long-context reasoner — Hugging Face, 2025-07-08 (blog) https://huggingface.co/blog/smollm3 [60] Announcing DistillKit — Arcee AI, 2024-08-01 (blog) https://arcee.ai/blog/announcing-distillkit/ [61] arcee-ai/DistillKit — Arcee AI, 2026 (docs) https://github.com/arcee-ai/DistillKit [62] Grok 4 Fast — xAI, 2025-09-19 (blog) https://x.ai/news/grok-4-fast [63] xAI Terms of Service - Consumer — xAI, undisclosed (law) https://x.ai/legal/terms-of-service [64] xAI Acceptable Use Policy — xAI, undisclosed (law) https://x.ai/legal/acceptable-use-policy [65] Musk admits distilling OpenAI data for his xAI — Forbes Africa, 2026-05-01 (news) https://www.forbesafrica.com/current-affairs/2026/05/01/musk-admits-distilling-openai-data-for-his-xai-heres-why-thats-controversial [66] Kimi-K2-Instruct model card — Moonshot AI / Hugging Face, 2025-07 (docs) https://huggingface.co/moonshotai/Kimi-K2-Instruct [67] Kimi K2 Modified MIT License — Moonshot AI, 2025-07 (law) https://huggingface.co/moonshotai/Kimi-K2-Instruct/blob/main/LICENSE [68] MiniMax-M2 model card — MiniMax / Hugging Face, 2025 (docs) https://huggingface.co/MiniMaxAI/MiniMax-M2 [69] GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models — Zhipu AI / Z.ai, 2025-08-08 (paper) https://arxiv.org/abs/2508.06471 [70] GLM-4.5 model card and license — Zhipu AI / Hugging Face, 2025-07 (docs) https://huggingface.co/zai-org/GLM-4.5 [71] AINews: Anthropic accuses DeepSeek, Moonshot, and MiniMax — Latent Space, 2026-02-23 (news) https://www.latent.space/p/ainews-anthropic-accuses-deepseek [72] Apple Intelligence Foundation Language Models: Tech Report 2025 — Apple, 2025-07 (paper) https://arxiv.org/abs/2507.13575 [73] Command A: An Enterprise-Ready Large Language Model — Cohere, 2025-04-01 (paper) https://arxiv.org/abs/2504.00698 [74] Cohere Terms of Use — Cohere, 2022-09-07 (law) https://cohere.com/terms-of-use [75] C4AI Command R7B model card — Cohere Labs / Hugging Face, 2024-12 (docs) https://huggingface.co/CohereLabs/c4ai-command-r7b-12-2024 ================================================================================================ 7. DEVELOPER PERSPECTIVE — DEVELOPER VIEW OF AI DISTILLATION: TOOLING, PLATFORMS AND RECIPES (2026) ================================================================================================ Page: https://global-distillation.com/developer Data: https://global-distillation.com/data/developer.json Updated: 2026-09-03 Content: 8 key figures, 7 tables, 8 charts, 20 dated events, 14 glossary terms, 58 sources SUMMARY -------- By September 2026 a developer can distill a model three ways: with open libraries (Hugging Face TRL now ships four distillation trainers, plus Arcee DistillKit, torchtune, NVIDIA Model Optimizer/NeMo and an Axolotl KD plugin), with managed cloud pipelines (Amazon Bedrock Model Distillation, Azure Foundry stored completions, and OpenAI's now-sunsetting fine-tuning platform), or with per-token fine-tuning APIs (Together, Fireworks, Databricks) fed by teacher-generated data. The field has converged on on-policy distillation where the student generates and the teacher grades every token; Qwen3 reports this needs roughly 1/10 of the GPU hours of RL for a better 8B model. Open reasoning datasets (OpenThoughts3-1.2M, OpenR1-Math-220k, Bespoke-Stratos-17k) plus cheap QLoRA mean a 7B reasoning student can be trained for well under $200 of GPU time, and DeepSeek's R1-Distill family alone has passed 97 million Hugging Face downloads. Managed options are in flux: OpenAI stops accepting new fine-tuning jobs on 2027-01-06, Azure retires stored completions on 2026-10-15, Bedrock currently lists no Anthropic teacher, and Vertex AI documents distillation only for open-model tuning. Every number below carries a source URL; where a vendor no longer publishes a figure it is marked undisclosed. KEY FIGURES ----------- - DeepSeek-R1-Distill family, all-time HF downloads: 97,824,225 downloads (2.23M in the last 30 days) Sum of the six official R1-Distill repos (1.5B, 7B, 8B, 14B, 32B, 70B) via the Hugging Face API on 2026-09-03 Source: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B - vLLM GitHub stars: 90,892 stars (queried 2026-09-03) Default serving engine for distilled students and EAGLE-3 draft heads Source: https://github.com/vllm-project/vllm - Unsloth GitHub stars: 75,556 stars (vs 74,553 for LLaMA-Factory) Most-starred fine-tuning library; QLoRA 7B student needs ~5 GB VRAM Source: https://github.com/unslothai/unsloth - Qwen3-8B: distillation vs RL GPU-hours: 1,800 GPU-hours (vs 17,920 for RL-only (≈1/10)) Table 21 of the Qwen3 technical report; distilled model also scored higher on AIME'24/'25 Source: https://arxiv.org/abs/2505.09388 - Cheapest managed LoRA SFT (≤16B student): 0.48 USD per 1M tokens (Together AI; Fireworks $0.50) Full-parameter SFT is $1.00-$1.20/M tokens on the same tiers Source: https://www.together.ai/pricing - s1-32B reasoning distillation compute: 7 H100 GPU-hours (26 min on 16 H100s, 1,000 samples) Traces from Gemini Flash Thinking; beat o1-preview on AIME24 (56.7 vs 44.6) Source: https://arxiv.org/abs/2501.19393 - OpenThoughts3-1.2M rows: 1,200,000 rows (850k math / 250k code / 100k science) Largest open reasoning-distillation set; teacher QwQ-32B; Apache-2.0 Source: https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M - Days left to start a new OpenAI fine-tuning job: 125 days (new orgs blocked since 2026-05-07) Last day for new jobs is 2027-01-06, counted from 2026-09-03; inference on existing fine-tunes continues until base models are deprecated Source: https://developers.openai.com/api/docs/deprecations KEY FINDINGS ------------ 1. On-policy distillation is now the default recipe, and TRL ships it out of the box TRL v1.12 offers four distillation trainers: GKDTrainer (generalized JSD, lmbda/beta/seq_kd), DistillationTrainer (on-policy, chunked JSD, vLLM-accelerated generation, tool-use and VLM support), GOLDTrainer (cross-tokenizer via Universal Logit Distillation) and MiniLLMTrainer (reverse-KL policy gradient). Hugging Face's July 2026 survey finds Qwen3, DeepSeek-V4, GLM-5, Nemotron 3 Ultra and MiMo-V2-Flash all use some form of on-policy distillation where the teacher grades the student's own rollouts. Sources: https://huggingface.co/docs/trl/distillation_trainer https://huggingface.co/docs/trl/gold_trainer https://huggingface.co/blog/sergiopaniego/distillation-2026 2. Distillation is roughly 10x cheaper than RL for small reasoning models Qwen3's technical report (Table 21) reports Qwen3-8B reaching AIME'24 74.4 / AIME'25 65.5 with 1,800 GPU-hours of on-policy distillation from Qwen3-32B and Qwen3-235B-A22B teachers on top of an off-policy-distilled checkpoint (the same starting point as the RL run; that checkpoint’s cost is excluded from both figures), versus 67.6 / 55.5 with 17,920 GPU-hours of RL from that checkpoint. Independent open runs corroborate the cheapness: s1-32B used 7 H100-hours, Sky-T1-32B about $450 (152 H100-hours). Sources: https://arxiv.org/abs/2505.09388 https://arxiv.org/abs/2501.19393 https://novasky-ai.github.io/posts/sky-t1/ 3. Managed distillation on the big clouds is shrinking, not growing OpenAI is winding down its fine-tuning platform in phases (2026-05-07, 2026-07-02, 2027-01-06); Azure retires stored completions on 2026-10-15; Amazon Bedrock's supported-model table states distillation is not currently available for Anthropic models with no restoration timeline; Google's Vertex AI still documents teacher-to-student distillation fine-tuning, but only for open models (Llama 3.1, Qwen) via the GenAI SDK, while Gemini tuning is limited to supervised, RL and preference tuning. The action has moved to per-token fine-tuning APIs (Together, Fireworks, Databricks) and to open libraries. Sources: https://developers.openai.com/api/docs/deprecations https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/stored-completions https://docs.aws.amazon.com/bedrock/latest/userguide/prequisites-model-distillation.html https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning/open-model-tuning 4. A 7B reasoning student fits on one consumer GPU with QLoRA Unsloth's published minimums are 5 GB VRAM for a 7B QLoRA run and 41 GB for 70B; a 32B student needs 26 GB (QLoRA) or 76 GB (LoRA 16-bit). Long reasoning traces (8k-16k tokens) raise activation memory well above these floors, so an 80 GB A100/H100 at $2.50-$3.95/h (Modal) or $1.99/h preemptible (Together) is the practical single-GPU tier for R1-style distillation. Sources: https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements https://modal.com/pricing https://www.together.ai/pricing 5. Open reasoning datasets have made the teacher's API bill optional Bespoke-Stratos-17k cost about $800 of DeepSeek-R1 calls to generate; OpenR1-Math-220k, OpenThoughts3-1.2M (QwQ-32B traces), Mixture-of-Thoughts and NVIDIA's OpenMathReasoning are all Apache-2.0 or similar and together see hundreds of thousands of downloads a month. OpenThinker3-7B, trained on OpenThoughts3 from Qwen2.5-7B-Instruct, reports AIME25 53.3 and LiveCodeBench 51.7, above DeepSeek-R1-Distill-Qwen-32B on the same dataset card. Sources: https://huggingface.co/datasets/bespokelabs/Bespoke-Stratos-17k https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M https://huggingface.co/datasets/open-r1/OpenR1-Math-220k 6. Logit access is the dividing line between recipes White-box KD (torchtune forward-KL, DistillKit logit/hidden-state losses, NeMo Model Optimizer logit and intermediate-layer losses, Axolotl's KD plugin consuming vLLM top-k logprobs) needs the teacher's weights or logprobs. Black-box distillation from an API (OpenAI stored completions, Bedrock synthetic data, Bespoke Curator or distilabel pipelines) only needs text, which is why it dominates for closed teachers. GOLD/ULD removes the same-tokenizer restriction that torchtune and Axolotl still carry. Sources: https://pytorch.org/blog/llama-into-torchtune/ https://github.com/arcee-ai/DistillKit https://github.com/axolotl-ai-cloud/axolotl/tree/main/src/axolotl/integrations/kd https://huggingface.co/docs/trl/gold_trainer 7. Speculative-decoding draft heads are a second, cheaper kind of distillation EAGLE-3 trains a one-layer draft head on the target model's hidden features; Red Hat/vLLM report an up to 2.5x headline speedup, with measured latency gains of 1.6x-2.1x, and the May 2026 EAGLE 3.1 release reports 2.03x per-user throughput on Kimi-K2.6 at concurrency 1. SGLang's SpecForge (1.1k stars) and TorchSpec train these heads, and they are served with a single --speculative-config flag, so a distilled student can itself be paired with an even smaller draft. Sources: https://vllm.ai/blog/2026-05-26-eagle-3-1 https://developers.redhat.com/articles/2025/07/01/fly-eagle3-fly-faster-inference-vllm-speculative-decoding https://github.com/sgl-project/SpecForge 8. Pruning plus distillation is the enterprise path to a model family NVIDIA's Minitron recipe (teacher correction on 94B tokens, depth or width pruning of Llama-3.1-8B, then distillation on 94B (width) or 1.4T (depth) tokens) is reported alongside NVIDIA’s earlier Nemotron prune-and-distill work, from which the up to 40x fewer training tokens per additional model and 1.8x compute savings for a full family are taken, with the depth-pruned 4B running about 2.7x the throughput of the 8B on H100 under TensorRT-LLM. The same pipeline is exposed through NeMo Framework and Model Optimizer (3.7k stars). Sources: https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model https://docs.nvidia.com/nemo-framework/user-guide/latest/model-optimization/distillation/distillation.html TABLES -------- TABLE: Managed distillation and fine-tuning platforms (September 2026) [id: platform-matrix, 10 rows] What each hosted platform actually offers a developer who wants to distill, with published prices where they exist. Platform | Distillation product | Teacher options | Student options | Published training price | Status | Source URL -------- | -------------------- | --------------- | --------------- | ------------------------ | ------ | ---------- OpenAI | Stored completions -> Evals -> fine-tune (dashboard 'Distill') | Any OpenAI model called with store=true | gpt-4.1, gpt-4.1-mini, gpt-4.1-nano (SFT/DPO); o4-mini (RFT) | $25 / $5 / $1.50 per 1M training tokens (4.1 / mini / nano); o4-mini RFT $100/hour | Winding down: no new orgs since 2026-05-07; no new jobs after 2027-01-06 | https://developers.openai.com/api/docs/pricing Amazon Bedrock | Model Distillation (synthetic data + student fine-tune, one job) | Nova Pro, Nova Premier, Llama 3.1 405B, Llama 3.1 70B, Llama 3.3 70B | Nova Micro/Lite/Pro; Llama 3.1 8B/70B, 3.2 1B/3B, 3.3 70B | Teacher calls at on-demand rate (synthesis up to 15k pairs) + student at customization rate (per-model rates not on public page excerpt) | GA; Anthropic teachers 'not currently available', no timeline | https://docs.aws.amazon.com/bedrock/latest/userguide/prequisites-model-distillation.html Azure / Microsoft Foundry | Stored completions -> Distill -> Azure OpenAI fine-tune (classic portal) | Any Azure OpenAI model with store=true | Azure OpenAI fine-tunable models | Azure OpenAI fine-tuning rates (not on this page) | Stored completions retire 2026-10-15; min 10 completions | https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/stored-completions Google Vertex AI | Distillation fine-tuning for open models (teacher -> student via GenAI SDK); supervised, RL and preference tuning for Gemini | Any supported open/Gemini model used as teacher for open-model distillation | Llama 3.1, Qwen open models (distillation); Gemini 2.5/3.5 Flash-Lite/Pro (SFT/RL/preference) | undisclosed on tuning overview page | GA; legacy text-model distillation page returns 404 | https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning/open-model-tuning Together AI | Per-token SFT/DPO API (bring teacher-generated data) | Any (you generate data) | Open models up to 100B+ (DeepSeek-V4 Flash, GLM-5, Qwen 3.5 priced separately) | LoRA SFT $0.48/M (<=16B), $1.50 (17-69B), $2.90 (70-100B); full SFT $1.20/$3.75/$7.25; $4 minimum | GA | https://www.together.ai/pricing Fireworks AI | Per-token SFT/DPO, RFT per GPU-hour | Any (you generate data) | Open models up to >300B | LoRA SFT $0.50/M (<=16B), $3 (16-80B), $6 (80-300B), $10 (>300B); full 2x; RFT = GPU rate ($8/h H100 from 2026-09-01) | GA; dataset 3 to 3M examples (per fine-tuning docs) | https://fireworks.ai/pricing Predibase | SFT/Turbo LoRA/RFT fine-tuning | Any (you generate data) | Open models | undisclosed (predibase.com and docs now redirect to Rubrik) | Folded into Rubrik Agent Cloud | https://predibase.com/pricing Databricks Mosaic AI | Foundation Model Fine-tuning (DBU-priced) | Any (you generate data) | Llama 3.x family and others | $0.65/DBU; e.g. Llama 3.1 8B ~100 DBU (~$65) per 10M words, Llama 3.3 70B ~225 DBU (~$146) | GA | https://www.databricks.com/product/pricing/mosaic-foundation-model-training Modal | Serverless GPUs (run TRL/Unsloth yourself) | Any | Any | H100 $3.95/h, A100-80GB $2.50/h, L40S $1.95/h, per-second billing, $30/month free | GA | https://modal.com/pricing Anyscale | Ray-based post-training (LLaMA-Factory, SkyRL, Ray Train) | Any | Any | undisclosed on docs page | GA | https://docs.anyscale.com/llm/fine-tuning Notes: Bedrock student-training rates for Nova and Llama 3.x are set at 'model customization' rates but the public pricing page excerpt only shows Llama 2 ($1.49/M tokens, 13B) and gpt-oss-20b ($80/training hour). Together and Fireworks size buckets differ slightly (17-69B vs 16.1-80B). Sources: https://developers.openai.com/api/docs/deprecations https://docs.aws.amazon.com/bedrock/latest/userguide/model-distillation.html https://aws.amazon.com/bedrock/pricing/ https://www.together.ai/pricing https://fireworks.ai/pricing https://docs.fireworks.ai/fine-tuning/fine-tuning-models https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning/open-model-tuning https://modal.com/pricing TABLE: Open-source distillation library feature matrix [id: library-feature-matrix, 11 rows] Which library gives you which knob. Stars and licenses from the GitHub API on 2026-09-03. Library | GitHub stars (stars) | License | Logit / white-box KD | On-policy KD | Cross-tokenizer | Pruning | LoRA/QLoRA | Notes | Source URL ------- | -------------------- | ------- | -------------------- | ------------ | --------------- | ------- | ---------- | ----- | ---------- Hugging Face TRL | 19210 | Apache-2.0 | Yes (GKD, DistillationTrainer, MiniLLM) | Yes (vLLM colocate/server) | Yes (GOLD/ULD, experimental) | No | Yes | Tool-calling and VLM distillation supported | https://huggingface.co/docs/trl/distillation_trainer Arcee DistillKit | 1052 | Apache-2.0 | Yes (KL, JSD, TVD, ranking, hidden-state MSE/cosine) | Online teacher inference | Yes (via mergekit-tokensurgeon embedding surgery) | No | Yes | Offline logit capture compressed to ~300 bytes/token | https://github.com/arcee-ai/DistillKit torchtune | 5802 | BSD-3-Clause | Yes (forward KL + CE, kd_ratio) | No | No | No | Yes (LoRA recipes) | knowledge_distillation_single_device / _distributed recipes | https://pytorch.org/blog/llama-into-torchtune/ NVIDIA NeMo + Model Optimizer | 3724 | Apache-2.0 | Yes (logit_layers, intermediate_layer_pairs cosine) | No | No | Yes (depth/width, Minitron) | Via NeMo | NeMo 2.0 GPT checkpoints only; stars are for NVIDIA/Model-Optimizer | https://docs.nvidia.com/nemo-framework/user-guide/latest/model-optimization/distillation/distillation.html Axolotl | 12436 | Apache-2.0 | Yes (KD plugin; top-k teacher logprobs in dataset) | No | No | No | Yes | kd_ce_alpha / kd_alpha / kd_temperature; not in main docs index | https://github.com/axolotl-ai-cloud/axolotl/tree/main/src/axolotl/integrations/kd Unsloth | 75556 | Apache-2.0 | No native KD loss | No | n/a | No | Yes (QLoRA focus) | Fastest path for SFT on teacher-generated text; 7B QLoRA in 5 GB | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements LLaMA-Factory | 74553 | Apache-2.0 | No native KD loss | No | n/a | No | Yes (2-8 bit QLoRA) | 100+ models, LLaMA Board UI; used to train Sky-T1 | https://github.com/hiyouga/LlamaFactory SGLang SpecForge | 1147 | MIT | Draft-head on target hidden states | Online/offline modes | n/a | No | n/a | EAGLE3, P-EAGLE, DFlash draft training for speculative decoding | https://github.com/sgl-project/SpecForge Bespoke Curator | 1722 | Apache-2.0 | No (black-box data synthesis) | No | n/a | No | n/a | Built Bespoke-Stratos-17k from DeepSeek-R1 in ~1.5 h | https://github.com/bespokelabsai/curator Argilla distilabel | 3385 | Apache-2.0 | No (black-box data synthesis) | No | n/a | No | n/a | Pipelines for synthetic instruction/preference data | https://github.com/argilla-io/distilabel Hugging Face Open-R1 | 26447 | Apache-2.0 | No (SFT on R1 traces + GRPO) | No | n/a | No | Via TRL | OpenR1-Distill-7B: AIME24 52.7, MATH-500 89.0 | https://github.com/huggingface/open-r1 Notes: 'No native KD loss' means the library trains on teacher text (sequence-level KD) but has no logit-matching objective. NVIDIA-NeMo/NeMo redirects to a Speech repo on GitHub (18,382 stars); the LLM distillation code now lives in Model Optimizer. Sources: https://huggingface.co/docs/trl/gkd_trainer https://github.com/arcee-ai/DistillKit https://github.com/meta-pytorch/torchtune https://github.com/NVIDIA/Model-Optimizer https://github.com/axolotl-ai-cloud/axolotl https://github.com/unslothai/unsloth https://github.com/hiyouga/LlamaFactory TABLE: Open datasets for distillation [id: dataset-comparison, 12 rows] Teacher-generated corpora a developer can train a student on today. Download counts from the Hugging Face API, 2026-09-03. Dataset | Rows | Teacher | Domain | License | Downloads (30d) (downloads) | Downloads (all-time) (downloads) | Source URL ------- | ---- | ------- | ------ | ------- | ------------------------ | ------------------------ | ---------- open-thoughts/OpenThoughts3-1.2M | 1,200,000 | QwQ-32B | Math 850k / code 250k / science 100k | Apache-2.0 | 19174 | 247476 | https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M open-thoughts/OpenThoughts-114k | 114k | DeepSeek-R1 | Math, code, science, puzzles | Apache-2.0 | 109485 | 1477621 | https://huggingface.co/datasets/open-thoughts/OpenThoughts-114k bespokelabs/Bespoke-Stratos-17k | 16,710 | DeepSeek-R1 (via Curator) | 5k code, 10k math, 1k science/puzzle | Apache-2.0 | 19434 | 321082 | https://huggingface.co/datasets/bespokelabs/Bespoke-Stratos-17k open-r1/OpenR1-Math-220k | 220k problems (94k 'default' split), 2-4 traces each | DeepSeek-R1 | Math (NuminaMath-1.5 problems) | Apache-2.0 | 105083 | 572723 | https://huggingface.co/datasets/open-r1/OpenR1-Math-220k open-r1/Mixture-of-Thoughts | 350k verified traces | DeepSeek-R1 | Math, code, science | not declared on card (science split derives from CC-BY-4.0 Llama-Nemotron) | 10735 | 131591 | https://huggingface.co/datasets/open-r1/Mixture-of-Thoughts AI-MO/NuminaMath-1.5 | ~900k problems | Human/CoT (base for R1 traces) | Competition math | see card | 33515 | 106494 | https://huggingface.co/datasets/AI-MO/NuminaMath-1.5 nvidia/OpenMathReasoning | see card | DeepSeek-R1 / QwQ | Math (CoT + tool-integrated) | CC-BY-4.0 | 69484 | 335233 | https://huggingface.co/datasets/nvidia/OpenMathReasoning nvidia/Llama-Nemotron-Post-Training-Dataset | see card | DeepSeek-R1, Qwen, Llama | Reasoning + chat + safety | CC-BY-4.0 | 7531 | 106953 | https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset simplescaling/s1K-1.1 | 1,000 | DeepSeek-R1 (1.1); Gemini Flash Thinking (1.0) | Hard math/science questions | MIT | 7133 | 78174 | https://huggingface.co/datasets/simplescaling/s1K-1.1 a-m-team/AM-DeepSeek-R1-Distilled-1.4M | 1.4M | DeepSeek-R1 | General reasoning | CC-BY-NC-4.0 (non-commercial) | 1793 | 51544 | https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-Distilled-1.4M Magpie-Align/Magpie-Pro-300K-Filtered | 300k | Llama-3-70B-Instruct (self-synthesized prompts) | General instruction/chat | Llama 3 | 3349 | 27342 | https://huggingface.co/datasets/Magpie-Align/Magpie-Pro-300K-Filtered BAAI/Infinity-Instruct | 7M+ foundational + chat subsets | Compiled + synthesized | General instruction | CC-BY-SA-4.0 | 3929 | 119827 | https://huggingface.co/datasets/BAAI/Infinity-Instruct Notes: Row counts and licenses for NuminaMath-1.5, OpenMathReasoning, Nemotron and AM-1.4M are taken from dataset names/cards as listed; verify the card before commercial use. R1-derived traces come from a model released under DeepSeek's MIT licence, but each dataset card sets its own terms: AM-DeepSeek-R1-Distilled-1.4M is CC-BY-NC-4.0 (non-commercial) and Mixture-of-Thoughts declares no licence at all. Sources: https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M https://huggingface.co/datasets/bespokelabs/Bespoke-Stratos-17k https://huggingface.co/datasets/open-r1/OpenR1-Math-220k https://arxiv.org/abs/2406.08464 https://arxiv.org/html/2506.11116 TABLE: Minimum GPU memory by student size (SFT on teacher outputs) [id: gpu-requirements, 13 rows] Unsloth's published 'absolute minimum' VRAM for QLoRA (4-bit) and LoRA (16-bit) at default sequence length, mapped to the cheapest serverless GPU that fits. Student params | QLoRA 4-bit VRAM (GB) | LoRA 16-bit VRAM (GB) | Cheapest fitting GPU (Modal) | Price (USD/hour) | Source URL -------------- | --------------------- | --------------------- | ------------------------ | ---------------- | ---------- 3B | 3.5 | 8 | T4 16GB | 0.59 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 7B | 5 | 19 | T4 (QLoRA) / L4 24GB (LoRA) | 0.8 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 8B | 6 | 22 | L4 24GB | 0.8 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 9B | 6.5 | 24 | L4 24GB (QLoRA) / L40S 48GB (LoRA) | 1.95 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 11B | 7.5 | 29 | L40S 48GB | 1.95 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 14B | 8.5 | 33 | L40S 48GB | 1.95 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 27B | 22 | 64 | L40S (QLoRA) / A100-80GB (LoRA) | 2.5 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 32B | 26 | 76 | L40S (QLoRA) / A100-80GB (LoRA) | 2.5 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 40B | 30 | 96 | L40S (QLoRA) / 2x A100-80GB (LoRA) | 5 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 70B | 41 | 164 | L40S/A100-80GB (QLoRA) / 3x H100 (LoRA) | 11.85 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 81B | 48 | 192 | A100-80GB (QLoRA) / 3x H100 (LoRA) | 11.85 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 90B | 53 | 212 | A100-80GB (QLoRA) / 3x H100 (LoRA) | 11.85 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements 405B | 237 | 950 | 4x H100 (QLoRA) / 12x H100 (LoRA) | 47.4 | https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements Notes: GPU mapping and hourly price use Modal's per-second rates (T4 $0.59, L4 $0.80, L40S $1.95, A100-80GB $2.50, H100 $3.95) and are the author's mapping of Unsloth's VRAM floors, priced for the LoRA column where it needs the larger GPU. Long reasoning traces (8k-16k tokens) and on-policy generation raise memory well above these floors; on-policy trainers also need memory for the teacher and a vLLM engine. Sources: https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements https://modal.com/pricing TABLE: Published distillation runs with disclosed compute or cost [id: published-distillation-runs, 10 rows] What it actually took to produce well-known distilled students, as reported by their authors. Student | Teacher / data | Training samples | Compute | Reported cost | Headline result | Source URL ------- | -------------- | ---------------- | ------- | ------------- | --------------- | ---------- s1-32B (Qwen2.5-32B-Instruct) | Gemini Flash Thinking traces | 1,000 (s1K) | 26 min on 16 H100 (7 GPU-hours) | undisclosed | AIME24 56.7 vs o1-preview 44.6; MATH500 93.0 | https://arxiv.org/abs/2501.19393 Sky-T1-32B-Preview (Qwen2.5-32B-Instruct) | QwQ-32B-Preview + GPT-4o-mini reformatting | 17k | 19 h on 8 H100 (DeepSpeed ZeRO-3), LLaMA-Factory | ~$450 (Lambda pricing) | AIME24 43.3 vs o1-preview 40.0 | https://novasky-ai.github.io/posts/sky-t1/ Bespoke-Stratos-32B | DeepSeek-R1 via Bespoke Curator | 16,710 | undisclosed (data generation ~1.5 h) | ~$800 for data generation | AIME24 63.3; MATH500 93.0; GPQA-D 58.1 | https://huggingface.co/datasets/bespokelabs/Bespoke-Stratos-17k OpenThinker3-7B (Qwen2.5-7B-Instruct) | QwQ-32B (OpenThoughts3-1.2M) | 1.2M | undisclosed on card | undisclosed | AIME25 53.3; HMMT 42.7; LCB 51.7; GPQA-D 53.7 | https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M OpenR1-Distill-7B (Qwen2.5-Math-7B-RoPE-300k) | DeepSeek-R1 (Mixture-of-Thoughts) | 350k | 8x H100 80GB node; duration undisclosed | undisclosed | AIME24 52.7; MATH-500 89.0; GPQA-D 52.8; LCB v5 39.4 | https://github.com/huggingface/open-r1 DeepSeek-R1-Distill-Qwen-7B | DeepSeek-R1 | 800k curated | undisclosed | undisclosed | AIME24 55.5; MATH-500 92.8; GPQA-D 49.1; LCB 37.6 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B Qwen3-8B (strong-to-weak distillation) | Qwen3-32B and Qwen3-235B-A22B, off- then on-policy | undisclosed | 1,800 GPU-hours (vs 17,920 RL-only) | undisclosed | AIME24 74.4 / AIME25 65.5 (RL-only: 67.6 / 55.5) | https://arxiv.org/abs/2505.09388 Llama-3.1-Minitron-4B (pruned from 8B) | Llama-3.1-8B (teacher-corrected on 94B tokens) | 94B tokens (width-pruned) / 1.4T tokens (depth-pruned) | undisclosed GPU-hours; 40x fewer tokens than from scratch | undisclosed (1.8x family savings claimed) | MMLU 60.53 (width); ~2.7x 8B throughput (depth) on H100 | https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model Llama-3.2-1B KD (torchtune case study) | LoRA-tuned Llama-3.1-8B | alpaca_cleaned | 1x A100 80GB; duration undisclosed | undisclosed | commonsense_qa 0.5717 (kd_ratio 1.0) vs 0.5536 base | https://pytorch.org/blog/llama-into-torchtune/ gpt-4o-mini distilled from gpt-4o (OpenAI cookbook) | gpt-4o stored completions | 500 wine reviews | OpenAI managed | undisclosed (gpt-4o-mini training was $3/M tokens) | Validation accuracy 79.33% vs 64.67% base, 79.67% teacher | https://developers.openai.com/cookbook/examples/leveraging_model_distillation_to_fine-tune_a_model Notes: Where authors did not publish GPU-hours or dollars the cell says undisclosed rather than an estimate. Sources: https://arxiv.org/abs/2501.19393 https://novasky-ai.github.io/posts/sky-t1/ https://arxiv.org/abs/2505.09388 https://pytorch.org/blog/llama-into-torchtune/ TABLE: DeepSeek-R1-Distill students: size vs score vs adoption [id: r1-distill-students, 6 rows] The canonical 2025 reasoning-distillation family, with Hugging Face download counts (API, 2026-09-03). Model | Params (B) | AIME 2024 pass@1 (%) | MATH-500 (%) | GPQA Diamond (%) | LiveCodeBench (%) | Downloads (30d) (downloads) | Downloads (all-time) (downloads) | Source URL ----- | ---------- | -------------------- | ------------ | ---------------- | ----------------- | ------------------------ | ------------------------ | ---------- DeepSeek-R1-Distill-Qwen-1.5B | 1.5 | 28.9 | 83.9 | 33.8 | 16.9 | 453242 | 20845014 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B DeepSeek-R1-Distill-Qwen-7B | 7 | 55.5 | 92.8 | 49.1 | 37.6 | 364966 | 15138582 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B DeepSeek-R1-Distill-Llama-8B | 8 | 50.4 | 89.1 | 49 | 39.6 | 377151 | 19567548 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-8B DeepSeek-R1-Distill-Qwen-14B | 14 | 69.7 | 93.9 | 59.1 | 53.1 | 396595 | 9133965 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-14B DeepSeek-R1-Distill-Qwen-32B | 32 | 72.6 | 94.3 | 62.1 | 57.2 | 548072 | 27257094 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Llama-70B | 70 | 70 | 94.5 | 65.2 | 57.5 | 87147 | 5882022 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B Notes: All six are MIT-licensed and were fine-tuned on 800k R1-curated samples. Downloads exclude community GGUF/MLX re-uploads (e.g. unsloth's 7B GGUF adds 27k/month). Sources: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B TABLE: Evaluation harnesses used to score distilled students [id: eval-harnesses, 3 rows] The three frameworks that appear in nearly every open distillation paper, and what each is good for. Harness | GitHub stars (stars) | License | Strength | Used by | Source URL ------- | -------------------- | ------- | -------- | ------- | ---------- EleutherAI lm-evaluation-harness | 13879 | MIT | Industry-standard task library (MMLU, GSM8K, HellaSwag, TruthfulQA); vLLM/HF backends | torchtune KD case study, most model cards | https://github.com/EleutherAI/lm-evaluation-harness Hugging Face lighteval | 2534 | MIT | 1000+ tasks incl. AIME24/25, MATH500, GPQA, LCB; sample-level result dumps; serves via HF endpoints | Open-R1, Open-RS | https://github.com/huggingface/lighteval mlfoundations evalchemy | 610 | not declared on GitHub | Reasoning benchmarks (AIME24/25, AMC23, MATH500, LiveCodeBench, GPQADiamond) with multi-GPU data-parallel sharding and completion caching; API models via Curator | OpenThinker, Bespoke-Stratos | https://github.com/mlfoundations/evalchemy Notes: Stars from the GitHub API on 2026-09-03. Evalchemy builds on lm-evaluation-harness. Sources: https://github.com/EleutherAI/lm-evaluation-harness https://github.com/huggingface/lighteval https://github.com/mlfoundations/evalchemy CHART DATA ---------- CHART: GitHub stars of the distillation toolchain [id: tool-github-stars, type: bar, unit: stars] Repository | Stars (2026-09-03) (stars) llama.cpp | 126910 vLLM | 90892 Unsloth | 75556 LLaMA-Factory | 74553 SGLang | 33881 Open-R1 | 26447 PEFT | 21626 TRL | 19210 lm-eval-harness | 13879 Axolotl | 12436 torchtune | 5802 NVIDIA Model Optimizer | 3724 distilabel | 3385 lighteval | 2534 Bespoke Curator | 1722 SpecForge | 1147 DistillKit | 1052 evalchemy | 610 Notes: Queried via api.github.com on 2026-09-03. Serving engines dwarf training libraries; dedicated distillation toolkits remain niche (DistillKit ~1k). Sources: https://github.com/vllm-project/vllm https://github.com/unslothai/unsloth https://github.com/huggingface/trl https://github.com/arcee-ai/DistillKit CHART: Managed fine-tuning price per 1M training tokens by student size [id: platform-sft-price, type: bar, unit: USD] Student size bucket | Together LoRA SFT (USD) <=16B | 0.48 17-69B (T) / 16.1-80B (F) | 1.5 70-100B (T) / 80-300B (F) | 2.9 Student size bucket | Together full SFT (USD) <=16B | 1.2 17-69B (T) / 16.1-80B (F) | 3.75 70-100B (T) / 80-300B (F) | 7.25 Student size bucket | Fireworks LoRA SFT (USD) <=16B | 0.5 17-69B (T) / 16.1-80B (F) | 3 70-100B (T) / 80-300B (F) | 6 Student size bucket | Fireworks full SFT (USD) <=16B | 1 17-69B (T) / 16.1-80B (F) | 6 70-100B (T) / 80-300B (F) | 12 Notes: The vendors’ size buckets do not line up, so categories 2 and 3 pair different ranges and are labelled with both: Together (T) buckets are <=16B / 17-69B / 70-100B; Fireworks (F) buckets are <=16B / 16.1-80B / 80-300B (plus $10 LoRA / $20 full above 300B). Bars within those categories are therefore adjacent, not equivalent. For comparison OpenAI charges $1.50 (gpt-4.1-nano), $5 (gpt-4.1-mini) and $25 (gpt-4.1) per 1M training tokens while its platform winds down. Sources: https://www.together.ai/pricing https://fireworks.ai/pricing https://developers.openai.com/api/docs/pricing CHART: Closed-model student training price (OpenAI fine-tuning) [id: closed-student-training-price, type: bar, unit: USD] Student model | Training price (USD) gpt-4.1-nano | 1.5 gpt-4o-mini | 3 gpt-4.1-mini | 5 gpt-4.1 | 25 gpt-4o | 25 Notes: o4-mini RFT is priced at $100/hour rather than per token. Batch pricing halves these rates. No GPT-5.x model is fine-tunable. Sources: https://developers.openai.com/api/docs/pricing CHART: Minimum VRAM vs student size (Unsloth) [id: vram-vs-student-size, type: line, unit: GB] Student parameters (B) | QLoRA 4-bit (GB) 3 | 3.5 7 | 5 8 | 6 9 | 6.5 11 | 7.5 14 | 8.5 27 | 22 32 | 26 40 | 30 70 | 41 81 | 48 90 | 53 405 | 237 Student parameters (B) | LoRA 16-bit (GB) 3 | 8 7 | 19 8 | 22 9 | 24 11 | 29 14 | 33 27 | 64 32 | 76 40 | 96 70 | 164 81 | 192 90 | 212 405 | 950 Notes: Unsloth calls these 'absolute minimum' figures; long reasoning traces and on-policy generation need more. Sources: https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements CHART: Reported GPU-hours of distillation runs vs student size [id: gpu-hours-vs-student-size, type: scatter, unit: GPU-hours] Student parameters (B) | Distillation (SFT on traces) (GPU-hours) 32 | 7 32 | 152 Student parameters (B) | Distillation (off- + on-policy) (GPU-hours) 8 | 1800 Student parameters (B) | RL-only baseline (GPU-hours) 8 | 17920 Notes: Only runs whose authors published GPU-hours are plotted; Sky-T1 = 8 H100 x 19 h. Qwen3 figures are GPU-hours as stated in the report without hardware detail. Sources: https://arxiv.org/abs/2501.19393 https://novasky-ai.github.io/posts/sky-t1/ https://arxiv.org/abs/2505.09388 CHART: DeepSeek-R1-Distill downloads in the last 30 days by student size [id: r1-distill-downloads, type: bar, unit: downloads] Student | Downloads last 30 days (downloads) Qwen-1.5B | 453242 Qwen-7B | 364966 Llama-8B | 377151 Qwen-14B | 396595 Qwen-32B | 548072 Llama-70B | 87147 Notes: Hugging Face API, 2026-09-03. The 32B student is the most downloaded; the 70B is the least, consistent with developers preferring students that fit one GPU. Sources: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B CHART: AIME 2024 pass@1 vs student size (R1-Distill family) [id: r1-distill-aime-vs-size, type: line, unit: %] Student parameters (B) | AIME 2024 (%) 1.5 | 28.9 7 | 55.5 8 | 50.4 14 | 69.7 32 | 72.6 70 | 70 Student parameters (B) | GPQA Diamond (%) 1.5 | 33.8 7 | 49.1 8 | 49 14 | 59.1 32 | 62.1 70 | 65.2 Notes: Returns diminish sharply above 14B for AIME; the 8B Llama student trails the 7B Qwen student, showing base-model choice matters as much as size. Sources: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B CHART: Open distillation datasets: downloads in the last 30 days [id: dataset-downloads, type: bar, unit: downloads] Dataset | Downloads last 30 days (downloads) OpenThoughts-114k | 109485 OpenR1-Math-220k | 105083 OpenMathReasoning | 69484 NuminaMath-1.5 | 33515 Bespoke-Stratos-17k | 19434 OpenThoughts3-1.2M | 19174 Mixture-of-Thoughts | 10735 Llama-Nemotron-PT | 7531 s1K-1.1 | 7133 Infinity-Instruct | 3929 Magpie-Pro-300K-F | 3349 AM-R1-Distilled-1.4M | 1793 Notes: Hugging Face API, 2026-09-03. Sources: https://huggingface.co/datasets/open-thoughts/OpenThoughts-114k https://huggingface.co/datasets/open-r1/OpenR1-Math-220k DETAIL RECORDS -------------- TOOLS AND PLATFORMS (27) - Hugging Face TRL (GKDTrainer, DistillationTrainer, GOLDTrainer, MiniLLMTrainer, SFTTrainer) — Hugging Face (library, open source) The most complete open implementation of modern (on-policy, cross-tokenizer, multimodal) distillation, with a `trl distillation` CLI. Techniques: Generalized JSD (forward/reverse KL interpolation); On-policy generation with vLLM colocate/server; Sequence-level KD (seq_kd); Universal Logit Distillation (cross-tokenizer); Reverse-KL policy gradient (MiniLLM); Tool-calling agent distillation; SFT on teacher outputs Pricing: Free, Apache-2.0; you pay for GPUs Pros: Four distillation objectives under one API; vLLM-backed generation and Liger fused JSD kernel; Same trainers used to reproduce frontier-lab recipes Cons: GKD, GOLD and MiniLLM live in trl.experimental and may change; Teacher must fit in memory alongside student unless served remotely; Default learning rates (1e-6 / 1e-7) differ sharply from SFT defaults URL: https://huggingface.co/docs/trl/distillation_trainer - Arcee DistillKit — Arcee AI (library, open source) Production toolkit behind Arcee's Virtuoso, SuperNova-Medius and Blitz models, notable for compressed offline logit capture. Techniques: Logit KD with KL, JSD, TVD; Ranking losses (hinge, logistic); Hidden-state alignment (MSE, cosine); Logit compression (~300 bytes/token) Pricing: Free, Apache-2.0 Pros: Composable loss menu; Offline mode decouples teacher inference from training; Battle-tested on real product models Cons: 1.1k stars, small community; Cross-tokenizer distillation needs a separate mergekit-tokensurgeon pre-step, not an in-trainer loss; No on-policy student sampling URL: https://github.com/arcee-ai/DistillKit - torchtune knowledge_distillation recipes — Meta / PyTorch (library, open source) Hackable pure-PyTorch KD recipe with a published ablation study (teacher fine-tuning, kd_ratio, learning rate). Techniques: Forward KL + cross-entropy with kd_ratio; LoRA student training; Teacher correction via LoRA fine-tune Pricing: Free, BSD-3-Clause Pros: Simple, readable recipe code; Documented results on truthfulqa/hellaswag/commonsense_qa; Runs on one A100 80GB Cons: Same tokenizer required; Forward KL only (no reverse KL / JSD / on-policy); Smaller ecosystem than TRL URL: https://github.com/meta-pytorch/torchtune - NVIDIA NeMo Framework + Model Optimizer (Minitron) — NVIDIA (library, open source) The prune-and-distill pipeline that produced Llama-3.1-Minitron and Nemotron families at 40x fewer tokens than training from scratch. Techniques: Logit KD (logit_layers); Intermediate-layer cosine losses (intermediate_layer_pairs); Depth and width pruning; Teacher correction Pricing: Free, Apache-2.0; NeMo containers via NGC Pros: Only mainstream stack with integrated pruning + KD; Megatron parallelism for 70B+ teachers; TensorRT-LLM deployment path Cons: NeMo checkpoint conversion required; Cluster-scale (trillion-token) recipes, not laptop-scale; NeMo repo reorganized; discoverability suffers URL: https://docs.nvidia.com/nemo-framework/user-guide/latest/model-optimization/distillation/distillation.html - Axolotl (KD plugin) — Axolotl AI (library, open source) YAML-driven trainer whose KDPlugin consumes vLLM-generated top-k teacher logprobs stored alongside the chat data. Techniques: Offline top-k logprob KD with kd_ce_alpha / kd_alpha / kd_temperature; SFT on teacher text Pricing: Free, Apache-2.0 Pros: Config-file workflow, multi-GPU/multi-node; Teacher inference done once, offline; Sample dataset published (evolkit-logprobs-pipeline-75k-v2-sample) Cons: KD not listed in the main docs index; Same tokenizer required; Top-k truncation approximates the teacher distribution URL: https://github.com/axolotl-ai-cloud/axolotl/tree/main/src/axolotl/integrations/kd - Unsloth — Unsloth AI (library, open source) Fastest, lowest-VRAM way to fine-tune a student on distilled traces: 7B in 5 GB, 70B in 41 GB. Techniques: SFT / SeqKD on teacher outputs; GRPO and DPO; 4-bit QLoRA with custom Triton kernels Pricing: Free, Apache-2.0 Pros: Lowest published VRAM floors; Colab/Kaggle notebooks for R1-Distill style SFT; Exports GGUF for llama.cpp Cons: No logit-matching KD loss; Single-GPU focus historically; Pro/enterprise tiers for multi-GPU features URL: https://github.com/unslothai/unsloth - LLaMA-Factory — hiyouga (open source) (library, open source) The trainer NovaSky used for Sky-T1-32B; broadest model coverage plus the LLaMA Board web UI. Techniques: SFT / SeqKD; Pre-training, reward modeling, PPO, DPO, KTO, ORPO, SimPO Pricing: Free, Apache-2.0 Pros: Very wide model and quantization backend support; Web UI for non-experts; Supported on Anyscale Cons: No native logit KD; Config surface is large; Less kernel-level optimization than Unsloth URL: https://github.com/hiyouga/LlamaFactory - OpenAI Model Distillation (stored completions + Evals + fine-tuning) — OpenAI (platform, proprietary) Dashboard-native distillation from a frontier OpenAI teacher into a smaller OpenAI student, now in phased shutdown. Techniques: Black-box SeqKD via stored completions; SFT, DPO; Reinforcement fine-tuning with graders Pricing: Training $25 / $5 / $1.50 per 1M tokens (gpt-4.1 / mini / nano); gpt-4o-mini $3; o4-mini RFT $100/hour; batch 50% off Pros: Zero infrastructure; cookbook shows 64.7% -> 79.3% accuracy on 500 examples; Evals integrated; Fine-tuned inference continues after wind-down Cons: No new jobs after 2027-01-06; new orgs already blocked; No GPT-5.x students; Student stays inside OpenAI; weights never leave URL: https://developers.openai.com/api/docs/guides/model-optimization - Amazon Bedrock Model Distillation — AWS (platform, proprietary) One-job managed pipeline that generates teacher data and trains a Nova or Llama student, run in us-east-1 (Nova) or us-west-2 (Llama). Techniques: Teacher synthetic-data generation with proprietary augmentation (max 15k pairs); Reuse of production invocation logs; Student fine-tuning Pricing: Teacher inference at on-demand rates; student fine-tuning at model-customization rates; custom model storage $1.95/month per Custom Model Unit (per pricing page); Llama 2 and gpt-oss-20b rates published, Nova/Llama 3.x rates not shown on public page excerpt Pros: Can distill straight from CloudWatch invocation logs with metadata filters; Data never used to train AWS models; AWS cites up to 75% cost reduction with <2% accuracy loss for RAG Cons: No Anthropic or Mistral teachers today; Region-locked; distilled Llama needs provisioned throughput or model copy; Opaque data synthesis adds teacher-inference charges URL: https://docs.aws.amazon.com/bedrock/latest/userguide/model-distillation.html - Azure OpenAI / Microsoft Foundry stored completions and distillation — Microsoft (platform, proprietary) Azure's mirror of OpenAI's distill-from-logs flow, retiring with stored completions on 2026-10-15. Techniques: Black-box SeqKD from stored completions (min 10); Evaluation on stored completions Pricing: Azure OpenAI fine-tuning rates (not on this page); stored completions capped at 10 GB Pros: Works for every Azure OpenAI model and region; Portal 'Distill' button builds the JSONL for you; Enterprise RBAC controls Cons: Stored completions retire 2026-10-15; migrate to Responses API traces; Training files cannot be exported; Classic portal only URL: https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/stored-completions - Google Vertex AI tuning (Gemini and open models) — Google Cloud (platform, proprietary) Vertex documents teacher-to-student distillation fine-tuning for open models via the GenAI SDK; the legacy text-model distillation page returns 404. Techniques: Distillation fine-tuning for open models (teacher -> student); Supervised fine-tuning; Reinforcement learning fine-tuning; Preference tuning Pricing: undisclosed on tuning overview page Pros: Teacher-to-student distillation fine-tuning for Llama 3.1 / Qwen students; Multimodal tuning of Gemini Flash-class students; RL and preference tuning available Cons: Distillation covers supported open models only, not Gemini students; Gemini ToS constraints on using outputs to train competing models apply; Legacy distillation docs removed URL: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning/open-model-tuning - Together AI Fine-tuning — Together AI (api, proprietary) Cheapest published per-token SFT for open students, plus preemptible H100s for DIY distillation. Techniques: LoRA and full SFT; DPO; Serverless deployment of fine-tuned models Pricing: LoRA SFT $0.48/M (<=16B), $1.50 (17-69B), $2.90 (70-100B); full $1.20/$3.75/$7.25; DPO ~12% higher; $4 minimum; GPU clusters H100 $1.99 preemptible / $3.99 on-demand Pros: Lowest LoRA rate found; No GPU management; Preemptible cluster pricing Cons: No logit-level KD (text only); Specialized-model surcharges up to $40/M (GLM-5.x); Buckets not aligned with Fireworks URL: https://www.together.ai/pricing - Fireworks AI Fine-tuning — Fireworks AI (api, proprietary) Per-token SFT with free serverless deployment of the resulting LoRA and RFT for reasoning students. Techniques: LoRA and full SFT; DPO; Reinforcement fine-tuning (GPU-hour billed) Pricing: LoRA SFT $0.50/M (<=16B), $3 (16.1-80B), $6 (80-300B), $10 (>300B); full 2x; RFT at $8/h H100/H200, $13/h B200 from 2026-09-01; datasets 3 to 3M examples (from the fine-tuning docs, not the pricing page) Pros: Free serverless deployment of fine-tunes; RFT available; Large dataset ceiling (3M) Cons: Mid-size bucket is 2x Together's; GPU prices rose 14-30% on 2026-09-01; No logit KD URL: https://fireworks.ai/pricing - Predibase (now Rubrik Agent Cloud) — Predibase / Rubrik (platform, proprietary) Early RFT platform whose public pricing disappeared after folding into Rubrik. Techniques: SFT LoRA; Turbo LoRA; Reinforcement fine-tuning (GRPO); LoRAX multi-adapter serving Pricing: undisclosed: predibase.com and docs.predibase.com redirect to rubrik.com; third-party trackers list SFT LoRA $0.50/M (<=16B) and RFT GRPO $10/M, unverified Pros: Pioneered hosted RFT; LoRAX serving of many adapters Cons: Public docs and pricing gone; Roadmap under Rubrik unclear URL: https://predibase.com/pricing - Databricks Mosaic AI Foundation Model Fine-tuning — Databricks (platform, proprietary) DBU-metered fine-tuning inside the lakehouse, priced per words rather than tokens. Techniques: Supervised fine-tuning; Continued pre-training; Chat completion tuning Pricing: $0.65/DBU; Llama 3.2 1B ~25 DBU (~$16) per 10M words, Llama 3.1 8B ~100 DBU (~$65), Llama 3.3 70B ~225 DBU (~$146); 500M words: $715 to $7,150 Pros: Governance via Unity Catalog; Predictable per-word estimates published; Serving with provisioned throughput Cons: DBU accounting is opaque; No KD-specific features; Requires Databricks workspace URL: https://www.databricks.com/product/pricing/mosaic-foundation-model-training - Modal — Modal Labs (platform, proprietary) Per-second GPU billing that suits bursty distillation experiments better than reserved clusters. Techniques: Serverless GPU functions for TRL/Unsloth/torchtune jobs; Batch teacher inference with vLLM Pricing: Per-second: B200 $6.25/h, H200 $4.54/h, H100 $3.95/h, A100-80GB $2.50/h, A100-40GB $2.10/h, L40S $1.95/h, A10 $1.10/h, L4 $0.80/h, T4 $0.59/h; $30/month free credits Pros: No idle cost, fast cold starts; Free monthly credit covers small QLoRA runs; Full GPU range from T4 to B200 Cons: H100 dearer than Together preemptible ($1.99); Region pinning multiplies price; You manage the training code URL: https://modal.com/pricing - Anyscale (Ray) post-training — Anyscale (platform, proprietary) Ray-native orchestration for distributed distillation and RL, with Ray Data for large-scale teacher generation. Techniques: CPT, SFT; RLHF (PPO, DPO, KTO, ORPO); RLVR (GRPO, DAPO); LoRA/QLoRA, FSDP, DeepSpeed, Megatron Pricing: undisclosed on docs page Pros: Scales teacher inference and student training on one substrate; Multiple open frameworks supported Cons: No dedicated distillation product; Pricing not public URL: https://docs.anyscale.com/llm/fine-tuning - EleutherAI lm-evaluation-harness — EleutherAI (library, open source) The default harness for reporting student-vs-teacher retention on classic benchmarks. Techniques: Standardized benchmarks (MMLU, GSM8K, HellaSwag, TruthfulQA, commonsense_qa) Pricing: Free, MIT Pros: 13.9k stars, industry standard; Backend-agnostic Cons: Reasoning benchmarks with long CoT are better served by evalchemy/lighteval URL: https://github.com/EleutherAI/lm-evaluation-harness - Hugging Face lighteval — Hugging Face (library, open source) Open-R1's evaluation backend; the quickest way to score a distilled reasoning model on AIME/MATH500. Techniques: 1000+ tasks incl. AIME24/25, MATH500, GPQA, LiveCodeBench, HLE; Sample-level result dumps Pricing: Free, MIT Pros: Reasoning benchmarks built in; Custom tasks/metrics; Open Benchmark Index Cons: Smaller community than lm-eval (2.5k stars) URL: https://github.com/huggingface/lighteval - Evalchemy — DataComp / Bespoke Labs (library, open source) The harness behind OpenThinker and Bespoke-Stratos numbers, optimized for long chain-of-thought evaluation. Techniques: AIME24/25, AMC23, MATH500, LiveCodeBench, GPQADiamond, HumanEvalPlus, BigCodeBench; Data-parallel sharding and completion caching Pricing: Free (license not declared) Pros: Caches full completions to avoid re-inference; Multi-GPU sharding Cons: 610 stars; last push Feb 2026; No SPDX license on GitHub URL: https://github.com/mlfoundations/evalchemy - vLLM — vLLM project (library, open source) 90.9k-star engine that both generates distillation data and serves the resulting student, optionally with an EAGLE draft. Techniques: High-throughput teacher sampling; Speculative decoding with distilled draft heads; Weight sync for online distillation Pricing: Free, Apache-2.0 Pros: Native TRL integration for on-policy KD; EAGLE 3.1: 2.03x per-user throughput on Kimi-K2.6 Cons: GPU-only; Colocated training/inference can contend for memory URL: https://github.com/vllm-project/vllm - SGLang + SpecForge — SGLang project (library, open source) Purpose-built trainer for speculative-decoding draft models, claiming up to 4x inference speedup with SGLang. Techniques: Online disaggregated and offline colocated/disaggregated draft training; `specforge train --config` Pricing: Free, MIT Pros: One command for all training topologies; Directly compatible with SGLang serving Cons: Training cost/time not published; 1.1k stars, young project URL: https://github.com/sgl-project/SpecForge - llama.cpp — ggml (library, open source) Where most distilled 1.5B-14B students actually run: the 126.9k-star local inference engine. Techniques: GGUF quantization; Speculative decoding with a small draft GGUF Pricing: Free, MIT Pros: Runs on laptops and phones; Huge community of GGUF re-uploads Cons: No training; Throughput far below vLLM/SGLang on servers URL: https://github.com/ggml-org/llama.cpp - Bespoke Curator — Bespoke Labs (library, open source) The tool that produced Bespoke-Stratos-17k from DeepSeek-R1 in ~1.5 hours for ~$800. Techniques: Batched synthetic data generation with caching and structured outputs; Rejection sampling pipelines Pricing: Free, Apache-2.0; you pay teacher API costs Pros: Provider-agnostic; Cost-aware batching; Also used by evalchemy for API models Cons: Black-box text only; 1.7k stars URL: https://github.com/bespokelabsai/curator - Argilla distilabel — Argilla / Hugging Face (library, open source) Pipeline framework for teacher-generated SFT/DPO data that plugs into Argilla for human review. Techniques: Synthetic instruction, preference and Magpie-style prompt generation; LLM-as-judge labeling Pricing: Free, Apache-2.0 Pros: Magpie pipeline built in; Pushes directly to the Hub Cons: No training component; Pipelines can be verbose URL: https://github.com/argilla-io/distilabel - Hugging Face Open-R1 — Hugging Face (library, open source) Reference reproduction of R1-Distill: OpenR1-Distill-7B scores AIME24 52.7 / MATH-500 89.0 on an 8xH100 node. Techniques: SFT distillation recipes (accelerate + DeepSpeed ZeRO-3); GRPO Pricing: Free, Apache-2.0 Pros: End-to-end scripts and datasets; 26.4k stars Cons: 8xH100 assumed; Training durations not documented URL: https://github.com/huggingface/open-r1 - Hugging Face PEFT — Hugging Face (library, open source) The adapter layer (21.6k stars) that lets TRL distill into a LoRA student instead of full weights. Techniques: LoRA, QLoRA, DoRA and other adapters Pricing: Free, Apache-2.0 Pros: Universal; Tiny artifacts to ship Cons: TRL's DistillationTrainer rejects adapters on lm_head and prompt-learning methods URL: https://github.com/huggingface/peft COSTED RECIPES (8) - Distill a frontier API teacher into Llama-3.1-8B for classification (black-box SeqKD) Tool: Unsloth / TRL SFTTrainer or Together AI fine-tuning 1. Collect 5k-20k representative unlabeled inputs from production logs; hold out 500 for evaluation. 2. Label them with the teacher (e.g. deepseek-v4-pro at $0.66-$1.32/M input and $1.98-$3.96/M output, or an OpenAI gpt-5.6 model at the rates on OpenAI’s published pricing page) using a strict JSON schema; 20k inputs x ~600 tokens is ~12M input tokens, roughly $8-$16 of teacher calls at DeepSeek's V4-Pro rates plus <$5 of output. Check the teacher's terms of service on training competing models first. 3. Filter: drop malformed labels and, where a rubric exists, self-consistency-check by sampling twice and keeping agreements (rejection sampling). 4. Train Llama-3.1-8B-Instruct with QLoRA in Unsloth or TRL SFTTrainer (6 GB VRAM minimum; an L4 or A100 on Modal) for 2-3 epochs, or upload JSONL to Together (LoRA SFT $0.48/M tokens, $4 minimum) or Fireworks ($0.50/M). 5. Evaluate with lm-evaluation-harness or a custom accuracy script on the 500 held-out teacher labels; target >95% agreement with the teacher. 6. Merge the adapter, quantize to GGUF or AWQ, serve with vLLM (or llama.cpp at the edge) and compare per-request cost against the teacher API. Estimated cost: $20-$40 total: ~$10-$20 teacher labels (DeepSeek V4-Pro rates) + ~$10-$20 training (Together: 20k x 650 tokens x 3 epochs = 39M tokens x $0.48 = ~$19; or ~2-4 h on a $2.50/h A100) Estimated time: Half a day: 1-2 h teacher labeling, 2-4 h training, 1 h evaluation (author estimate from listed prices) Sources: https://www.together.ai/pricing https://api-docs.deepseek.com/quick_start/pricing https://developers.openai.com/api/docs/pricing - Reasoning distillation of a 7B model from DeepSeek-R1 traces on one GPU Tool: Unsloth or TRL SFTTrainer + evalchemy 1. Pick an open trace set: Bespoke-Stratos-17k (16,710 verified R1 traces, Apache-2.0) for a fast run, or the 94k 'default' split of OpenR1-Math-220k for math-heavy students. 2. Start from Qwen2.5-7B-Instruct (the base used by OpenThinker3-7B and Bespoke-Stratos-7B); set max sequence length to 16k so long chains of thought are not truncated. 3. Train QLoRA (rank 64) with Unsloth or TRL SFTTrainer on a single 80 GB A100/H100 (Unsloth's 5 GB floor does not cover 16k-token traces); use packing, gradient checkpointing, lr 1e-5 to 2e-5, 1-3 epochs. 4. Evaluate AIME24/25, MATH500, GPQA-Diamond and LiveCodeBench with evalchemy or lighteval (the same harnesses the source datasets used) and compare to DeepSeek-R1-Distill-Qwen-7B (AIME24 55.5, MATH-500 92.8). 5. If short on compute, replicate s1 instead: 1,000 curated samples took 26 minutes on 16 H100s (7 GPU-hours) for a 32B student, so a 7B student on 1k samples is well under 2 GPU-hours. 6. Export to GGUF for llama.cpp or serve with vLLM; consider a 1.5B EAGLE-style draft for speed. Estimated cost: $70-$140 of GPU rental (Together preemptible H100 $1.99/h to Modal H100 $3.95/h); dataset free. Scaled from Sky-T1's 152 H100-hours for a 32B student on 17k samples (~4.5x fewer FLOPs at 7B); actual throughput varies with sequence length Estimated time: ~1-1.5 days wall-clock on one H100 for 17k samples x 3 epochs; ~2 hours for an s1-style 1k-sample run Sources: https://novasky-ai.github.io/posts/sky-t1/ - Amazon Bedrock Model Distillation walkthrough (Llama 3.3 70B -> Llama 3.2 3B) Tool: Amazon Bedrock Model Distillation 1. In us-west-2, create an IAM service role with S3 access to your training bucket and permission to invoke the teacher (and its cross-region inference profile if you choose one). 2. Prepare prompts as JSONL (Converse-format messages); optionally add a few golden prompt-response pairs that Bedrock uses to steer the teacher. Alternatively enable CloudWatch invocation logging and tag production calls with requestMetadata so the job can reuse real teacher responses. 3. Create the distillation job in the console or CreateModelCustomizationJob API with teacher meta.llama3-3-70b-instruct-v1:0 and student meta.llama3-2-3b-instruct-v1:0:128k; Bedrock generates and augments responses (up to 15k pairs), splits train/validation, and fine-tunes. 4. Expect two charges: teacher inference at on-demand rates for every synthesized response, and student training at Llama customization rates; the pricing page shows $1.95/month storage per Custom Model Unit, and a model can need more than one unit. 5. When the job finishes, buy provisioned throughput in us-west-2 or copy the model to another region, then evaluate with Bedrock Evaluations or your own harness. 6. Note that Claude teachers are not currently offered and there is no announced restoration date; use Nova Premier -> Nova Pro/Lite/Micro (us-east-1) if you need an Amazon teacher. Estimated cost: Teacher synthesis: up to 15k prompts x ~1k tokens = ~15M tokens at Llama 3.3 70B on-demand rates, plus student training at Llama customization rates (Llama 3.x per-token training rates are not shown on the public pricing page excerpt; Llama 2 13B is $1.49/M tokens) plus $1.95/month storage per Custom Model Unit; provisioned throughput billed hourly for serving Estimated time: Hours per job (job runtime undisclosed); data prep 1-2 h Sources: https://docs.aws.amazon.com/bedrock/latest/userguide/model-distillation.html - On-policy GKD with TRL (Qwen3-8B teacher -> Qwen3-1.7B student) Tool: Hugging Face TRL DistillationTrainer / GKDTrainer 1. pip install trl[vllm]; prepare a prompt-only conversational dataset (e.g. trl-lib/ultrafeedback-prompt or your own prompts); no teacher answers are needed because the student generates on-policy. 2. Use DistillationTrainer(model='Qwen/Qwen3-1.7B', teacher_model='Qwen/Qwen3-8B') with DistillationConfig(use_vllm=True, vllm_mode='colocate', vllm_gpu_memory_utilization=0.3, max_completion_length=512, beta=1.0 for reverse KL or 0.5 for JSD, learning_rate=1e-6, bf16=True). For the classic GKD knobs use trl.experimental.gkd.GKDTrainer with lmbda=0.5 (half on-policy), beta=0.5, temperature=0.9, seq_kd=False. 3. Launch with `accelerate launch train_distillation.py` or the CLI: `trl distillation --model_name_or_path ... --teacher_model_name_or_path ... --use_peft --lora_r 64` for a LoRA student. 4. Watch completions/mean_length, clipped_ratio and entropy; if the student's completions are truncated raise max_completion_length. Enable use_liger_kernel=True for the fused JSD unless the model uses logit soft-capping (Gemma). 5. vllm_mode='server' moves the student's generation to a separate vLLM server (`vllm serve `), not the teacher's. For a teacher that does not fit alongside the student, use AsyncDistillationTrainer, which scores against a teacher served over HTTP: run `vllm serve --logprobs-mode processed_logprobs --max-logprobs -1` on separate GPUs. 6. Evaluate against the teacher on lighteval tasks; Qwen3's report saw distilled 8B students beat RL-only training at ~1/10 of the GPU hours. Estimated cost: $10-$40: 8B teacher (bf16 ~16 GB) + 1.7B student + vLLM engine fit on one 80 GB A100 ($2.50/h on Modal) for a few hours; author estimate Estimated time: 2-8 hours for 10k-50k prompts at 512-token completions on one A100 (author estimate; generation dominates step time) Sources: https://huggingface.co/docs/trl/distillation_trainer - Train an EAGLE-3 draft model for speculative decoding and serve it in vLLM Tool: SpecForge (SGLang) or TorchSpec + vLLM 1. Choose the target (the distilled student itself, or a large model such as Llama-3.3-70B). Check whether a pre-trained head already exists on the Hub (e.g. yuhuili/EAGLE3-LLaMA3.3-Instruct-70B, lightseekorg/kimi-k2.6-eagle3.1-mla) before training. 2. Install SpecForge (`pip install specforge` from sgl-project/SpecForge) and assemble a prompt corpus resembling production traffic (chat-style datasets such as ShareGPT/UltraChat are the usual starting point). 3. Run `specforge train --config ` in online mode (target runs live to produce hidden states) or offline mode (hidden states cached first); the draft is a single-layer head that reads the target's feature vectors. 4. Serve with vLLM: `vllm serve --speculative-config '{"model":"","method":"eagle3","num_speculative_tokens":3}'`, or with SGLang's speculative flags. 5. Benchmark acceptance length and tokens/s at your real concurrency: EAGLE 3.1 reports 2.03x per-user throughput at concurrency 1 on Kimi-K2.6-NVFP4 falling to 1.66x at concurrency 16; gains shrink as batch size grows. 6. Re-train the head whenever the target is re-distilled or its chat template changes, since the draft is tied to the target's hidden-state distribution. Estimated cost: undisclosed by SpecForge/TorchSpec; budget multi-GPU hours because the target model must run over the whole corpus (70B targets need >=2 H100s at $3.95/h each on Modal) Estimated time: Hours to a day depending on corpus size and target size (author estimate) Sources: https://vllm.ai/blog/2026-05-26-eagle-3-1 - Prune-and-distill a Llama-3.1-8B into a 4B (NVIDIA Minitron recipe) Tool: NVIDIA NeMo Framework + Model Optimizer 1. Convert Llama-3.1-8B to a NeMo 2.0 checkpoint and run teacher correction: continue-train the 8B on your distillation corpus (NVIDIA used 94B tokens) so the teacher's distribution matches the data. 2. Estimate importance with ~1,024 calibration samples, then prune: depth (drop 16 of 32 layers, e.g. layers 16-31) for maximum speed, or width (hidden 4096 -> 3072, MLP 14336 -> 9216) for maximum accuracy. 3. Distill with NeMo/Model Optimizer using logit KD (logit_layers) plus optional intermediate-layer cosine losses; NVIDIA used 94B tokens for the width-pruned and 1.4T for the depth-pruned student. 4. Evaluate: the width-pruned 4B reached MMLU 60.53 and HellaSwag 76.06; the depth-pruned variant runs ~2.7x the 8B's throughput under TensorRT-LLM on H100 (~1.8x for width), with FP8 adding ~1.3x. 5. Export to TensorRT-LLM or vLLM; repeat pruning from the same corrected teacher to build a family (NVIDIA reports 40x fewer tokens per additional model and 1.8x total compute savings from its earlier Nemotron family work). Estimated cost: undisclosed by NVIDIA; trillion-token retraining implies a multi-node H100 cluster, i.e. tens of thousands of GPU-hours Estimated time: Days to weeks on a cluster (author estimate) Sources: https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model - OpenAI stored-completions distillation (gpt-4o -> gpt-4.1-mini) before the platform closes Tool: OpenAI Model Distillation (stored completions + Evals + fine-tuning) 1. Confirm eligibility: since 2026-05-07 only organizations that have previously fine-tuned can create jobs, and after 2026-07-02 only those with fine-tuned inference in the last 60 days; all new jobs stop 2027-01-06. 2. Add store=True and metadata tags to production teacher calls; the cookbook stored 500 gpt-4o wine-classification completions. 3. In the dashboard filter stored completions by metadata/model, run an Eval to record the teacher baseline (cookbook: 79.67% validation accuracy vs 64.67% for base gpt-4o-mini), then click Distill and pick the student (gpt-4.1-mini at $5/M training tokens, gpt-4.1-nano at $1.50/M). 4. Launch the fine-tune (500 examples x ~300 tokens x 3 epochs = ~0.45M tokens, about $2.25 on gpt-4.1-mini) and re-run the same Eval on the fine-tuned model; the cookbook's distilled gpt-4o-mini reached 79.33%. 5. Serve the fine-tuned model (gpt-4.1-mini fine-tuned inference $0.80 in / $3.20 out per 1M) and plan a migration to an open student before base-model deprecation. Estimated cost: ~$2-$10 training for 500-2,000 examples on gpt-4.1-mini/nano; teacher calls at normal inference rates; evals billed as inference Estimated time: 1-2 hours of dashboard work plus fine-tune queue time Sources: https://developers.openai.com/cookbook/examples/leveraging_model_distillation_to_fine-tune_a_model - Cross-tokenizer distillation with TRL GOLD (Qwen teacher -> Llama student) Tool: Hugging Face TRL GOLDTrainer (experimental) 1. Install TRL and load a conversational dataset with assistant answers (GOLD needs messages, e.g. HuggingFaceTB/OpenR1-Math-220k-default-verified or trl-lib/chatbot_arena_completions). 2. Instantiate GOLDTrainer(model='meta-llama/Llama-3.2-1B-Instruct', teacher_model='Qwen/Qwen2.5-0.5B-Instruct' or a larger Qwen, args=GOLDConfig(use_uld_loss=True, teacher_tokenizer_name_or_path=, uld_use_hybrid_loss=True, lmbda=0.5, beta=0.5)). 3. GOLD aligns visible text spans across the two tokenizers and merges teacher probabilities per student token, so no completion tokens are dropped; hybrid loss uses exact JSD for shared vocabulary items and ULD for the rest. 4. Set dataloader_drop_last=True to avoid the undersized-final-batch warning; enable use_vllm=True for faster on-policy rollouts. 5. Evaluate with lighteval and compare against a same-tokenizer baseline; for VLMs, GOLD supports Qwen3-VL-8B -> Qwen3-VL-2B (JSD) or cross-family LFM2.5-VL students (ULD). Estimated cost: $5-$30 on a single A100/L40S for 1B-3B students (author estimate from Modal rates) Estimated time: 1-6 hours (author estimate) Sources: https://huggingface.co/docs/trl/gold_trainer TIMELINE OF THIS PERSPECTIVE (20 EVENTS) ---------------------------------------- 2024-08 | product | Arcee releases DistillKit v0.1 2024-08 | research | NVIDIA publishes the Llama-3.1-Minitron prune-and-distill recipe 2024-10-01 | product | OpenAI launches Model Distillation in the API 2024-12 | product | Amazon Bedrock Model Distillation announced (preview) 2025-01-10 | research | Sky-T1-32B-Preview: an o1-preview-class model for ~$450 2025-01-20 | product | DeepSeek-R1 ships six distilled students (1.5B-70B, MIT) 2025-01-22 | research | Bespoke-Stratos-17k released 2025-01-31 | research | s1: 1,000 samples and 26 minutes on 16 H100s beat o1-preview on AIME24 2025-02 | research | Hugging Face Open-R1 publishes OpenR1-Math-220k 2025-02 | research | torchtune KD case study: Llama-3.1-8B into Llama-3.2-1B on one A100 2025-05 | research | Qwen3 report: distillation beats RL at 1/10 the GPU-hours 2025-06-05 | research | OpenThoughts3-1.2M and OpenThinker3-7B 2025-07-01 | product | EAGLE-3 speculative decoding lands in vLLM 2026-02-17 | product | Bedrock adds reinforcement fine-tuning for gpt-oss-20b and Qwen3-32B 2026-05-07 | product | OpenAI begins winding down the fine-tuning platform 2026-05-26 | product | EAGLE 3.1 released with vLLM and TorchSpec 2026-07-08 | research | Hugging Face surveys distillation in 2026 frontier models 2026-09-01 | market | Fireworks raises on-demand GPU prices 2026-10-15 | product | Azure OpenAI stored completions retire 2027-01-06 | product | OpenAI stops all new fine-tuning jobs SOURCES CITED BY THIS SECTION (58) ---------------------------------- [1] Generalized Knowledge Distillation Trainer (TRL docs) — Hugging Face, 2026 (docs) https://huggingface.co/docs/trl/gkd_trainer [2] Distillation Trainer (TRL docs) — Hugging Face, 2026 (docs) https://huggingface.co/docs/trl/distillation_trainer [3] GOLD Trainer (TRL docs) — Hugging Face, 2026 (docs) https://huggingface.co/docs/trl/gold_trainer [4] MiniLLM Trainer (TRL docs) — Hugging Face, 2026 (docs) https://huggingface.co/docs/trl/main/minillm [5] Distillation in 2026 (so far): which frontier models use it and how — Hugging Face, 2026-07-08 (blog) https://huggingface.co/blog/sergiopaniego/distillation-2026 [6] arcee-ai/DistillKit — Arcee AI, 2026 (docs) https://github.com/arcee-ai/DistillKit [7] DistillKit v0.1 technical paper — Arcee AI, 2024-08 (blog) https://blog.arcee.ai/distillkit-v0-1-by-arcee-ai [8] Distilling Llama3.1 8B into 1B in torchtune — PyTorch, 2025-02 (blog) https://pytorch.org/blog/llama-into-torchtune/ [9] meta-pytorch/torchtune — Meta, 2026 (docs) https://github.com/meta-pytorch/torchtune [10] How to Prune and Distill Llama-3.1 8B to an NVIDIA Llama-3.1-Minitron 4B Model — NVIDIA, 2024-08 (blog) https://developer.nvidia.com/blog/how-to-prune-and-distill-llama-3-1-8b-to-an-nvidia-llama-3-1-minitron-4b-model [11] NeMo Framework: Knowledge Distillation — NVIDIA, 2026 (docs) https://docs.nvidia.com/nemo-framework/user-guide/latest/model-optimization/distillation/distillation.html [12] NVIDIA/Model-Optimizer — NVIDIA, 2026 (docs) https://github.com/NVIDIA/Model-Optimizer [13] Axolotl KD plugin README — Axolotl AI, 2026 (docs) https://github.com/axolotl-ai-cloud/axolotl/tree/main/src/axolotl/integrations/kd [14] Unsloth requirements (VRAM table) — Unsloth, 2026 (docs) https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements [15] hiyouga/LlamaFactory — LLaMA-Factory, 2026 (docs) https://github.com/hiyouga/LlamaFactory [16] OpenAI API deprecations (fine-tuning platform wind-down) — OpenAI, 2026 (docs) https://developers.openai.com/api/docs/deprecations [17] OpenAI model optimization guide — OpenAI, 2026 (docs) https://developers.openai.com/api/docs/guides/model-optimization [18] OpenAI API pricing (fine-tuning section) — OpenAI, 2026 (pricing) https://developers.openai.com/api/docs/pricing [19] Model Distillation in the API — OpenAI, 2024-10-01 (blog) https://openai.com/index/api-model-distillation/ [20] Leveraging model distillation to fine-tune a model (cookbook) — OpenAI, 2024-10 (docs) https://developers.openai.com/cookbook/examples/leveraging_model_distillation_to_fine-tune_a_model [21] Customize a model with distillation in Amazon Bedrock — AWS, 2026 (docs) https://docs.aws.amazon.com/bedrock/latest/userguide/model-distillation.html [22] Prerequisites for model distillation (supported teacher/student pairs) — AWS, 2026 (docs) https://docs.aws.amazon.com/bedrock/latest/userguide/prequisites-model-distillation.html [23] Amazon Bedrock pricing — AWS, 2026 (pricing) https://aws.amazon.com/bedrock/pricing/ [24] Amazon Bedrock reinforcement fine-tuning adds open-weight models — AWS, 2026-02-17 (news) https://aws.amazon.com/about-aws/whats-new/2026/02/amazon-bedrock-reinforcement-fine-tuning-openai/ [25] Stored completions and distillation (Foundry classic) — Microsoft, 2026-07-06 (docs) https://learn.microsoft.com/en-us/azure/foundry-classic/openai/how-to/stored-completions [26] Distillation in Azure AI Foundry (blog) — Microsoft, 2024-12 (blog) https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/distillation-turning-smaller-models-into-high-performance-cost-effective-solutio/4355029 [27] Vertex AI: tune Gemini models (overview) — Google Cloud, 2026 (docs) https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/tune-models [28] Supervised and distillation fine-tuning for open models — Google Cloud, 2026-09-03 (docs) https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning/open-model-tuning [29] Together AI pricing — Together AI, 2026 (pricing) https://www.together.ai/pricing [30] Fireworks AI pricing — Fireworks AI, 2026-09 (pricing) https://fireworks.ai/pricing [31] Fireworks fine-tuning docs — Fireworks AI, 2026 (docs) https://docs.fireworks.ai/fine-tuning/fine-tuning-models [32] Databricks Foundation Model Fine-tuning pricing — Databricks, 2026 (pricing) https://www.databricks.com/product/pricing/mosaic-foundation-model-training [33] Modal pricing — Modal, 2026 (pricing) https://modal.com/pricing [34] Post-training for LLMs on Anyscale — Anyscale, 2026 (docs) https://docs.anyscale.com/llm/fine-tuning [35] Predibase pricing (redirects to Rubrik) — Predibase / Rubrik, 2026 (pricing) https://predibase.com/pricing [36] DeepSeek API models and pricing — DeepSeek, 2026 (pricing) https://api-docs.deepseek.com/quick_start/pricing [37] DeepSeek-R1-Distill-Qwen-7B model card — DeepSeek, 2025-01-20 (docs) https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B [38] OpenThoughts3-1.2M dataset card — Open Thoughts, 2025-06 (docs) https://huggingface.co/datasets/open-thoughts/OpenThoughts3-1.2M [39] OpenThoughts: Data Recipes for Reasoning Models — arXiv, 2025-06 (paper) https://arxiv.org/pdf/2506.04178 [40] Bespoke-Stratos-17k dataset card — Bespoke Labs, 2025-01 (docs) https://huggingface.co/datasets/bespokelabs/Bespoke-Stratos-17k [41] OpenR1-Math-220k dataset card — Hugging Face, 2025-02 (docs) https://huggingface.co/datasets/open-r1/OpenR1-Math-220k [42] huggingface/open-r1 — Hugging Face, 2026 (docs) https://github.com/huggingface/open-r1 [43] Magpie: Alignment Data Synthesis from Scratch — arXiv / ICLR 2025, 2024-06 (paper) https://arxiv.org/abs/2406.08464 [44] Infinity Instruct: Scaling Instruction Selection and Synthesis — arXiv / BAAI, 2025-06 (paper) https://arxiv.org/html/2506.11116 [45] Sky-T1: Train your own O1 preview model within $450 — NovaSky (UC Berkeley), 2025-01-10 (blog) https://novasky-ai.github.io/posts/sky-t1/ [46] s1: Simple test-time scaling — arXiv, 2025-01-31 (paper) https://arxiv.org/abs/2501.19393 [47] Qwen3 Technical Report — arXiv / Alibaba, 2025-05 (paper) https://arxiv.org/abs/2505.09388 [48] EleutherAI/lm-evaluation-harness — EleutherAI, 2026 (docs) https://github.com/EleutherAI/lm-evaluation-harness [49] huggingface/lighteval — Hugging Face, 2026 (docs) https://github.com/huggingface/lighteval [50] mlfoundations/evalchemy — DataComp / Bespoke Labs, 2026 (docs) https://github.com/mlfoundations/evalchemy [51] vllm-project/vllm — vLLM, 2026 (docs) https://github.com/vllm-project/vllm [52] EAGLE 3.1: Advancing Speculative Decoding Through Collaboration — vLLM, 2026-05-26 (blog) https://vllm.ai/blog/2026-05-26-eagle-3-1 [53] Fly Eagle(3) fly: faster inference with vLLM and speculative decoding — Red Hat, 2025-07-01 (blog) https://developers.redhat.com/articles/2025/07/01/fly-eagle3-fly-faster-inference-vllm-speculative-decoding [54] sgl-project/SpecForge — SGLang, 2026 (docs) https://github.com/sgl-project/SpecForge [55] sgl-project/sglang — SGLang, 2026 (docs) https://github.com/sgl-project/sglang [56] ggml-org/llama.cpp — ggml, 2026 (docs) https://github.com/ggml-org/llama.cpp [57] bespokelabsai/curator — Bespoke Labs, 2026 (docs) https://github.com/bespokelabsai/curator [58] argilla-io/distilabel — Argilla / Hugging Face, 2026 (docs) https://github.com/argilla-io/distilabel ================================================================================================ 8. CUSTOMER PERSPECTIVE — THE BUYER’S VIEW OF AI DISTILLATION: WHICH SMALL MODEL TO ACTUALLY DEPLOY ================================================================================================ Page: https://global-distillation.com/customer Data: https://global-distillation.com/data/customer.json Updated: 2026-09-04 Content: 8 key figures, 5 tables, 6 charts, 21 dated events, 16 glossary terms, 78 sources SUMMARY -------- For an enterprise buyer in September 2026, the distillation question is no longer "is the small model good enough" but "which small model, and what does the licence let me do with it". Within a single generation, distilled and small-sibling tiers retain 88–98% of their larger sibling's benchmark score for 3–25% of the price: GPT-5.6 Luna holds about 94% of Sol's GPQA Diamond at 5% of the input rate, DeepSeek-V4-Flash scores 98% of V4-Pro's at a third of the price, and Gemini 2.5 Flash — which Google explicitly documents as a k-sparse logit distillation of 2.5 Pro — holds 96% at 24%; across a large parameter gap, retention falls to 47–76%, as DeepSeek's own R1 students and Google's Gemma E-series show. The gap that remains is not knowledge but agentic reliability: on SWE-bench Verified, Claude Sonnet 5 retains 89% of Opus 5 and Haiku 4.5 only 76%, so long-horizon coding and tool-use workloads are the last place a frontier tier still pays for itself. Open weights have become the buyer's real leverage — DeepSeek V4 (MIT), Gemma 4 (Apache 2.0), Qwen3.8-27B (Apache 2.0) and Mistral Small 4 (Apache 2.0) all clear 71.2–90.1 GPQA Diamond with no per-token fee and no anti-distillation clause — which is why the Silicon Data enterprise inference index fell from $2.04/MTok on 31 May 2026 to $1.16–1.18 in early August. The three things that should drive the decision, in order, are: whether your traffic needs frontier-grade agentic reliability on more than 15% of calls; whether the vendor's terms let you distil its outputs into your own model (OpenAI and Anthropic say no, DeepSeek's MIT licence says yes); and whether you can absorb the migration cost when a tier is repriced or retired — which, on 2026 evidence, happens roughly every quarter. KEY FIGURES ----------- - Enterprise inference price index: 1.17 USD / M tokens (-43% since 31 May 2026) Silicon Data index cited by Jefferies; hit a 2026 low of $1.16–$1.18 on 6–8 Aug 2026, down from $2.04 on 31 May and $1.45 in late July. Source: https://www.scmp.com/tech/tech-trends/article/3363549/enterprise-ai-costs-hit-2026-low-driven-price-wars-chinese-open-source-models-research - Cheapest hosted model per 1M tokens (in): 0.035 USD (Amazon Nova Micro) Amazon's cheapest Nova tier. AWS supports it as a distillation student in Bedrock Model Distillation, but does not state that the shipped weights were distilled from Premier or Pro. Text-only, 128K context. Source: https://pricepertoken.com/pricing-page/model/amazon-nova-micro-v1 - GPQA retained by GPT-5.6 Luna vs Sol: 94.2 % (at 5% of the input price (94.2–97.3% across providers)) 87.0 vs 92.4 GPQA Diamond; $0.20 vs $4.00 per M input tokens. Both figures are third-party (OpenRouter auto-routing); provider-specific Luna scores on the same page run to 89.9, and DataLearner reports Sol at 93.5 at max thinking, so retention is best read as a 94.2–97.3% range pending the GPT-5.6 system card. Source: https://openrouter.ai/openai/gpt-5.6-luna - GPQA of DeepSeek-V4-Flash vs V4-Pro: 97.8 % (at 33% of the price) 88.1 vs 90.1 GPQA Diamond, both MIT-licensed open weights. Same-family comparison: DeepSeek documents V4-Flash as a consolidation of V4 domain experts via on-policy distillation, not as a distillation of V4-Pro. Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash - SWE-bench Verified retained by Claude Haiku 4.5 vs Opus 5: 76.4 % (at 20% of the price) 73.3 vs 96.0. Agentic coding is where the distilled tiers still lose most. Source: https://datanorth.ai/news/claude-opus-5-by-anthropic - Cost per 1M requests, GPT-5.6 Sol vs Luna: 18 x cheaper ($3,600 → $200) At 400 input + 100 output tokens per request, list price, no caching. Source: https://developers.openai.com/api/docs/pricing - Best open-weight GPQA Diamond in this dataset: 90.1 % (DeepSeek-V4-Pro, MIT licence (open-weight range 71.2–90.1)) 1.6T-parameter MoE with 49B active; within 4.2 points of Gemini 3.1 Pro. Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro - Bedrock Model Distillation claim: 75 % cheaper (and up to 500% faster) AWS states distilled models are up to 500% faster and up to 75% less expensive than the original, with under 2% accuracy loss on RAG-style use cases. Source: https://aws.amazon.com/bedrock/model-distillation/ KEY FINDINGS ------------ 1. The cheap tier is now good enough for roughly 85% of enterprise traffic Within a single current generation the small sibling retains 88–98% of its teacher's knowledge benchmark. GPT-5.6 Luna scores 87.0 GPQA Diamond against Sol's 92.4; GPT-5.4 mini scores 88.0 against 93.0; Gemini 2.5 Flash scores 82.8 against 2.5 Pro's 86.4; DeepSeek-V4-Flash scores 88.1 against V4-Pro's 90.1. Across a large parameter gap retention falls to 47–76% — DeepSeek's own R1 students range from 91.2% at 70B down to 47.3% at 1.5B, and Gemma 4 31B to E4B retains 69.5%. The practical rule that falls out of the numbers is to route the hardest 5–15% of traffic to a frontier tier and everything else down to a same-generation small tier, not to the smallest model available. Sources: https://openrouter.ai/openai/gpt-5.6-luna https://openrouter.ai/openai/gpt-5.6-sol https://arxiv.org/html/2507.06261v1/ https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash 2. Agentic coding is the one place the frontier tier still earns its price Knowledge benchmarks compress; long-horizon agent benchmarks do not. Claude Opus 5 scores 96.0 on SWE-bench Verified, Sonnet 5 85.2 (89% retention) and Haiku 4.5 73.3 (76% retention). Llama 4 Scout retains 92% of Maverick’s MMLU-Pro but only 76% of its LiveCodeBench. If your workload is a coding agent that must finish a multi-step task unsupervised, the retention curve is much steeper than the GPQA curve suggests and a cheap tier will show up as retries, not as wrong answers. Sources: https://datanorth.ai/news/claude-opus-5-by-anthropic https://www.morphllm.com/claude-benchmarks https://www.anthropic.com/news/claude-haiku-4-5 https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md 3. Open weights are the buyer’s only real negotiating position Four open-weight families now clear 71.2–90.1 GPQA Diamond with no per-token fee and no anti-distillation clause: DeepSeek V4-Pro/Flash (MIT, 90.1/88.1), Qwen3.8-27B (Apache 2.0, 89.2), Gemma 4 31B (Apache 2.0, 84.3 GPQA / 85.2 MMLU-Pro) and Mistral Small 4 (Apache 2.0, 71.2 GPQA / 78.0 MMLU-Pro). Having a credible self-host fallback is what makes an API price cut stick — and 2026 has been a year of price cuts. Sources: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro https://huggingface.co/Qwen/Qwen3.8-27B https://ai.google.dev/gemma/docs/core/model_card_4 https://openrouter.ai/mistralai/mistral-small-2603 4. You may not distil the models you buy — but you may distil the ones you download Anthropic’s Commercial Terms section D.4 bars customers from accessing the Services "to build a competing product or service, including to train competing AI models". OpenAI’s Services Agreement carries an equivalent restriction. Both apply to enterprise accounts. By contrast DeepSeek R1 and V4 ship under MIT, and the R1 release explicitly encourages distillation; Qwen3.8, Gemma 4, Mistral Small 4 and Ministral 3 are Apache 2.0. If your roadmap includes training a small in-house model on model outputs, the licence question decides your teacher before any benchmark does. Sources: https://www.anthropic.com/legal/commercial-terms https://huggingface.co/deepseek-ai/DeepSeek-R1 https://ai.google.dev/gemma/docs/core/model_card_4 5. Distillation is now a purchasable service, not just a research technique Amazon Bedrock Model Distillation takes only your prompts: it generates synthetic teacher responses and fine-tunes the student. Nova Premier, Claude 3.5 Sonnet v2 and Llama 3.3 70B are supported teachers; Amazon Nova Pro and Llama 3.2 1B/3B are supported students. AWS claims up to 500% faster and up to 75% less expensive inference than the original models, with less than 2% accuracy loss for use cases like RAG. For most buyers this is a cheaper path to a task-specific small model than running a distillation pipeline in-house. Sources: https://aws.amazon.com/bedrock/model-distillation/ https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html 6. Distillation transfers narrow skills far better than broad knowledge The DeepSeek R1 student series is the cleanest natural experiment available. R1-Distill-Qwen-1.5B keeps 86% of R1’s MATH-500 (83.9 vs 97.3-class teacher performance) but only 47% of its GPQA Diamond (33.8 vs 71.5). At 7B the split is still stark: MATH-500 92.8 but GPQA 49.1. Buyers evaluating a tiny distilled model on a maths or format-following eval will systematically overestimate how it behaves on open-domain knowledge work. Sources: https://huggingface.co/deepseek-ai/DeepSeek-R1 https://huggingface.co/deepseek-ai/DeepSeek-R1-0528 7. Advertised latency and measured latency have diverged because of adaptive thinking Vendor tables still say "fastest", but the number a buyer feels now depends on the reasoning effort setting. Artificial Analysis measures GPT-5.6 Luna at 1.70s time-to-first-token at low effort and 19.87s at high; Claude Haiku 4.5 with reasoning on measures 19.92s; Claude Sonnet 5 at max effort measures 177.77s. Gemini 2.5 Flash-Lite in non-reasoning mode measures 0.30s. Any latency SLA written against a reasoning model has to pin the effort level or it is not a specification. Sources: https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gpt-5-6-luna-high https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5 https://artificialanalysis.ai/models 8. Legacy tiers are the biggest silent line item in most 2026 AI bills GPT-4o still lists at $2.50/$10 — the same rate as in 2024 — while GPT-5.6 Luna lists at $0.20/$1.20 with a far larger context window. Claude Sonnet 4.6 costs 50% more than the newer, better Sonnet 5. Gemini 3.5 Flash at $1.50/$9.00 costs twice Gemini 3.8 Flash's $0.75/$3.75 — but that Flash rate is promotional through 2026-12-31 and lists at $1.50/$7.50 from 2027-01-01, i.e. the same input rate as 3.5 Flash. Nothing forces a migration, so pinned model IDs from 2024–25 quietly bill at up to 12x the current market rate for the same job. Sources: https://developers.openai.com/api/docs/pricing https://platform.claude.com/docs/en/about-claude/pricing https://ai.google.dev/gemini-api/docs/pricing 9. Price stability is not something you can assume any more In 2026 alone: OpenAI cut GPT-5.6 Luna 80% and Terra 20% on 30 July, then cut Sol over 20% on 21 August for a three-month window; Anthropic cancelled a scheduled Sonnet 5 increase from $2/$10 to $3/$15 and made the lower price permanent; DeepSeek raised V4 standard rates roughly 3–4.7x on 16 August — and cache-hit input rates by up to 11x — as demand strained capacity. Contract for the workload, not for the price — and keep a second vendor wired up. Sources: https://www.axios.com/2026/07/30/openai-cuts-prices-gpt-terra-luna5 https://platform.claude.com/docs/en/about-claude/pricing https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html 10. Context window is no longer a reason to pay frontier prices Every GPT-5.6 tier including the $0.20 Luna carries a 1.05M-token window. Claude Sonnet 5 and Opus 5 both carry 1M. Gemini 3.5 Flash-Lite carries 1M at $0.30/MTok. Llama 4 Scout carries 10M with open weights. Two caveats matter for whole-repository or whole-contract passes: Claude Haiku 4.5 is still capped at 200K, and long prompts are repriced above 200K input tokens — Gemini 3.1 Pro and 2.5 Pro double the input rate above that threshold (and GPT-5.4-class models apply 2x input / 1.5x output), so the cheap headline rate is not the rate you pay on a million-token prompt. Sources: https://developers.openai.com/api/docs/models https://platform.claude.com/docs/en/about-claude/models/overview https://ai.google.dev/gemini-api/docs/pricing https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md TABLES -------- TABLE: Master comparison: 74 models a buyer could shortlist in September 2026 [id: model-master, 74 rows] Price, quality, context and licence for frontier teachers, vendor small siblings and openly-licensed distilled students. "Quality (GPQA-D)" is GPQA Diamond; MMLU-Pro is shown where the vendor publishes it. Blank cells mean the figure is not published — nothing here is estimated. Model | Vendor | Role | Distilled? | Teacher | Input (USD/MTok) | Output (USD/MTok) | GPQA-D (%) | MMLU-Pro (%) | Coding | Context (K tokens) | Licence | Source URL ----- | ------ | ---- | ---------- | ------- | ---------------- | ----------------- | ---------- | ------------ | ------ | ------------------ | ------- | ---------- GPT-6 Astra | OpenAI | Frontier teacher | No | n/a | 10 | 50 | n/a | n/a | n/a | 1050 | Proprietary API | https://developers.openai.com/api/docs/pricing GPT-5.6 Sol | OpenAI | Frontier teacher | No | n/a | 4 | 20 | 92.4 | n/a | n/a | 1050 | Proprietary API | https://openrouter.ai/openai/gpt-5.6-sol GPT-5.6 Terra | OpenAI | Small sibling | Undisclosed | undisclosed | 2 | 12 | 88.4 | n/a | n/a | 1050 | Proprietary API | https://openrouter.ai/openai/gpt-5.6-terra GPT-5.6 Luna | OpenAI | Small sibling | Undisclosed | undisclosed | 0.2 | 1.2 | 87 | n/a | n/a | 1050 | Proprietary API | https://openrouter.ai/openai/gpt-5.6-luna GPT-5.5 | OpenAI | Frontier teacher | No | n/a | 5 | 30 | n/a | n/a | n/a | n/a | Proprietary API | https://developers.openai.com/api/docs/pricing GPT-5.4 | OpenAI | Frontier teacher | No | n/a | 2.5 | 15 | 93 | n/a | 57.7 (SWE-bench Pro) | 1050 | Proprietary API | https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ GPT-5.4 mini | OpenAI | Small sibling | Undisclosed | undisclosed | 0.75 | 4.5 | 88 | n/a | 54.4 (SWE-bench Pro) | 400 | Proprietary API | https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ GPT-5.4 nano | OpenAI | Small sibling | Undisclosed | undisclosed | 0.2 | 1.25 | 82.8 | n/a | 52.4 (SWE-bench Pro) | 400 | Proprietary API | https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ GPT-5 | OpenAI | Frontier teacher | No | n/a | 1.25 | 10 | n/a | n/a | 74.9 (SWE-bench Verified) | 400 | Proprietary API | https://arxiv.org/pdf/2601.03267 GPT-5 mini | OpenAI | Small sibling | Undisclosed | undisclosed | 0.25 | 2 | 80.3 | n/a | 45.7 (SWE-bench Pro) | 400 | Proprietary API | https://openrouter.ai/openai/gpt-5-mini GPT-5 nano | OpenAI | Small sibling | Undisclosed | undisclosed | 0.05 | 0.4 | 70.9 | n/a | n/a | 400 | Proprietary API | https://openrouter.ai/openai/gpt-5-nano GPT-4o | OpenAI | Frontier teacher | No | n/a | 2.5 | 10 | n/a | n/a | n/a | 128 | Proprietary API | https://openrouter.ai/openai/gpt-4o GPT-4o mini | OpenAI | Small sibling | Undisclosed | undisclosed | 0.15 | 0.6 | n/a | n/a | 87.2 (HumanEval) | 128 | Proprietary API | https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ o3 | OpenAI | Frontier teacher | No | n/a | 2 | 8 | 83.3 | n/a | 69.1 (SWE-bench Verified) | 200 | Proprietary API | https://www.datacamp.com/blog/o4-mini o4-mini | OpenAI | Small sibling | Undisclosed | undisclosed | 1.1 | 4.4 | 81.4 | n/a | 68.1 (SWE-bench Verified) | 200 | Proprietary API | https://www.datacamp.com/blog/o4-mini gpt-oss-120b | OpenAI (open weights) | Open weights | Undisclosed | undisclosed | 0.15 | 0.6 | n/a | n/a | n/a | 128 | Apache 2.0 | https://huggingface.co/openai/gpt-oss-20b gpt-oss-20b | OpenAI (open weights) | Open weights | Undisclosed | undisclosed | 0.075 | 0.3 | 58.59 | n/a | 53.2 (SWE-bench Verified) | 128 | Apache 2.0 | https://huggingface.co/openai/gpt-oss-20b Claude Fable 5.1 | Anthropic | Frontier teacher | No | n/a | 10 | 50 | n/a | n/a | n/a | 1000 | Proprietary API | https://platform.claude.com/docs/en/about-claude/models/overview Claude Opus 5 | Anthropic | Frontier teacher | No | n/a | 5 | 25 | n/a | n/a | 96 (SWE-bench Verified) | 1000 | Proprietary API | https://datanorth.ai/news/claude-opus-5-by-anthropic Claude Sonnet 5 | Anthropic | Small sibling | Undisclosed | undisclosed | 2 | 10 | n/a | n/a | 85.2 (SWE-bench Verified) | 1000 | Proprietary API | https://www.morphllm.com/claude-benchmarks Claude Sonnet 4.6 | Anthropic | Small sibling | Undisclosed | undisclosed | 3 | 15 | n/a | n/a | n/a | 1000 | Proprietary API | https://platform.claude.com/docs/en/about-claude/models/overview Claude Haiku 4.5 | Anthropic | Small sibling | Undisclosed | undisclosed | 1 | 5 | n/a | n/a | 73.3 (SWE-bench Verified) | 200 | Proprietary API | https://www.anthropic.com/news/claude-haiku-4-5 Claude Haiku 3.5 | Anthropic | Small sibling | Undisclosed | undisclosed | 0.8 | 4 | n/a | n/a | n/a | 200 | Proprietary API | https://platform.claude.com/docs/en/about-claude/pricing Gemini 3.1 Pro (Preview) | Google | Frontier teacher | No | n/a | 2 | 12 | 94.3 | 92.6 | 80.6 (SWE-bench Verified) | 1000 | Proprietary API | https://deepmind.google/models/gemini/pro/ Gemini 3.8 Flash | Google | Small sibling | Undisclosed | undisclosed | 0.75 | 3.75 | n/a | n/a | 73.7 (DeepSWE v1.1) | 1000 | Proprietary API | https://deepmind.google/models/gemini/flash/ Gemini 3.5 Flash | Google | Small sibling | Undisclosed | undisclosed | 1.5 | 9 | n/a | n/a | n/a | 1000 | Proprietary API | https://ai.google.dev/gemini-api/docs/models Gemini 3.5 Flash-Lite | Google | Small sibling | Undisclosed | undisclosed | 0.3 | 2.5 | n/a | n/a | n/a | 1000 | Proprietary API | https://ai.google.dev/gemini-api/docs/models Gemini 3.1 Flash-Lite | Google | Small sibling | Undisclosed | undisclosed | 0.25 | 1.5 | 72.2 | 83 | n/a | 1000 | Proprietary API | https://layerlens.ai/blog/gemini-3-1-flash-lite-benchmark-results-efficiency-model-comparison Gemini 2.5 Pro | Google | Frontier teacher | No | n/a | 1.25 | 10 | 86.4 | n/a | 67.2 (SWE-bench Verified) | 1000 | Proprietary API | https://arxiv.org/html/2507.06261v1/ Gemini 2.5 Flash | Google | Documented distillation | Yes (documented) | Gemini 2.5 Pro (k-sparse logit distillation) | 0.3 | 2.5 | 82.8 | n/a | 60.3 (SWE-bench Verified) | 1000 | Proprietary API | https://arxiv.org/html/2507.06261v1/ Gemini 2.5 Flash-Lite | Google | Documented distillation | Yes (documented) | Gemini 2.5 Pro (k-sparse logit distillation) | 0.1 | 0.4 | n/a | n/a | n/a | 1000 | Proprietary API | https://arxiv.org/html/2507.06261v1/ Gemma 4 31B | Google | Open weights | Undisclosed | undisclosed | n/a | n/a | 84.3 | 85.2 | 80 (LiveCodeBench v6) | 256 | Apache 2.0 | https://ai.google.dev/gemma/docs/core/model_card_4 Gemma 4 26B A4B (MoE) | Google | Open weights | Undisclosed | undisclosed | n/a | n/a | 82.3 | 82.6 | 77.1 (LiveCodeBench v6) | 256 | Apache 2.0 | https://ai.google.dev/gemma/docs/core/model_card_4 Gemma 4 12B Unified | Google | Open weights | Undisclosed | undisclosed | n/a | n/a | 78.8 | 77.2 | 72 (LiveCodeBench v6) | 256 | Apache 2.0 | https://ai.google.dev/gemma/docs/core/model_card_4 Gemma 4 E4B | Google | Open weights | Undisclosed | undisclosed | n/a | n/a | 58.6 | 69.4 | 52 (LiveCodeBench v6) | 128 | Apache 2.0 | https://ai.google.dev/gemma/docs/core/model_card_4 Gemma 4 E2B | Google | Open weights | Undisclosed | undisclosed | n/a | n/a | 43.4 | 60 | 44 (LiveCodeBench v6) | 128 | Apache 2.0 | https://ai.google.dev/gemma/docs/core/model_card_4 Gemma 3 27B IT | Google | Open weights | Undisclosed | undisclosed | n/a | n/a | 24.3 | n/a | 48.8 (HumanEval) | 128 | Gemma Terms of Use (custom) | https://huggingface.co/google/gemma-3-27b-it Gemma 3 4B IT | Google | Open weights | Undisclosed | undisclosed | n/a | n/a | 15 | n/a | 36 (HumanEval) | 128 | Gemma Terms of Use (custom) | https://huggingface.co/google/gemma-3-27b-it Llama 4 Maverick | Meta | Documented distillation | Yes (documented) | Llama 4 Behemoth (codistillation) | n/a | n/a | 69.8 | 80.5 | 43.4 (LiveCodeBench) | 1000 | Llama 4 Community License (custom commercial) | https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md Llama 4 Scout | Meta | Documented distillation | Yes (documented) | Llama 4 Behemoth (codistillation) | n/a | n/a | 57.2 | 74.3 | 32.8 (LiveCodeBench) | 10000 | Llama 4 Community License (custom commercial) | https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md Llama 3.3 70B Instruct | Meta | Open weights | Undisclosed | undisclosed | 1.04 | 1.04 | 50.5 | 68.9 | 88.4 (HumanEval) | 128 | Llama 3.3 Community License (custom commercial) | https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md Llama 3.2 3B Instruct | Meta | Documented distillation | Yes (documented) | Llama 3.1 8B and 70B (token-level logit distillation after pruning) | n/a | n/a | 32.8 | n/a | n/a | 128 | Llama 3.2 Community License (custom commercial) | https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md Llama 3.2 1B Instruct | Meta | Documented distillation | Yes (documented) | Llama 3.1 8B and 70B (token-level logit distillation after pruning) | n/a | n/a | 27.2 | n/a | n/a | 128 | Llama 3.2 Community License (custom commercial) | https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md DeepSeek-V4-Pro | DeepSeek | Frontier teacher | No | n/a | 1.32 | 3.96 | 90.1 | 87.5 | 80.6 (SWE-bench Verified) | 1000 | MIT | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro DeepSeek-V4-Flash | DeepSeek | Documented distillation | Yes (documented) | DeepSeek V4 domain experts (on-policy distillation consolidation) | 0.44 | 1.32 | 88.1 | 86.2 | 79 (SWE-bench Verified) | 1000 | MIT | https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash DeepSeek-V3.2 | DeepSeek | Open weights | Yes (documented) | DeepSeek specialist models (specialist distillation into the generalist) | n/a | n/a | 82.4 | 85 | 73.1 (SWE-bench Verified) | 128 | MIT | https://arxiv.org/html/2512.02556 DeepSeek-R1 (0528) | DeepSeek | Frontier teacher | No | n/a | n/a | n/a | 81 | 85 | 73.3 (LiveCodeBench) | 128 | MIT | https://huggingface.co/deepseek-ai/DeepSeek-R1-0528 DeepSeek-R1-Distill-Llama-70B | DeepSeek / Meta base | Documented distillation | Yes (documented) | DeepSeek-R1 | n/a | n/a | 65.2 | n/a | 57.5 (LiveCodeBench) | 128 | MIT (weights) over Llama 3.3 Community License base | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1-Distill-Qwen-32B | DeepSeek / Qwen base | Documented distillation | Yes (documented) | DeepSeek-R1 | n/a | n/a | 62.1 | n/a | 57.2 (LiveCodeBench) | 128 | MIT (weights), Qwen2.5-32B base under Apache 2.0 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1-Distill-Qwen-14B | DeepSeek / Qwen base | Documented distillation | Yes (documented) | DeepSeek-R1 | n/a | n/a | 59.1 | n/a | 53.1 (LiveCodeBench) | 128 | MIT (weights), Qwen2.5-14B base under Apache 2.0 | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1-Distill-Llama-8B | DeepSeek / Meta base | Documented distillation | Yes (documented) | DeepSeek-R1 | n/a | n/a | 49 | n/a | 39.6 (LiveCodeBench) | 128 | MIT (weights) over Llama 3.1 Community License base | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1-Distill-Qwen-7B | DeepSeek / Qwen base | Documented distillation | Yes (documented) | DeepSeek-R1 | n/a | n/a | 49.1 | n/a | 37.6 (LiveCodeBench) | 128 | MIT (weights), Qwen2.5-Math-7B base under Apache 2.0 | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1-Distill-Qwen-1.5B | DeepSeek / Qwen base | Documented distillation | Yes (documented) | DeepSeek-R1 | n/a | n/a | 33.8 | n/a | 16.9 (LiveCodeBench) | 128 | MIT (weights), Qwen2.5-Math-1.5B base under Apache 2.0 | https://huggingface.co/deepseek-ai/DeepSeek-R1 Qwen3.8-27B | Alibaba | Open weights | Undisclosed | undisclosed | 0.5 | 3 | 89.2 | n/a | 61.7 (SWE-bench Pro) | 262 | Apache 2.0 | https://huggingface.co/Qwen/Qwen3.8-27B qwen3.8-max | Alibaba | Frontier teacher | No | n/a | 2 | 6 | n/a | n/a | n/a | n/a | Proprietary API | https://www.alibabacloud.com/help/en/model-studio/model-pricing qwen3.8-flash | Alibaba | Small sibling | Undisclosed | undisclosed | 0.15 | 0.47 | n/a | n/a | n/a | n/a | Proprietary API | https://www.alibabacloud.com/help/en/model-studio/model-pricing qwen-turbo | Alibaba | Small sibling | Undisclosed | undisclosed | 0.05 | 0.2 | n/a | n/a | n/a | n/a | Proprietary API | https://www.alibabacloud.com/help/en/model-studio/model-pricing Qwen3-4B-Instruct-2507 | Alibaba | Documented distillation | Yes (documented) | Qwen3-32B / Qwen3-235B-A22B (off-policy + on-policy strong-to-weak distillation) | n/a | n/a | 62 | 69.6 | 35.1 (LiveCodeBench v6) | 262 | Apache 2.0 | https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507 qwen3-8b (hosted) | Alibaba | Documented distillation | Yes (documented) | Qwen3-32B / Qwen3-235B-A22B (strong-to-weak distillation) | 0.18 | 0.7 | n/a | n/a | n/a | 128 | Apache 2.0 | https://arxiv.org/pdf/2505.09388 Phi-4 (14B) | Microsoft | Open weights | Undisclosed | undisclosed | n/a | n/a | 56.1 | 70.4 | 82.6 (HumanEval) | 16 | MIT | https://huggingface.co/microsoft/phi-4 Phi-4-mini-instruct (3.8B) | Microsoft | Open weights | Undisclosed | undisclosed | n/a | n/a | 25.2 | 52.8 | n/a | 128 | MIT | https://huggingface.co/microsoft/Phi-4-mini-instruct Mistral Medium 3.5 | Mistral AI | Frontier teacher | No | n/a | 1.5 | 7.5 | n/a | n/a | n/a | n/a | Modified MIT | https://docs.mistral.ai/getting-started/models/models_overview/ Mistral Large 3 | Mistral AI | Open weights | Undisclosed | undisclosed | 0.5 | 1.5 | n/a | n/a | n/a | n/a | Apache 2.0 | https://docs.mistral.ai/getting-started/models/models_overview/ Mistral Small 4 | Mistral AI | Small sibling | Undisclosed | undisclosed | 0.15 | 0.6 | 71.2 | 78 | n/a | n/a | Apache 2.0 | https://openrouter.ai/mistralai/mistral-small-2603 Ministral 3 14B Instruct | Mistral AI | Open weights | Undisclosed | undisclosed | 0.2 | 0.2 | 71.2 | n/a | 64.6 (LiveCodeBench) | 256 | Apache 2.0 | https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512 Ministral 3 8B | Mistral AI | Open weights | Undisclosed | undisclosed | 0.15 | 0.15 | n/a | n/a | n/a | 256 | Apache 2.0 | https://docs.mistral.ai/getting-started/models/models_overview/ Ministral 3 3B | Mistral AI | Open weights | Undisclosed | undisclosed | 0.1 | 0.1 | n/a | n/a | n/a | 256 | Apache 2.0 | https://docs.mistral.ai/getting-started/models/models_overview/ Amazon Nova Premier | Amazon | Frontier teacher | No | n/a | n/a | n/a | n/a | n/a | n/a | 1000 | Proprietary API | https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html Amazon Nova Pro | Amazon | Frontier teacher | Undisclosed (Bedrock distillation student) | undisclosed | 0.8 | 3.2 | 46.9 | n/a | n/a | 300 | Proprietary API | https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf Amazon Nova Lite | Amazon | Small tier | Undisclosed (Bedrock distillation student) | undisclosed | 0.06 | 0.24 | 42 | n/a | n/a | 300 | Proprietary API | https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf Amazon Nova Micro | Amazon | Small tier | Undisclosed (Bedrock distillation student) | undisclosed | 0.035 | 0.14 | 40 | n/a | n/a | 128 | Proprietary API | https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf Amazon Nova 2 Lite | Amazon | Small sibling | Undisclosed | undisclosed | 0.3 | 2.5 | n/a | n/a | n/a | 1000 | Proprietary API | https://docs.aws.amazon.com/nova/latest/nova2-userguide/whats-new.html SmolLM3-3B | Hugging Face | Open weights | Undisclosed | n/a | n/a | n/a | 41.7 | n/a | 30.48 (HumanEval+) | 128 | Apache 2.0 | https://huggingface.co/HuggingFaceTB/SmolLM3-3B grok-4.6 | xAI | Frontier teacher | No | n/a | 2 | 6 | n/a | n/a | n/a | 200 | Proprietary API | https://docs.x.ai/docs/models Notes: Prices are vendor list rates in USD per million tokens, before batch (typically -50%) or cache discounts. Open-weight rows with no price are self-host only in this dataset. DeepSeek prices are peak-hour; off-peak is half. Gemini 3.1 Pro and 2.5 Pro input prices double above 200K input tokens. Gemini 3.8 Flash's $0.75/$3.75 is promotional through 2026-12-31; it lists at $1.50/$7.50 from 2027-01-01. Qwen3.8-27B is shown at Alibaba Model Studio's list rate; Groq hosts the same weights at $0.80/$4.00. Gemini 3.7 Flash shipped in August 2026 but is not included in this shortlist. Sources: https://developers.openai.com/api/docs/pricing https://platform.claude.com/docs/en/about-claude/pricing https://ai.google.dev/gemini-api/docs/pricing https://api-docs.deepseek.com/quick_start/pricing/ https://mistral.ai/pricing/api https://www.alibabacloud.com/help/en/model-studio/model-pricing https://www.together.ai/pricing https://console.groq.com/docs/models https://ai.google.dev/gemma/docs/core/model_card_4 https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md https://huggingface.co/deepseek-ai/DeepSeek-R1 https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf TABLE: Price vs quality: GPQA Diamond points per dollar of blended token price [id: price-vs-quality, 24 rows] Blended price uses the 3:1 input:output weighting (0.75 x input + 0.25 x output). "GPQA per $" is the crude but decisive ranking metric for knowledge-heavy workloads. Cost per 1M requests assumes 400 input + 100 output tokens per request at list price with no caching. Model | Vendor | Tier | GPQA Diamond (%) | Blended price (USD/MTok) | GPQA points per $ (pts/USD) | Cost / 1M requests (USD) | Source URL ----- | ------ | ---- | ---------------- | ------------------------ | ------------------------ | ------------------------ | ---------- Amazon Nova Micro | Amazon | Small tier | 40 | 0.0613 | 652.5 | 28 | https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf GPT-5 nano | OpenAI | Small sibling | 70.9 | 0.1375 | 515.6 | 60 | https://openrouter.ai/openai/gpt-5-nano gpt-oss-20b | OpenAI (open weights) | Open weights | 58.59 | 0.1312 | 446.6 | 60 | https://huggingface.co/openai/gpt-oss-20b Amazon Nova Lite | Amazon | Small tier | 42 | 0.105 | 400 | 48 | https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf Ministral 3 14B Instruct | Mistral AI | Open weights | 71.2 | 0.2 | 356 | 100 | https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512 Mistral Small 4 | Mistral AI | Small sibling | 71.2 | 0.2625 | 271.2 | 120 | https://openrouter.ai/mistralai/mistral-small-2603 GPT-5.6 Luna | OpenAI | Small sibling | 87 | 0.45 | 193.3 | 200 | https://openrouter.ai/openai/gpt-5.6-luna GPT-5.4 nano | OpenAI | Small sibling | 82.8 | 0.4625 | 179 | 205 | https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ Gemini 3.1 Flash-Lite | Google | Small sibling | 72.2 | 0.5625 | 128.4 | 250 | https://layerlens.ai/blog/gemini-3-1-flash-lite-benchmark-results-efficiency-model-comparison DeepSeek-V4-Flash | DeepSeek | Distilled | 88.1 | 0.66 | 133.5 | 308 | https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash GPT-5 mini | OpenAI | Small sibling | 80.3 | 0.6875 | 116.8 | 300 | https://openrouter.ai/openai/gpt-5-mini Gemini 2.5 Flash | Google | Distilled | 82.8 | 0.85 | 97.4 | 370 | https://arxiv.org/html/2507.06261v1/ Qwen3.8-27B | Alibaba | Open weights | 89.2 | 1.125 | 79.3 | 500 | https://huggingface.co/Qwen/Qwen3.8-27B GPT-5.4 mini | OpenAI | Small sibling | 88 | 1.6875 | 52.1 | 750 | https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ Llama 3.3 70B Instruct | Meta | Open weights | 50.5 | 1.04 | 48.6 | 520 | https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md DeepSeek-V4-Pro | DeepSeek | Teacher | 90.1 | 1.98 | 45.5 | 924 | https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro o4-mini | OpenAI | Small sibling | 81.4 | 1.925 | 42.3 | 880 | https://openrouter.ai/openai/o4-mini Amazon Nova Pro | Amazon | Teacher | 46.9 | 1.4 | 33.5 | 640 | https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf o3 | OpenAI | Teacher | 83.3 | 3.5 | 23.8 | 1600 | https://www.datacamp.com/blog/o4-mini Gemini 2.5 Pro | Google | Teacher | 86.4 | 3.4375 | 25.1 | 1500 | https://arxiv.org/html/2507.06261v1/ Gemini 3.1 Pro (Preview) | Google | Teacher | 94.3 | 4.5 | 21 | 2000 | https://deepmind.google/models/gemini/pro/ GPT-5.6 Terra | OpenAI | Small sibling | 88.4 | 4.5 | 19.6 | 2000 | https://openrouter.ai/openai/gpt-5.6-terra GPT-5.4 | OpenAI | Teacher | 93 | 5.625 | 16.5 | 2500 | https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ GPT-5.6 Sol | OpenAI | Teacher | 92.4 | 8 | 11.6 | 3600 | https://openrouter.ai/openai/gpt-5.6-sol Notes: GPT-5 nano tops this ranking on raw efficiency but scores only 70.9 GPQA; GPT-5.6 Luna is the highest-scoring model in the top three, which is why it dominates most 2026 routing configurations. Amazon Nova figures use MMLU-era GPQA methodology from the Nova technical report and are not directly comparable with the 2026 reasoning-model scores. Sources: https://developers.openai.com/api/docs/pricing https://openrouter.ai/openai/gpt-5.6-luna https://api-docs.deepseek.com/quick_start/pricing/ https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash https://ai.google.dev/gemini-api/docs/pricing https://arxiv.org/html/2507.06261v1/ https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf https://console.groq.com/docs/models TABLE: Measured latency and throughput (Artificial Analysis, September 2026) [id: latency, 12 rows] Independently measured time-to-first-token and output speed. On adaptive-thinking models TTFT includes reasoning time, so the effort setting is stated for every row — without it these numbers are not comparable. Model (setting) | Time to first token (s) | Output speed (tokens/s) | AA Intelligence Index | Blended price (USD/MTok) | Source URL --------------- | ----------------------- | ----------------------- | --------------------- | ------------------------ | ---------- Gemini 2.5 Flash-Lite (non-reasoning) | 0.3 | n/a | n/a | 0.18 | https://artificialanalysis.ai/models gpt-oss-120b (high) | 0.85 | 151 | 24 | 0.2 | https://artificialanalysis.ai/models/comparisons/gpt-oss-120b-vs-llama-4-maverick Llama 4 Maverick | 0.92 | 82 | 14 | 0.31 | https://artificialanalysis.ai/models/comparisons/gpt-oss-120b-vs-llama-4-maverick DeepSeek V4-Flash 0731 (reasoning, max) | 1.19 | 140 | 52 | 0.23 | https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-deepseek-v4-pro GPT-5.6 Luna (low) | 1.7 | 109 | 34 | 0.17 | https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning DeepSeek V4-Pro 0813 (reasoning, max) | 1.9 | 60.2 | 53 | 0.69 | https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-deepseek-v4-pro Gemini 3.5 Flash-Lite | 6.48 | 391 | 37 | 0.33 | https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gpt-5-6-luna-high GPT-5.6 Luna (high) | 19.87 | 123 | 47 | 0.17 | https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gpt-5-6-luna-high Claude Haiku 4.5 (reasoning) | 19.92 | 90 | 30 | 0.77 | https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning Claude Opus 5 (adaptive, max effort) | 77.25 | 57 | 63 | 3.85 | https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5 Claude Sonnet 5 (adaptive, max effort) | 177.77 | 78 | 55 | 1.54 | https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5 Amazon Nova 2 Lite (non-reasoning) | n/a | 149 | 12 | 0.85 | https://pricepertoken.com/pricing-page/model/amazon-nova-2-lite-v1 Notes: Gemini 3.5 Flash-Lite is the throughput leader at 391 tokens/s. GPT-5.6 Luna is the only model here that spans both ends of the latency range purely through its effort parameter (1.70s to 19.87s), which makes it unusually easy to run interactive and batch traffic on one model ID. Claude Sonnet 5’s 177.77s figure is max-effort adaptive thinking, not a typical production setting. Sources: https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gpt-5-6-luna-high https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5 https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-deepseek-v4-pro https://artificialanalysis.ai/models/comparisons/gpt-oss-120b-vs-llama-4-maverick https://artificialanalysis.ai/models TABLE: Licence and deployment: what you are allowed to do with each family [id: license-deployment, 14 rows] The question that decides procurement before any benchmark does. "Distil outputs?" means: may you legally train your own model on this model’s outputs. Family | Licence | Weights | Self-host / air-gap | Distil outputs? | Data residency options | Buyer note | Source URL ------ | ------- | ------- | ------------------- | --------------- | ---------------------- | ---------- | ---------- OpenAI GPT-5.x / GPT-6 / o-series | Proprietary API | Closed | No | No — Services Agreement bars using Output to develop competing models | Azure/Foundry regions | Deepest price ladder in the market and 1.05M context on every 5.6 tier. | https://openai.com/policies/services-agreement/ OpenAI gpt-oss 20b / 120b | Apache 2.0 | Open | Yes (120b on one 80GB GPU; 20b in 16GB) | Yes | Anywhere | The escape hatch inside the OpenAI ecosystem. Served by Groq at ~500–1000 tok/s. | https://huggingface.co/openai/gpt-oss-20b Anthropic Claude (Fable / Opus / Sonnet / Haiku) | Proprietary API | Closed | No | No — Commercial Terms D.4 bars building a competing product or training competing AI models | inference_geo:"us" at a 1.1x multiplier; Bedrock/Vertex regional endpoints at +10% | Best published SWE-bench Verified (Opus 5, 96.0). 1M context on Fable/Opus/Sonnet, 200K on Haiku 4.5. | https://www.anthropic.com/legal/commercial-terms Google Gemini 2.5 / 3.x | Proprietary API | Closed | No | No — Gemini API Additional Terms restrict competitive model development | Vertex AI regions | Only closed vendor that has publicly documented distilling its own small tier (Gemini 2.5 report). | https://arxiv.org/html/2507.06261v1/ Google Gemma 4 | Apache 2.0 | Open | Yes (2.3B–31B) | Yes | Anywhere | MMLU-Pro 85.2 at 31B under Apache 2.0 — the strongest permissively-licensed model you can put on one node. | https://ai.google.dev/gemma/docs/core/model_card_4 Google Gemma 3 | Gemma Terms of Use (custom, use restrictions apply) | Open | Yes | Yes, subject to the Gemma prohibited-use policy | Anywhere | Superseded by Gemma 4 on both quality and licence terms. | https://huggingface.co/google/gemma-3-27b-it Meta Llama 3.x / 4 | Llama Community License (custom commercial; 700M MAU clause) | Open | Yes | Yes, with Llama attribution and naming obligations on derivatives | Anywhere | Llama 3.2 1B/3B are explicitly documented distillations; Meta's Llama 4 launch post describes Scout/Maverick as codistilled from the unreleased Behemoth (the model card itself does not mention it). | https://ai.meta.com/blog/llama-4-multimodal-intelligence/ DeepSeek R1 / V3.2 / V4 | MIT | Open | Yes (V4-Pro is a 1.6T MoE — non-trivial) | Yes — the R1 release shipped six distilled students itself | Anywhere; first-party API is PRC-hosted | Highest open-weight GPQA Diamond here (90.1). Many enterprises self-host rather than use the PRC-hosted API. | https://huggingface.co/deepseek-ai/DeepSeek-R1 Alibaba Qwen3 / Qwen3.8 (open weights) | Apache 2.0 | Open | Yes | Yes | Anywhere; Model Studio API is PRC/Singapore | Qwen3 technical report documents strong-to-weak distillation for the 0.6B–14B dense sizes and 30B-A3B. | https://huggingface.co/Qwen/Qwen3.8-27B Microsoft Phi-4 / Phi-4-mini | MIT | Open | Yes | Yes | Anywhere | Least legally encumbered small models in this dataset. Phi-4’s 16K context is the catch. | https://huggingface.co/microsoft/phi-4 Mistral Small 4 / Ministral 3 / Large 3 | Apache 2.0 | Open | Yes | Yes | EU-headquartered vendor; EU hosting available | The default answer when the requirement is EU sovereignty plus a permissive licence. | https://docs.mistral.ai/getting-started/models/models_overview/ Mistral Medium 3.5 | Modified MIT | Open | Yes | Yes, subject to the modified terms | EU hosting available | Frontier-class tier of the Mistral line; read the modification before assuming MIT. | https://docs.mistral.ai/getting-started/models/models_overview/ Amazon Nova / Nova 2 | Proprietary API | Closed | No | Only through Bedrock Model Distillation, into another Amazon-supported student | AWS regions incl. GovCloud (US-West) | The only vendor that documents its teacher/student graph in product docs: Premier → Pro/Lite/Micro, Pro → Lite/Micro. | https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html Hugging Face SmolLM3 | Apache 2.0 | Open (plus data mixture and training configs) | Yes | Yes | Anywhere | Fully reproducible supply chain — relevant where an auditor asks what the model was trained on. | https://huggingface.co/HuggingFaceTB/SmolLM3-3B Notes: Anti-distillation clauses bind the enterprise account holder, not just individual developers, and survive termination in most of these agreements. If your roadmap includes training an in-house model on teacher outputs, pick an MIT or Apache 2.0 teacher up front rather than seeking a waiver later. Sources: https://www.anthropic.com/legal/commercial-terms https://openai.com/policies/services-agreement/ https://huggingface.co/openai/gpt-oss-20b https://ai.google.dev/gemma/docs/core/model_card_4 https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md https://huggingface.co/deepseek-ai/DeepSeek-R1 https://huggingface.co/Qwen/Qwen3.8-27B https://huggingface.co/microsoft/phi-4 https://docs.mistral.ai/getting-started/models/models_overview/ https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html https://huggingface.co/HuggingFaceTB/SmolLM3-3B TABLE: Quality retention by pair (documented distillations and same-family size comparisons) [id: retention, 22 rows] What fraction of the teacher’s score the cheaper model keeps, and what fraction of the teacher’s input price it costs. Rows where the student is documented as a distillation are the ones with a named teacher in the master table; the rest are same-family price ladders. Teacher → student | Relationship | Benchmark | Teacher (%) | Student (%) | Retention (%) | Student price / teacher price (input) (x) | Source URL ----------------- | ------------ | --------- | ----------- | ----------- | ------------- | ------------------------ | ---------- GPT-5.6 Sol → Terra | Same-generation tier comparison (method undisclosed) | GPQA Diamond | 92.4 | 88.4 | 95.7 | 0.5 | https://openrouter.ai/openai/gpt-5.6-terra GPT-5.6 Sol → Luna | Same-generation tier comparison (method undisclosed); third-party benchmark, 94.2–97.3% across providers | GPQA Diamond | 92.4 | 87 | 94.2 | 0.05 | https://openrouter.ai/openai/gpt-5.6-luna GPT-5.4 → GPT-5.4 mini | Same-generation tier comparison (method undisclosed) | GPQA Diamond | 93 | 88 | 94.6 | 0.3 | https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ GPT-5.4 → GPT-5.4 nano | Same-generation tier comparison (method undisclosed) | GPQA Diamond | 93 | 82.8 | 89 | 0.08 | https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ o3 → o4-mini | Same-generation tier comparison (method undisclosed) | GPQA Diamond | 83.3 | 81.4 | 97.7 | 0.55 | https://www.datacamp.com/blog/o4-mini Gemini 2.5 Pro → 2.5 Flash | Documented distillation | GPQA Diamond | 86.4 | 82.8 | 95.8 | 0.24 | https://arxiv.org/html/2507.06261v1/ DeepSeek V4-Pro vs V4-Flash | Same-family price/quality comparison — V4-Flash is a consolidation of V4 domain experts via on-policy distillation, not a compression of V4-Pro | GPQA Diamond | 90.1 | 88.1 | 97.8 | 0.333 | https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash DeepSeek-R1 → R1-Distill-Llama-70B | Documented distillation | GPQA Diamond | 71.5 | 65.2 | 91.2 | n/a | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1 → R1-Distill-Qwen-32B | Documented distillation | GPQA Diamond | 71.5 | 62.1 | 86.9 | n/a | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1 → R1-Distill-Qwen-14B | Documented distillation | GPQA Diamond | 71.5 | 59.1 | 82.7 | n/a | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1 → R1-Distill-Llama-8B | Documented distillation | GPQA Diamond | 71.5 | 49 | 68.5 | n/a | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1 → R1-Distill-Qwen-7B | Documented distillation | GPQA Diamond | 71.5 | 49.1 | 68.7 | n/a | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1 → R1-Distill-Qwen-1.5B | Documented distillation | GPQA Diamond | 71.5 | 33.8 | 47.3 | n/a | https://huggingface.co/deepseek-ai/DeepSeek-R1 Gemma 4 31B vs Gemma 4 E4B | Same-family size comparison — Google does not describe E4B as a distillation of 31B | GPQA Diamond | 84.3 | 58.6 | 69.5 | n/a | https://ai.google.dev/gemma/docs/core/model_card_4 Llama 4 Maverick → Scout | Documented codistillation (both from Behemoth), compared here by size | MMLU-Pro | 80.5 | 74.3 | 92.3 | n/a | https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md Nova Pro vs Nova Lite | Same-family size comparison — Bedrock supports both as distillation students; shipped provenance undisclosed | MMLU | 85.9 | 80.5 | 93.7 | 0.075 | https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf Nova Pro vs Nova Micro | Same-family size comparison — Bedrock supports both as distillation students; shipped provenance undisclosed | MMLU | 85.9 | 77.6 | 90.3 | 0.044 | https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf Claude Opus 5 → Sonnet 5 | Same-generation tier comparison (method undisclosed) | SWE-bench Verified | 96 | 85.2 | 88.8 | 0.4 | https://www.morphllm.com/claude-benchmarks Claude Opus 5 → Haiku 4.5 | Cross-generation tier comparison (method undisclosed) | SWE-bench Verified | 96 | 73.3 | 76.4 | 0.2 | https://datanorth.ai/news/claude-opus-5-by-anthropic Llama 4 Maverick → Scout (coding) | Documented codistillation (both from Behemoth), compared here by size | LiveCodeBench | 43.4 | 32.8 | 75.6 | n/a | https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md DeepSeek V4-Pro vs V4-Flash (coding) | Same-family price/quality comparison — see row 6 | SWE-bench Verified | 80.6 | 79 | 98 | 0.333 | https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash Gemini 2.5 Pro → 2.5 Flash (coding) | Documented distillation | SWE-bench Verified | 67.2 | 60.3 | 89.7 | 0.24 | https://arxiv.org/html/2507.06261v1/ Notes: Two patterns stand out. First, knowledge retention above 94% is now routine inside a family, and DeepSeek V4-Flash reaches 97.8% for a third of the price. Second, retention falls off a cliff for very small students on broad-knowledge benchmarks — R1-Distill-Qwen-1.5B keeps only 47% of R1’s GPQA — while the same model keeps 86% of R1’s MATH-500. Distillation buys you narrow competence cheaply and broad competence expensively. Only rows marked "Documented distillation" carry a vendor statement that the student was trained from the teacher. The remaining rows are same-family or same-generation price/quality comparisons and should not be read as training provenance. Sources: https://openrouter.ai/openai/gpt-5.6-luna https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ https://arxiv.org/html/2507.06261v1/ https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash https://huggingface.co/deepseek-ai/DeepSeek-R1 https://ai.google.dev/gemma/docs/core/model_card_4 https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf https://www.morphllm.com/claude-benchmarks https://datanorth.ai/news/claude-opus-5-by-anthropic https://www.datacamp.com/blog/o4-mini CHART DATA ---------- CHART: Blended price vs MMLU-Pro [id: price-vs-mmlupro, type: scatter, unit: %] Blended price (USD per million tokens, 3:1 input:output) | Frontier teachers (%) 4.5 | 92.6 1.98 | 87.5 Blended price (USD per million tokens, 3:1 input:output) | Small siblings (%) 0.5625 | 83 0.2625 | 78 Blended price (USD per million tokens, 3:1 input:output) | Documented distillations (%) 0.66 | 86.2 Blended price (USD per million tokens, 3:1 input:output) | Open weights (hosted price) (%) 1.04 | 68.9 Notes: Only models that publish MMLU-Pro AND have a sourced token price appear here. The frontier is almost flat between $0.26 and $2.00: Mistral Small 4 buys 78.0 MMLU-Pro for $0.26 blended, DeepSeek-V4-Flash 86.2 for $0.66, DeepSeek-V4-Pro 87.5 for $1.98, and Gemini 3.1 Pro 92.6 for $4.50 — so the last 6 MMLU-Pro points cost roughly 7x. Gemma 4 31B (85.2) and Gemma 4 26B A4B (82.6) sit off this chart because they are self-host-only and therefore have no list token price. Sources: https://ai.google.dev/gemini-api/docs/pricing https://deepmind.google/models/gemini/pro/ https://layerlens.ai/blog/gemini-3-1-flash-lite-benchmark-results-efficiency-model-comparison https://api-docs.deepseek.com/quick_start/pricing/ https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash https://mistral.ai/pricing/api https://openrouter.ai/mistralai/mistral-small-2603 https://www.together.ai/pricing https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md CHART: Blended price vs GPQA Diamond — the whole shortlist on one plot [id: price-vs-gpqa, type: scatter, unit: %] Blended price (USD per million tokens, 3:1 input:output) | Frontier teachers (%) 8 | 92.4 5.625 | 93 3.5 | 83.3 4.5 | 94.3 3.4375 | 86.4 1.98 | 90.1 1.4 | 46.9 Blended price (USD per million tokens, 3:1 input:output) | Small siblings (method undisclosed) (%) 4.5 | 88.4 0.45 | 87 1.6875 | 88 0.4625 | 82.8 0.6875 | 80.3 0.1375 | 70.9 1.925 | 81.4 0.5625 | 72.2 0.2625 | 71.2 0.105 | 42 0.0613 | 40 Blended price (USD per million tokens, 3:1 input:output) | Documented distillations (%) 0.85 | 82.8 0.66 | 88.1 Blended price (USD per million tokens, 3:1 input:output) | Open weights (%) 0.1312 | 58.59 1.04 | 50.5 1.125 | 89.2 0.2 | 71.2 Notes: The upper-left corner is the whole story: GPT-5.6 Luna at $0.45 blended / 87.0 GPQA and DeepSeek-V4-Flash at $0.66 / 88.1 sit close to the frontier for a fraction of the price. Gemini 3.1 Flash-Lite is cheaper still at $0.56 blended but scores 72.2 GPQA Diamond, so it trades roughly 15 points of knowledge benchmark for the saving. Paying 8–18x more than Luna or V4-Flash moves you 5–7 GPQA points. Amazon Nova rows use the 2024 Nova technical report’s GPQA methodology and are not comparable with the reasoning-era scores. Sources: https://developers.openai.com/api/docs/pricing https://openrouter.ai/openai/gpt-5.6-sol https://openrouter.ai/openai/gpt-5.6-terra https://openrouter.ai/openai/gpt-5.6-luna https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ https://www.datacamp.com/blog/o4-mini https://ai.google.dev/gemini-api/docs/pricing https://arxiv.org/html/2507.06261v1/ https://layerlens.ai/blog/gemini-3-1-flash-lite-benchmark-results-efficiency-model-comparison https://api-docs.deepseek.com/quick_start/pricing/ https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash https://mistral.ai/pricing/api https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512 https://console.groq.com/docs/models https://huggingface.co/Qwen/Qwen3.8-27B https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf CHART: Larger vs smaller tier on GPQA Diamond (and MMLU for Nova) [id: teacher-vs-distilled-gpqa, type: bar, unit: %] Pair (documented distillations and same-family comparisons) | Larger / teacher tier (%) GPT-5.6 Sol → Luna | 92.4 GPT-5.4 → GPT-5.4 mini | 93 o3 → o4-mini | 83.3 Gemini 2.5 Pro → 2.5 Flash | 86.4 DeepSeek V4-Pro vs V4-Flash | 90.1 DeepSeek-R1 → R1-Distill-Llama-70B | 71.5 DeepSeek-R1 → R1-Distill-Qwen-32B | 71.5 DeepSeek-R1 → R1-Distill-Qwen-7B | 71.5 DeepSeek-R1 → R1-Distill-Qwen-1.5B | 71.5 Gemma 4 31B vs Gemma 4 E4B | 84.3 Nova Pro vs Nova Lite | 85.9 Pair (documented distillations and same-family comparisons) | Smaller tier (distilled or same-family) (%) GPT-5.6 Sol → Luna | 87 GPT-5.4 → GPT-5.4 mini | 88 o3 → o4-mini | 81.4 Gemini 2.5 Pro → 2.5 Flash | 82.8 DeepSeek V4-Pro vs V4-Flash | 88.1 DeepSeek-R1 → R1-Distill-Llama-70B | 65.2 DeepSeek-R1 → R1-Distill-Qwen-32B | 62.1 DeepSeek-R1 → R1-Distill-Qwen-7B | 49.1 DeepSeek-R1 → R1-Distill-Qwen-1.5B | 33.8 Gemma 4 31B vs Gemma 4 E4B | 58.6 Nova Pro vs Nova Lite | 80.5 Notes: Inside a single generation the bars are nearly the same height. Across a large parameter gap they are not: R1 → R1-Distill-Qwen-1.5B loses more than half the teacher's GPQA, and Gemma 4 31B vs E4B loses 30%. Only the DeepSeek-R1 and Gemini 2.5 pairs are vendor-documented distillations; DeepSeek V4-Pro vs V4-Flash, Gemma 4 and Nova are same-family comparisons — V4-Flash is documented as a consolidation of V4 domain experts via on-policy distillation, not a compression of V4-Pro. Nova uses MMLU rather than GPQA because that is what the Amazon technical report publishes. Sources: https://openrouter.ai/openai/gpt-5.6-luna https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ https://www.datacamp.com/blog/o4-mini https://arxiv.org/html/2507.06261v1/ https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash https://huggingface.co/deepseek-ai/DeepSeek-R1 https://ai.google.dev/gemma/docs/core/model_card_4 https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf CHART: Cost per 1,000,000 requests (400 input + 100 output tokens each) [id: cost-per-1m-requests, type: bar, unit: USD] Model | List price, no caching or batch discount (USD) Amazon Nova Micro | 28 qwen-turbo | 40 Amazon Nova Lite | 48 GPT-5 nano | 60 gpt-oss-20b | 60 Ministral 3 14B Instruct | 100 qwen3.8-flash | 107 GPT-4o mini | 120 Mistral Small 4 | 120 gpt-oss-120b | 120 GPT-5.6 Luna | 200 GPT-5 mini | 300 DeepSeek-V4-Flash | 308 Gemini 2.5 Flash | 370 Gemini 3.5 Flash-Lite | 370 Amazon Nova 2 Lite | 370 Llama 3.3 70B Instruct | 520 Gemini 3.8 Flash | 675 GPT-5.4 mini | 750 o4-mini | 880 Claude Haiku 4.5 | 900 DeepSeek-V4-Pro | 924 Claude Sonnet 5 | 1800 Gemini 3.1 Pro (Preview) | 2000 GPT-5.6 Terra | 2000 GPT-5.6 Sol | 3600 Claude Opus 5 | 4500 Claude Fable 5.1 | 9000 GPT-6 Astra | 9000 Notes: A classification or extraction workload at a million requests a month costs $28 on Amazon Nova Micro, $200 on GPT-5.6 Luna, $900 on Claude Haiku 4.5, $1,800 on Claude Sonnet 5 and $9,000 on the two most expensive tiers, Claude Fable 5.1 and GPT-6 Astra (both $10/$50) — a 320x spread across the same shortlist. Batch APIs cut these by 50% at both Anthropic and OpenAI, and prompt caching cuts the input component by up to 90% (97.5% on Claude Fable 5.1), which is usually a bigger lever than moving one tier down. Sources: https://developers.openai.com/api/docs/pricing https://platform.claude.com/docs/en/about-claude/pricing https://ai.google.dev/gemini-api/docs/pricing https://api-docs.deepseek.com/quick_start/pricing/ https://mistral.ai/pricing/api https://www.alibabacloud.com/help/en/model-studio/model-pricing https://console.groq.com/docs/models https://www.together.ai/pricing https://pricepertoken.com/pricing-page/model/amazon-nova-micro-v1 https://pricepertoken.com/pricing-page/model/amazon-nova-lite-v1 https://pricepertoken.com/pricing-page/model/amazon-nova-2-lite-v1 CHART: Quality retention: smaller tier score as a percentage of the larger tier [id: quality-retention, type: bar, unit: %] Pair (documented distillations and same-family comparisons) | Knowledge benchmarks (GPQA-D / MMLU / MMLU-Pro) (%) GPT-5.6 Sol → Terra | 95.7 GPT-5.6 Sol → Luna | 94.2 GPT-5.4 → GPT-5.4 mini | 94.6 GPT-5.4 → GPT-5.4 nano | 89 o3 → o4-mini | 97.7 Gemini 2.5 Pro → 2.5 Flash | 95.8 DeepSeek V4-Pro vs V4-Flash | 97.8 DeepSeek-R1 → R1-Distill-Llama-70B | 91.2 DeepSeek-R1 → R1-Distill-Qwen-32B | 86.9 DeepSeek-R1 → R1-Distill-Qwen-14B | 82.7 DeepSeek-R1 → R1-Distill-Llama-8B | 68.5 DeepSeek-R1 → R1-Distill-Qwen-7B | 68.7 DeepSeek-R1 → R1-Distill-Qwen-1.5B | 47.3 Gemma 4 31B vs Gemma 4 E4B | 69.5 Llama 4 Maverick → Scout | 92.3 Nova Pro vs Nova Lite | 93.7 Nova Pro vs Nova Micro | 90.3 Pair (documented distillations and same-family comparisons) | Coding / agentic benchmarks (%) Claude Opus 5 → Sonnet 5 | 88.8 Claude Opus 5 → Haiku 4.5 | 76.4 Llama 4 Maverick → Scout (coding) | 75.6 DeepSeek V4-Pro vs V4-Flash (coding) | 98 Gemini 2.5 Pro → 2.5 Flash (coding) | 89.7 Notes: Read the two series against each other. Within one generation, knowledge retention clusters at 94–98%. Coding and agentic retention is bimodal: DeepSeek V4-Flash keeps 98% of V4-Pro on SWE-bench Verified, but Claude Haiku 4.5 keeps only 76% of Opus 5 and Llama 4 Scout only 76% of Maverick on LiveCodeBench. Only the DeepSeek-R1, Gemini 2.5 and Llama 4 rows are vendor-documented distillations; DeepSeek V4-Pro vs V4-Flash, Gemma 4 and Nova are same-family size comparisons (V4-Flash is documented as a consolidation of V4 domain experts, not a compression of V4-Pro). That spread, not the price sheet, is what should decide whether a workload can be moved down a tier. Sources: https://huggingface.co/deepseek-ai/DeepSeek-R1 https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash https://arxiv.org/html/2507.06261v1/ https://www.morphllm.com/claude-benchmarks https://datanorth.ai/news/claude-opus-5-by-anthropic https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md https://ai.google.dev/gemma/docs/core/model_card_4 https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf https://openrouter.ai/openai/gpt-5.6-luna https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ CHART: Enterprise inference price index, 2026 [id: inference-price-index, type: line, unit: USD] Date | Silicon Data enterprise inference index (cited by Jefferies) (USD) 2026-05-31 | 2.04 2026-07-25 | 1.45 2026-08-07 | 1.17 Notes: A 43% fall in ten weeks. Jefferies attributes it to OpenAI cutting GPT-5.6 rates by up to 80%, Anthropic matching prior frontier performance at half the price with Claude Opus 5, and Chinese open-weight models such as DeepSeek V4-Flash-0731 competing at roughly $0.03 per task. The 7 August figure is the midpoint of the reported $1.16–$1.18 low. Sources: https://www.scmp.com/tech/tech-trends/article/3363549/enterprise-ai-costs-hit-2026-low-driven-price-wars-chinese-open-source-models-research https://www.axios.com/2026/07/30/openai-cuts-prices-gpt-terra-luna5 DETAIL RECORDS -------------- DECISION GUIDE (10 use cases) - Use case: High-volume classification, routing, tagging, PII redaction Recommendation: Amazon Nova Micro or GPT-5 nano; Ministral 3 3B if you must self-host Why: At 400 input + 100 output tokens a request, Nova Micro costs $28 per million requests and GPT-5 nano $60. Neither needs reasoning. Nova Micro is a documented Bedrock distillation student, so you can also distil it further on your own labels. Ministral 3 3B (Apache 2.0, 256K context) is the equivalent when the data cannot leave your network. Models: Amazon Nova Micro; GPT-5 nano; qwen-turbo; Ministral 3 3B; Gemma 4 E4B - Use case: Customer-facing chat and support deflection Recommendation: GPT-5.6 Luna at low effort, with Gemini 3.5 Flash-Lite as the latency fallback Why: Luna measures 1.70s time-to-first-token at low effort, holds 87.0 GPQA Diamond, and costs $200 per million requests. Gemini 3.5 Flash-Lite runs at 391 output tokens/s if perceived typing speed matters more than depth. Anthropic’s own worked example puts 10,000 support conversations on Claude Haiku 4.5 at about $37, which is the same order of magnitude — pick on latency and evals, not price. Models: GPT-5.6 Luna; Gemini 3.5 Flash-Lite; Claude Haiku 4.5; Amazon Nova 2 Lite - Use case: Retrieval-augmented Q&A over a large private corpus Recommendation: GPT-5.6 Luna or Gemini 3.1 Flash-Lite, with aggressive prompt caching Why: Both carry 1M+ context at $0.20–$0.25 per million input tokens, and RAG is input-heavy, so the input rate dominates. GPT-5.6 Luna is the stronger of the two on knowledge benchmarks (87.0 GPQA Diamond vs Gemini 3.1 Flash-Lite’s 72.2), so prefer Luna where answer quality over technical corpora matters and Flash-Lite where throughput and price dominate. AWS measures under 2% accuracy loss on RAG-style tasks from distillation, so this is the workload class where cheap tiers are safest. Cache reads at 10% of input beat any further tier reduction. Models: GPT-5.6 Luna; Gemini 3.1 Flash-Lite; DeepSeek-V4-Flash; Claude Haiku 4.5 - Use case: Autonomous coding agent that must finish multi-step tasks unsupervised Recommendation: Claude Opus 5; drop to Claude Sonnet 5 only after measuring your own retry rate Why: This is the one workload where the retention curve is genuinely steep: Opus 5 scores 96.0 on SWE-bench Verified, Sonnet 5 85.2 and Haiku 4.5 73.3. A 20% failure delta on a long-horizon task compounds into far more than 20% extra cost once retries and human review are counted. Gemini 3.1 Pro (80.6 SWE-bench Verified) and DeepSeek-V4-Pro (80.6, MIT) are the credible alternatives. Models: Claude Opus 5; Claude Sonnet 5; Gemini 3.1 Pro (Preview); DeepSeek-V4-Pro; GPT-5.6 Sol - Use case: Code completion and inline suggestions in an IDE Recommendation: Qwen3.8-27B on Groq, or gpt-oss-20b if you need the weights Why: Latency is the product here. Qwen3.8-27B posts LiveCodeBench v6 90.3 and SWE-bench Pro 61.7 under Apache 2.0 and is served at roughly 450 tokens/s; gpt-oss-20b runs at about 1000 tokens/s and fits in 16GB. Neither needs a frontier tier because the human reviews every suggestion. Models: Qwen3.8-27B; gpt-oss-20b; Ministral 3 14B Instruct; Mistral Small 4 - Use case: Regulated workload that cannot leave your infrastructure Recommendation: Gemma 4 31B or 26B A4B (Apache 2.0); Phi-4-mini (MIT) at the small end Why: Gemma 4 31B posts MMLU-Pro 85.2 and GPQA Diamond 84.3 under Apache 2.0 with a 256K window — self-hostable quality that did not exist a year ago. The 26B MoE gets 82.3 GPQA with 3.8B active parameters, so it serves cheaply. Phi-4-mini is MIT with a 128K window for the smallest footprint. Avoid Gemma 3 (custom licence) and Llama (700M MAU clause plus naming obligations) if legal review is the bottleneck. Models: Gemma 4 31B; Gemma 4 26B A4B (MoE); Phi-4-mini-instruct (3.8B); Mistral Small 4; SmolLM3-3B - Use case: EU data-sovereignty requirement Recommendation: Mistral Small 4 or Ministral 3 14B, both Apache 2.0 from an EU vendor Why: Mistral Small 4 posts MMLU-Pro 78.0 and GPQA Diamond 71.2 at $0.15/$0.60; Ministral 3 14B posts GPQA Diamond 71.2, AIME25 85.0 and MATH 90.4 with a 256K context at symmetric $0.20 pricing. If you must stay on a US hyperscaler, Anthropic’s inference_geo:"us" and Bedrock/Vertex regional endpoints carry a 10% premium. Models: Mistral Small 4; Ministral 3 14B Instruct; Mistral Medium 3.5; Gemma 4 12B Unified - Use case: On-device or offline assistant (laptop, handset, vehicle) Recommendation: Gemma 4 E4B, with Llama 3.2 3B as the proven-in-production alternative Why: Gemma 4 E4B runs at 4.5B effective parameters with MMLU-Pro 69.4 and a 128K window — roughly GPT-4o mini class, entirely offline, Apache 2.0. Llama 3.2 3B is the reference documented distillation (logits from Llama 3.1 8B and 70B as token-level targets after pruning) and has the widest edge-runtime support. Expect broad-knowledge accuracy to degrade much faster than format compliance at this size. Models: Gemma 4 E4B; Llama 3.2 3B Instruct; Ministral 3 3B; Phi-4-mini-instruct (3.8B); SmolLM3-3B - Use case: Maths, scientific reasoning and quantitative analysis at low cost Recommendation: DeepSeek-V4-Flash, or R1-Distill-Qwen-32B if the weights must be local Why: V4-Flash keeps 97.8% of V4-Pro’s GPQA Diamond (88.1 vs 90.1) at a third of the price under MIT. For self-hosting, R1-Distill-Qwen-32B fits on one 80GB GPU and beat o1-mini on AIME 2024, MATH-500 and GPQA Diamond. Gemma 4 31B’s AIME 2026 score of 89.2 is the Apache 2.0 alternative if MIT-from-a-PRC-vendor is a procurement problem. Models: DeepSeek-V4-Flash; DeepSeek-R1-Distill-Qwen-32B; Gemma 4 31B; Ministral 3 14B Instruct - Use case: You want to build your own distilled model on your own task data Recommendation: Amazon Bedrock Model Distillation for a managed path; DeepSeek V4 or Qwen3.8 as the teacher if you build it yourself Why: Bedrock takes only your prompts, generates teacher responses and fine-tunes the student, with Nova Premier/Pro, Claude and Llama 3.3 70B / 3.2 1B / 3.2 3B in the supported graph; AWS claims up to 500% faster and 75% cheaper inference with under 2% accuracy loss on RAG. If you build the pipeline yourself, the teacher must be one whose licence permits it: OpenAI’s Services Agreement and Anthropic’s Commercial Terms D.4 both prohibit training competing models on their outputs, while DeepSeek (MIT) and Qwen3.8/Gemma 4 (Apache 2.0) do not. Models: Amazon Nova Premier; Amazon Nova Pro; Llama 3.3 70B Instruct; DeepSeek-V4-Pro; Qwen3.8-27B BUYER CHECKLIST 1. Pin the effort/reasoning level in any latency SLA — the same model ID measures 1.70s and 19.87s to first token depending on it. 2. Price the workload at 400 input + 100 output tokens per request before comparing tiers; input-heavy RAG and output-heavy generation rank models differently. 3. Turn on prompt caching before changing model tier: cache reads cost 10% of input on most Claude models and 2.5% on Fable 5.1. 4. Use the batch API for anything not user-facing — 50% off input and output at both OpenAI and Anthropic. 5. Audit for pinned legacy model IDs. GPT-4o still lists at 2024 prices; Claude Sonnet 4.6 costs 50% more than the better Sonnet 5; Gemini 3.5 Flash costs 2x Gemini 3.8 Flash. 6. Check the anti-distillation clause before choosing a teacher if you plan to train anything on its outputs. 7. For open weights, check the base-model licence as well as the released weights licence — R1-Distill-Llama-70B is MIT weights over a Llama Community License base. 8. Confirm the context window on the cheap tier specifically. Claude Haiku 4.5 is 200K while every other current Claude is 1M. 9. Keep a second vendor integrated. In 2026 alone OpenAI cut prices twice, Anthropic cancelled an announced increase, and DeepSeek raised prices more than tenfold. 10. Evaluate small models on your own broad-knowledge tasks, not just format compliance — that is where distillation retention drops fastest. TABLE: Model register: every model a buyer could shortlist [74 rows, from data/customer.json $.extras] Model | Vendor | Role | Distilled | Teacher | Params (B) | Active params (B) | Input (USD/MTok) | Output (USD/MTok) | Blended (USD/MTok) | Cost per 1M requests (USD) | GPQA Diamond (%) | MMLU (%) | Coding (%) | AIME (%) | TTFT (ms) | Context (K tokens) | Licence | Released | Source URL ----- | ------ | ---- | --------- | ------- | ---------- | ----------------- | ---------------- | ----------------- | ------------------ | ------------------------ | ---------------- | -------- | ---------- | -------- | --------- | ------------------ | ------- | -------- | ---------- GPT-6 Astra | OpenAI | teacher | false | n/a | n/a | n/a | 10 | 50 | 20 | 9000 | n/a | n/a | n/a | n/a | n/a | 1050 | Proprietary API | 2026 | https://developers.openai.com/api/docs/pricing GPT-5.6 Sol | OpenAI | teacher | false | n/a | n/a | n/a | 4 | 20 | 8 | 3600 | 92.4 | n/a | n/a | n/a | n/a | 1050 | Proprietary API | 2026-07-09 | https://developers.openai.com/api/docs/pricing GPT-5.6 Terra | OpenAI | small-sibling | false | undisclosed | n/a | n/a | 2 | 12 | 4.5 | 2000 | 88.4 | n/a | n/a | n/a | n/a | 1050 | Proprietary API | 2026-07-09 | https://developers.openai.com/api/docs/pricing GPT-5.6 Luna | OpenAI | small-sibling | false | undisclosed | n/a | n/a | 0.2 | 1.2 | 0.45 | 200 | 87 | n/a | n/a | n/a | 1700 | 1050 | Proprietary API | 2026-07-09 | https://developers.openai.com/api/docs/pricing GPT-5.5 | OpenAI | teacher | false | n/a | n/a | n/a | 5 | 30 | 11.25 | 5000 | n/a | n/a | n/a | n/a | n/a | n/a | Proprietary API | 2026-04 | https://developers.openai.com/api/docs/pricing GPT-5.4 | OpenAI | teacher | false | n/a | n/a | n/a | 2.5 | 15 | 5.625 | 2500 | 93 | n/a | 57.7 | n/a | n/a | 1050 | Proprietary API | 2026-02 | https://developers.openai.com/api/docs/pricing GPT-5.4 mini | OpenAI | small-sibling | false | undisclosed | n/a | n/a | 0.75 | 4.5 | 1.6875 | 750 | 88 | n/a | 54.4 | n/a | n/a | 400 | Proprietary API | 2026-02 | https://developers.openai.com/api/docs/pricing GPT-5.4 nano | OpenAI | small-sibling | false | undisclosed | n/a | n/a | 0.2 | 1.25 | 0.4625 | 205 | 82.8 | n/a | 52.4 | n/a | n/a | 400 | Proprietary API | 2026-02 | https://developers.openai.com/api/docs/pricing GPT-5 | OpenAI | teacher | false | n/a | n/a | n/a | 1.25 | 10 | 3.4375 | 1500 | n/a | n/a | 74.9 | 94.6 | n/a | 400 | Proprietary API | 2025-08-07 | https://developers.openai.com/api/docs/pricing GPT-5 mini | OpenAI | small-sibling | false | undisclosed | n/a | n/a | 0.25 | 2 | 0.6875 | 300 | 80.3 | n/a | 45.7 | n/a | n/a | 400 | Proprietary API | 2025-08-07 | https://developers.openai.com/api/docs/pricing GPT-5 nano | OpenAI | small-sibling | false | undisclosed | n/a | n/a | 0.05 | 0.4 | 0.1375 | 60 | 70.9 | n/a | n/a | n/a | n/a | 400 | Proprietary API | 2025-08-07 | https://developers.openai.com/api/docs/pricing GPT-4o | OpenAI | teacher | false | n/a | n/a | n/a | 2.5 | 10 | 4.375 | 2000 | n/a | n/a | n/a | n/a | n/a | 128 | Proprietary API | 2024-05-13 | https://developers.openai.com/api/docs/pricing GPT-4o mini | OpenAI | small-sibling | false | undisclosed | n/a | n/a | 0.15 | 0.6 | 0.2625 | 120 | n/a | 82 | 87.2 | n/a | n/a | 128 | Proprietary API | 2024-07-18 | https://developers.openai.com/api/docs/pricing o3 | OpenAI | teacher | false | n/a | n/a | n/a | 2 | 8 | 3.5 | 1600 | 83.3 | n/a | 69.1 | n/a | n/a | 200 | Proprietary API | 2025-04-16 | https://developers.openai.com/api/docs/pricing o4-mini | OpenAI | small-sibling | false | undisclosed | n/a | n/a | 1.1 | 4.4 | 1.925 | 880 | 81.4 | n/a | 68.1 | 88.9 | n/a | 200 | Proprietary API | 2025-04-16 | https://developers.openai.com/api/docs/pricing gpt-oss-120b | OpenAI (open weights) | open | false | undisclosed | 116.8 | 5.1 | 0.15 | 0.6 | 0.2625 | 120 | n/a | n/a | n/a | n/a | 850 | 128 | Apache 2.0 | 2025-08 | https://console.groq.com/docs/models gpt-oss-20b | OpenAI (open weights) | open | false | undisclosed | 20.9 | 3.6 | 0.075 | 0.3 | 0.1312 | 60 | 58.59 | n/a | 53.2 | n/a | n/a | 128 | Apache 2.0 | 2025-08 | https://console.groq.com/docs/models Claude Fable 5.1 | Anthropic | teacher | false | n/a | n/a | n/a | 10 | 50 | 20 | 9000 | n/a | n/a | n/a | n/a | n/a | 1000 | Proprietary API | 2026 | https://platform.claude.com/docs/en/about-claude/pricing Claude Opus 5 | Anthropic | teacher | false | n/a | n/a | n/a | 5 | 25 | 10 | 4500 | n/a | n/a | 96 | n/a | 77250 | 1000 | Proprietary API | 2026-07-24 | https://platform.claude.com/docs/en/about-claude/pricing Claude Sonnet 5 | Anthropic | small-sibling | false | undisclosed | n/a | n/a | 2 | 10 | 4 | 1800 | n/a | n/a | 85.2 | n/a | 177770 | 1000 | Proprietary API | 2026-06-30 | https://platform.claude.com/docs/en/about-claude/pricing Claude Sonnet 4.6 | Anthropic | small-sibling | false | undisclosed | n/a | n/a | 3 | 15 | 6 | 2700 | n/a | n/a | n/a | n/a | n/a | 1000 | Proprietary API | 2026-02 | https://platform.claude.com/docs/en/about-claude/pricing Claude Haiku 4.5 | Anthropic | small-sibling | false | undisclosed | n/a | n/a | 1 | 5 | 2 | 900 | n/a | n/a | 73.3 | n/a | 19920 | 200 | Proprietary API | 2025-10-15 | https://platform.claude.com/docs/en/about-claude/pricing Claude Haiku 3.5 | Anthropic | small-sibling | false | undisclosed | n/a | n/a | 0.8 | 4 | 1.6 | 720 | n/a | n/a | n/a | n/a | n/a | 200 | Proprietary API | 2024-11 | https://platform.claude.com/docs/en/about-claude/pricing Gemini 3.1 Pro (Preview) | Google | teacher | false | n/a | n/a | n/a | 2 | 12 | 4.5 | 2000 | 94.3 | 92.6 | 80.6 | n/a | n/a | 1000 | Proprietary API | 2026-02 | https://ai.google.dev/gemini-api/docs/pricing Gemini 3.8 Flash | Google | small-sibling | false | undisclosed | n/a | n/a | 0.75 | 3.75 | 1.5 | 675 | n/a | n/a | 73.7 | n/a | n/a | 1000 | Proprietary API | 2026-09-02 | https://ai.google.dev/gemini-api/docs/pricing Gemini 3.5 Flash | Google | small-sibling | false | undisclosed | n/a | n/a | 1.5 | 9 | 3.375 | 1500 | n/a | n/a | n/a | n/a | n/a | 1000 | Proprietary API | 2026-06 | https://ai.google.dev/gemini-api/docs/pricing Gemini 3.5 Flash-Lite | Google | small-sibling | false | undisclosed | n/a | n/a | 0.3 | 2.5 | 0.85 | 370 | n/a | n/a | n/a | n/a | 6480 | 1000 | Proprietary API | 2026-06 | https://ai.google.dev/gemini-api/docs/pricing Gemini 3.1 Flash-Lite | Google | small-sibling | false | undisclosed | n/a | n/a | 0.25 | 1.5 | 0.5625 | 250 | 72.2 | 83 | n/a | n/a | n/a | 1000 | Proprietary API | 2026-03-03 | https://ai.google.dev/gemini-api/docs/pricing Gemini 2.5 Pro | Google | teacher | false | n/a | n/a | n/a | 1.25 | 10 | 3.4375 | 1500 | 86.4 | n/a | 67.2 | 88 | n/a | 1000 | Proprietary API | 2025-03 | https://ai.google.dev/gemini-api/docs/pricing Gemini 2.5 Flash | Google | distilled | true | Gemini 2.5 Pro (k-sparse logit distillation) | n/a | n/a | 0.3 | 2.5 | 0.85 | 370 | 82.8 | n/a | 60.3 | 72 | n/a | 1000 | Proprietary API | 2025-04 | https://ai.google.dev/gemini-api/docs/pricing Gemini 2.5 Flash-Lite | Google | distilled | true | Gemini 2.5 Pro (k-sparse logit distillation) | n/a | n/a | 0.1 | 0.4 | 0.175 | 80 | n/a | n/a | n/a | n/a | 300 | 1000 | Proprietary API | 2025-06 | https://ai.google.dev/gemini-api/docs/pricing Gemma 4 31B | Google | open | false | undisclosed | 30.7 | n/a | n/a | n/a | n/a | n/a | 84.3 | 85.2 | 80 | 89.2 | n/a | 256 | Apache 2.0 | 2026-04-02 | https://ai.google.dev/gemma/docs/core/model_card_4 Gemma 4 26B A4B (MoE) | Google | open | false | undisclosed | 25.2 | 3.8 | n/a | n/a | n/a | n/a | 82.3 | 82.6 | 77.1 | 88.3 | n/a | 256 | Apache 2.0 | 2026-04-02 | https://ai.google.dev/gemma/docs/core/model_card_4 Gemma 4 12B Unified | Google | open | false | undisclosed | 11.95 | n/a | n/a | n/a | n/a | n/a | 78.8 | 77.2 | 72 | 77.5 | n/a | 256 | Apache 2.0 | 2026-04-02 | https://ai.google.dev/gemma/docs/core/model_card_4 Gemma 4 E4B | Google | open | false | undisclosed | 4.5 | n/a | n/a | n/a | n/a | n/a | 58.6 | 69.4 | 52 | 42.5 | n/a | 128 | Apache 2.0 | 2026-04-02 | https://ai.google.dev/gemma/docs/core/model_card_4 Gemma 4 E2B | Google | open | false | undisclosed | 2.3 | n/a | n/a | n/a | n/a | n/a | 43.4 | 60 | 44 | 37.5 | n/a | 128 | Apache 2.0 | 2026-04-02 | https://ai.google.dev/gemma/docs/core/model_card_4 Gemma 3 27B IT | Google | open | false | undisclosed | 27 | n/a | n/a | n/a | n/a | n/a | 24.3 | 78.6 | 48.8 | n/a | n/a | 128 | Gemma Terms of Use (custom) | 2025-03 | https://huggingface.co/google/gemma-3-27b-it Gemma 3 4B IT | Google | open | false | undisclosed | 4 | n/a | n/a | n/a | n/a | n/a | 15 | 59.6 | 36 | n/a | n/a | 128 | Gemma Terms of Use (custom) | 2025-03 | https://huggingface.co/google/gemma-3-27b-it Llama 4 Maverick | Meta | distilled | true | Llama 4 Behemoth (codistillation) | 400 | 17 | n/a | n/a | n/a | n/a | 69.8 | 80.5 | 43.4 | n/a | 920 | 1000 | Llama 4 Community License (custom commercial) | 2025-04 | https://ai.meta.com/blog/llama-4-multimodal-intelligence/ Llama 4 Scout | Meta | distilled | true | Llama 4 Behemoth (codistillation) | 109 | 17 | n/a | n/a | n/a | n/a | 57.2 | 74.3 | 32.8 | n/a | n/a | 10000 | Llama 4 Community License (custom commercial) | 2025-04 | https://ai.meta.com/blog/llama-4-multimodal-intelligence/ Llama 3.3 70B Instruct | Meta | open | false | undisclosed | 70 | n/a | 1.04 | 1.04 | 1.04 | 520 | 50.5 | 68.9 | 88.4 | n/a | n/a | 128 | Llama 3.3 Community License (custom commercial) | 2024-12-06 | https://www.together.ai/pricing Llama 3.2 3B Instruct | Meta | distilled | true | Llama 3.1 8B and 70B (token-level logit distillation after pruning) | 3 | n/a | n/a | n/a | n/a | n/a | 32.8 | 63.4 | n/a | n/a | n/a | 128 | Llama 3.2 Community License (custom commercial) | 2024-09-25 | https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md Llama 3.2 1B Instruct | Meta | distilled | true | Llama 3.1 8B and 70B (token-level logit distillation after pruning) | 1 | n/a | n/a | n/a | n/a | n/a | 27.2 | 49.3 | n/a | n/a | n/a | 128 | Llama 3.2 Community License (custom commercial) | 2024-09-25 | https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md DeepSeek-V4-Pro | DeepSeek | teacher | false | n/a | 1600 | 49 | 1.32 | 3.96 | 1.98 | 924 | 90.1 | 87.5 | 80.6 | n/a | 1900 | 1000 | MIT | 2026-04-26 | https://api-docs.deepseek.com/quick_start/pricing/ DeepSeek-V4-Flash | DeepSeek | distilled | true | DeepSeek V4 domain experts (on-policy distillation consolidation) | 284 | 13 | 0.44 | 1.32 | 0.66 | 308 | 88.1 | 86.2 | 79 | n/a | 1190 | 1000 | MIT | 2026-07 | https://api-docs.deepseek.com/quick_start/pricing/ DeepSeek-V3.2 | DeepSeek | open | true | DeepSeek specialist models (specialist distillation into the generalist) | n/a | n/a | n/a | n/a | n/a | n/a | 82.4 | 85 | 73.1 | 93.1 | n/a | 128 | MIT | 2025-12 | https://arxiv.org/html/2512.02556 DeepSeek-R1 (0528) | DeepSeek | teacher | false | n/a | 685 | n/a | n/a | n/a | n/a | n/a | 81 | 85 | 73.3 | 87.5 | n/a | 128 | MIT | 2025-05-28 | https://huggingface.co/deepseek-ai/DeepSeek-R1-0528 DeepSeek-R1-Distill-Llama-70B | DeepSeek / Meta base | distilled | true | DeepSeek-R1 | 70 | n/a | n/a | n/a | n/a | n/a | 65.2 | n/a | 57.5 | 70 | n/a | 128 | MIT (weights) over Llama 3.3 Community License base | 2025-01-20 | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1-Distill-Qwen-32B | DeepSeek / Qwen base | distilled | true | DeepSeek-R1 | 32 | n/a | n/a | n/a | n/a | n/a | 62.1 | n/a | 57.2 | 72.6 | n/a | 128 | MIT (weights), Qwen2.5-32B base under Apache 2.0 | 2025-01-20 | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1-Distill-Qwen-14B | DeepSeek / Qwen base | distilled | true | DeepSeek-R1 | 14 | n/a | n/a | n/a | n/a | n/a | 59.1 | n/a | 53.1 | 69.7 | n/a | 128 | MIT (weights), Qwen2.5-14B base under Apache 2.0 | 2025-01-20 | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1-Distill-Llama-8B | DeepSeek / Meta base | distilled | true | DeepSeek-R1 | 8 | n/a | n/a | n/a | n/a | n/a | 49 | n/a | 39.6 | 50.4 | n/a | 128 | MIT (weights) over Llama 3.1 Community License base | 2025-01-20 | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1-Distill-Qwen-7B | DeepSeek / Qwen base | distilled | true | DeepSeek-R1 | 7 | n/a | n/a | n/a | n/a | n/a | 49.1 | n/a | 37.6 | 55.5 | n/a | 128 | MIT (weights), Qwen2.5-Math-7B base under Apache 2.0 | 2025-01-20 | https://huggingface.co/deepseek-ai/DeepSeek-R1 DeepSeek-R1-Distill-Qwen-1.5B | DeepSeek / Qwen base | distilled | true | DeepSeek-R1 | 1.5 | n/a | n/a | n/a | n/a | n/a | 33.8 | n/a | 16.9 | 28.9 | n/a | 128 | MIT (weights), Qwen2.5-Math-1.5B base under Apache 2.0 | 2025-01-20 | https://huggingface.co/deepseek-ai/DeepSeek-R1 Qwen3.8-27B | Alibaba | open | false | undisclosed | 27 | n/a | 0.5 | 3 | 1.125 | 500 | 89.2 | n/a | 61.7 | n/a | n/a | 262 | Apache 2.0 | 2026-08 | https://www.alibabacloud.com/help/en/model-studio/model-pricing qwen3.8-max | Alibaba | teacher | false | n/a | n/a | n/a | 2 | 6 | 3 | 1400 | n/a | n/a | n/a | n/a | n/a | n/a | Proprietary API | 2026 | https://www.alibabacloud.com/help/en/model-studio/model-pricing qwen3.8-flash | Alibaba | small-sibling | false | undisclosed | n/a | n/a | 0.15 | 0.47 | 0.23 | 107 | n/a | n/a | n/a | n/a | n/a | n/a | Proprietary API | 2026 | https://www.alibabacloud.com/help/en/model-studio/model-pricing qwen-turbo | Alibaba | small-sibling | false | undisclosed | n/a | n/a | 0.05 | 0.2 | 0.0875 | 40 | n/a | n/a | n/a | n/a | n/a | n/a | Proprietary API | 2025 | https://www.alibabacloud.com/help/en/model-studio/model-pricing Qwen3-4B-Instruct-2507 | Alibaba | distilled | true | Qwen3-32B / Qwen3-235B-A22B (off-policy + on-policy strong-to-weak distillation) | 4 | n/a | n/a | n/a | n/a | n/a | 62 | 69.6 | 35.1 | 47.4 | n/a | 262 | Apache 2.0 | 2025-07 | https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507 qwen3-8b (hosted) | Alibaba | distilled | true | Qwen3-32B / Qwen3-235B-A22B (strong-to-weak distillation) | 8 | n/a | 0.18 | 0.7 | 0.31 | 142 | n/a | n/a | n/a | n/a | n/a | 128 | Apache 2.0 | 2025-04 | https://www.alibabacloud.com/help/en/model-studio/model-pricing Phi-4 (14B) | Microsoft | open | false | undisclosed | 14 | n/a | n/a | n/a | n/a | n/a | 56.1 | 70.4 | 82.6 | n/a | n/a | 16 | MIT | 2024-12 | https://huggingface.co/microsoft/phi-4 Phi-4-mini-instruct (3.8B) | Microsoft | open | false | undisclosed | 3.8 | n/a | n/a | n/a | n/a | n/a | 25.2 | 52.8 | n/a | n/a | n/a | 128 | MIT | 2025-02 | https://huggingface.co/microsoft/Phi-4-mini-instruct Mistral Medium 3.5 | Mistral AI | teacher | false | n/a | n/a | n/a | 1.5 | 7.5 | 3 | 1350 | n/a | n/a | n/a | n/a | n/a | n/a | Modified MIT | 2026-04 | https://mistral.ai/pricing/api Mistral Large 3 | Mistral AI | open | false | undisclosed | n/a | n/a | 0.5 | 1.5 | 0.75 | 350 | n/a | n/a | n/a | n/a | n/a | n/a | Apache 2.0 | 2025-12 | https://mistral.ai/pricing/api Mistral Small 4 | Mistral AI | small-sibling | false | undisclosed | n/a | n/a | 0.15 | 0.6 | 0.2625 | 120 | 71.2 | 78 | n/a | n/a | n/a | n/a | Apache 2.0 | 2026-03-16 | https://mistral.ai/pricing/api Ministral 3 14B Instruct | Mistral AI | open | false | undisclosed | 14 | n/a | 0.2 | 0.2 | 0.2 | 100 | 71.2 | n/a | 64.6 | 85 | n/a | 256 | Apache 2.0 | 2025-12-02 | https://mistral.ai/pricing/api Ministral 3 8B | Mistral AI | open | false | undisclosed | 8 | n/a | 0.15 | 0.15 | 0.15 | 75 | n/a | n/a | n/a | n/a | n/a | 256 | Apache 2.0 | 2025-12-02 | https://mistral.ai/pricing/api Ministral 3 3B | Mistral AI | open | false | undisclosed | 3 | n/a | 0.1 | 0.1 | 0.1 | 50 | n/a | n/a | n/a | n/a | n/a | 256 | Apache 2.0 | 2025-12-02 | https://mistral.ai/pricing/api Amazon Nova Premier | Amazon | teacher | false | n/a | n/a | n/a | n/a | n/a | n/a | n/a | n/a | n/a | n/a | n/a | n/a | 1000 | Proprietary API | 2025-04 | https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html Amazon Nova Pro | Amazon | teacher | false | n/a | n/a | n/a | 0.8 | 3.2 | 1.4 | 640 | 46.9 | 85.9 | n/a | n/a | n/a | 300 | Proprietary API | 2024-12 | https://ucstrategies.com/news/amazon-nova-pro-aws-bedrock-model-guide-specs-pricing-2026/ Amazon Nova Lite | Amazon | distilled | false | n/a | n/a | n/a | 0.06 | 0.24 | 0.105 | 48 | 42 | 80.5 | n/a | n/a | n/a | 300 | Proprietary API | 2024-12 | https://pricepertoken.com/pricing-page/model/amazon-nova-lite-v1 Amazon Nova Micro | Amazon | distilled | false | n/a | n/a | n/a | 0.035 | 0.14 | 0.0613 | 28 | 40 | 77.6 | n/a | n/a | n/a | 128 | Proprietary API | 2024-12 | https://pricepertoken.com/pricing-page/model/amazon-nova-micro-v1 Amazon Nova 2 Lite | Amazon | small-sibling | false | undisclosed | n/a | n/a | 0.3 | 2.5 | 0.85 | 370 | n/a | n/a | n/a | n/a | n/a | 1000 | Proprietary API | 2025-12 | https://pricepertoken.com/pricing-page/model/amazon-nova-2-lite-v1 SmolLM3-3B | Hugging Face | open | false | n/a | 3 | n/a | n/a | n/a | n/a | n/a | 41.7 | n/a | 30.48 | 36.7 | n/a | 128 | Apache 2.0 | 2025-07 | https://huggingface.co/HuggingFaceTB/SmolLM3-3B grok-4.6 | xAI | teacher | false | n/a | n/a | n/a | 2 | 6 | 3 | 1400 | n/a | n/a | n/a | n/a | n/a | 200 | Proprietary API | 2026-08 | https://docs.x.ai/docs/models TIMELINE OF THIS PERSPECTIVE (21 EVENTS) ---------------------------------------- 2024-07-18 | product | GPT-4o mini launches at $0.15/$0.60 2024-09-25 | product | Llama 3.2 1B/3B ship as documented distillations 2024-12-03 | product | Amazon Bedrock Model Distillation announced 2024-12-06 | product | Llama 3.3 70B released 2025-01-20 | product | DeepSeek ships R1 plus six distilled students under MIT 2025-04 | product | Llama 4 Scout and Maverick ship as codistilled models 2025-06 | research | Gemini 2.5 report confirms the Flash line is distilled 2025-08-07 | product | GPT-5, GPT-5 mini and GPT-5 nano launch together 2025-10-15 | product | Claude Haiku 4.5 released at $1/$5 2025-12-02 | product | Ministral 3 (3B/8B/14B) ships Apache 2.0 with 256K context 2026-03-16 | product | Mistral Small 4 released under Apache 2.0 2026-04-02 | product | Gemma 4 ships Apache 2.0 with frontier-class small models 2026-04-26 | product | DeepSeek-V4-Pro released under MIT 2026-06-30 | product | Claude Sonnet 5 launches at $2/$10 with 1M context 2026-07-09 | product | GPT-5.6 Sol, Terra and Luna launch as one price ladder 2026-07-24 | product | Claude Opus 5 released at $5/$25 2026-07-30 | market | OpenAI cuts GPT-5.6 Luna by 80% and Terra by 20% 2026-08-11 | market | Anthropic makes Sonnet 5 $2/$10 permanent 2026-08-08 | market | Enterprise inference index hits a 2026 low of $1.16–$1.18/MTok 2026-08-16 | market | DeepSeek raises V4 standard rates ~3–4.7x, and cache-hit input rates by up to 11x 2026-09-02 | product | Gemini 3.8 Flash ships at $0.75/$3.75 SOURCES CITED BY THIS SECTION (78) ---------------------------------- [1] OpenAI API pricing — OpenAI, 2026-09 (pricing) https://developers.openai.com/api/docs/pricing [2] OpenAI API models reference — OpenAI, 2026-09 (docs) https://developers.openai.com/api/docs/models [3] GPT-4o mini: advancing cost-efficient intelligence — OpenAI, 2024-07-18 (blog) https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ [4] OpenAI GPT-5 System Card — OpenAI / arXiv, 2026-01 (paper) https://arxiv.org/pdf/2601.03267 [5] OpenAI Services Agreement — OpenAI, 2026-01-01 (law) https://openai.com/policies/services-agreement/ [6] GPT-5.6 Sol model page — OpenRouter, 2026-07-09 (docs) https://openrouter.ai/openai/gpt-5.6-sol [7] GPT-5.6 Terra model page — OpenRouter, 2026-07-09 (docs) https://openrouter.ai/openai/gpt-5.6-terra [8] GPT-5.6 Luna model page — OpenRouter, 2026-07-09 (docs) https://openrouter.ai/openai/gpt-5.6-luna [9] GPT-5 mini model page — OpenRouter, 2025-08-07 (docs) https://openrouter.ai/openai/gpt-5-mini [10] GPT-5 nano model page — OpenRouter, 2025-08-07 (docs) https://openrouter.ai/openai/gpt-5-nano [11] GPT-4o model page — OpenRouter, 2024-05-13 (docs) https://openrouter.ai/openai/gpt-4o [12] o4-mini model page — OpenRouter, 2025-04-16 (docs) https://openrouter.ai/openai/o4-mini [13] OpenAI ships GPT-5.4 mini and nano — The Decoder, 2026-02 (news) https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/ [14] o4-mini: tests, features, o3 comparison, benchmarks — DataCamp, 2025-04 (news) https://www.datacamp.com/blog/o4-mini [15] gpt-oss-20b model card — OpenAI / Hugging Face, 2025-08 (docs) https://huggingface.co/openai/gpt-oss-20b [16] Claude platform pricing — Anthropic, 2026-09 (pricing) https://platform.claude.com/docs/en/about-claude/pricing [17] Claude models overview — Anthropic, 2026-09 (docs) https://platform.claude.com/docs/en/about-claude/models/overview [18] Claude Opus 5 model page — Anthropic, 2026-07-24 (docs) https://platform.claude.com/docs/en/models/opus-5/overview [19] Introducing Claude Sonnet 5 — Anthropic, 2026-06-30 (blog) https://www.anthropic.com/news/claude-sonnet-5 [20] Introducing Claude Haiku 4.5 — Anthropic, 2025-10-15 (blog) https://www.anthropic.com/news/claude-haiku-4-5 [21] Anthropic Commercial Terms of Service — Anthropic, 2025-06-17 (law) https://www.anthropic.com/legal/commercial-terms [22] Claude Opus 5 by Anthropic: benchmarks and pricing — DataNorth, 2026-07 (news) https://datanorth.ai/news/claude-opus-5-by-anthropic [23] Claude benchmarks 2026 — Morph, 2026-09 (news) https://www.morphllm.com/claude-benchmarks [24] Anthropic launches Opus 5 — TechCrunch, 2026-07-24 (news) https://techcrunch.com/2026/07/24/anthropic-launches-opus-5/ [25] Gemini API pricing — Google, 2026-09 (pricing) https://ai.google.dev/gemini-api/docs/pricing [26] Gemini API models — Google, 2026-09 (docs) https://ai.google.dev/gemini-api/docs/models [27] Gemini 3.1 Pro model page — Google DeepMind, 2026-02 (docs) https://deepmind.google/models/gemini/pro/ [28] Gemini Flash model page — Google DeepMind, 2026-09-02 (docs) https://deepmind.google/models/gemini/flash/ [29] Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context and Next Generation Agentic Capabilities — Google DeepMind / arXiv, 2025-07 (paper) https://arxiv.org/html/2507.06261v1/ [30] Gemma 4 model card — Google, 2026-04-02 (docs) https://ai.google.dev/gemma/docs/core/model_card_4 [31] Gemma 4 — Google DeepMind, 2026-04-02 (blog) https://deepmind.google/models/gemma/gemma-4/ [32] gemma-3-27b-it model card — Google / Hugging Face, 2025-03 (docs) https://huggingface.co/google/gemma-3-27b-it [33] Gemini 3.8 Flash rolling out three weeks after last release — 9to5Google, 2026-09-02 (news) https://9to5google.com/2026/09/02/gemini-3-8-flash-launch/ [34] Gemini 3.8 Flash model stats — LLM-Stats, 2026-09-02 (news) https://llm-stats.com/models/gemini-3.8-flash [35] Gemini 3.1 Flash-Lite benchmark results — LayerLens, 2026-03 (news) https://layerlens.ai/blog/gemini-3-1-flash-lite-benchmark-results-efficiency-model-comparison [36] DeepSeek API pricing — DeepSeek, 2026-08 (pricing) https://api-docs.deepseek.com/quick_start/pricing/ [37] DeepSeek-R1 model card (with distilled model evaluations) — DeepSeek / Hugging Face, 2025-01-20 (docs) https://huggingface.co/deepseek-ai/DeepSeek-R1 [38] DeepSeek-R1-0528 model card — DeepSeek / Hugging Face, 2025-05-28 (docs) https://huggingface.co/deepseek-ai/DeepSeek-R1-0528 [39] DeepSeek-R1-Distill-Qwen-32B model card — DeepSeek / Hugging Face, 2025-01-20 (docs) https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B [40] DeepSeek-V4-Pro model card — DeepSeek / Hugging Face, 2026-04-26 (docs) https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro [41] DeepSeek-V4-Flash model card — DeepSeek / Hugging Face, 2026-07 (docs) https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash [42] DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models — DeepSeek / arXiv, 2025-12 (paper) https://arxiv.org/html/2512.02556 [43] DeepSeek raises some V4 prices by more than 10x — InfoWorld, 2026-08 (news) https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html [44] Llama 4 model card — Meta, 2025-04 (docs) https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md [45] Llama 3.3 model card — Meta, 2024-12-06 (docs) https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md [46] Llama 3.2 model card — Meta, 2024-09-25 (docs) https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md [47] Qwen3 Technical Report — Alibaba Qwen Team / arXiv, 2025-05 (paper) https://arxiv.org/pdf/2505.09388 [48] Qwen3.8-27B model card — Alibaba / Hugging Face, 2026-08 (docs) https://huggingface.co/Qwen/Qwen3.8-27B [49] Qwen3-4B-Instruct-2507 model card — Alibaba / Hugging Face, 2025-07 (docs) https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507 [50] Alibaba Model Studio model pricing — Alibaba Cloud, 2026-09 (pricing) https://www.alibabacloud.com/help/en/model-studio/model-pricing [51] phi-4 model card — Microsoft / Hugging Face, 2024-12 (docs) https://huggingface.co/microsoft/phi-4 [52] Phi-4-mini-instruct model card — Microsoft / Hugging Face, 2025-02 (docs) https://huggingface.co/microsoft/Phi-4-mini-instruct [53] SmolLM3-3B model card — Hugging Face, 2025-07 (docs) https://huggingface.co/HuggingFaceTB/SmolLM3-3B [54] Mistral AI API pricing — Mistral AI, 2026-09 (pricing) https://mistral.ai/pricing/api [55] Mistral models overview — Mistral AI, 2026-09 (docs) https://docs.mistral.ai/getting-started/models/models_overview/ [56] Ministral-3-14B-Instruct-2512 model card — Mistral AI / Hugging Face, 2025-12-02 (docs) https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512 [57] Mistral Small 4 model page — OpenRouter, 2026-03-16 (docs) https://openrouter.ai/mistralai/mistral-small-2603 [58] What is Amazon Nova? (teacher/student distillation matrix) — Amazon Web Services, 2026 (docs) https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html [59] What’s new in Amazon Nova 2 — Amazon Web Services, 2025-12 (docs) https://docs.aws.amazon.com/nova/latest/nova2-userguide/whats-new.html [60] The Amazon Nova Family of Models: Technical Report and Model Card — Amazon Science, 2025-03-17 (paper) https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf [61] Amazon Bedrock Model Distillation — Amazon Web Services, 2025-05 (docs) https://aws.amazon.com/bedrock/model-distillation/ [62] Amazon Bedrock pricing — Amazon Web Services, 2026-09 (pricing) https://aws.amazon.com/bedrock/pricing/ [63] Nova Micro API pricing — PricePerToken, 2026 (pricing) https://pricepertoken.com/pricing-page/model/amazon-nova-micro-v1 [64] Nova Lite API pricing — PricePerToken, 2026 (pricing) https://pricepertoken.com/pricing-page/model/amazon-nova-lite-v1 [65] Nova 2 Lite API pricing — PricePerToken, 2026 (pricing) https://pricepertoken.com/pricing-page/model/amazon-nova-2-lite-v1 [66] Amazon Nova Pro: AWS Bedrock model guide, specs and pricing (2026) — UC Strategies, 2026 (news) https://ucstrategies.com/news/amazon-nova-pro-aws-bedrock-model-guide-specs-pricing-2026/ [67] Together AI pricing — Together AI, 2026-09 (pricing) https://www.together.ai/pricing [68] Groq supported models and pricing — Groq, 2026-09 (pricing) https://console.groq.com/docs/models [69] xAI models and pricing — xAI, 2026-08 (pricing) https://docs.x.ai/docs/models [70] GPT-5.6 Luna (low) vs Claude 4.5 Haiku (reasoning) — Artificial Analysis, 2026-09 (news) https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning [71] Gemini 3.5 Flash-Lite vs GPT-5.6 Luna (high) — Artificial Analysis, 2026-09 (news) https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gpt-5-6-luna-high [72] Claude Sonnet 5 vs Claude Opus 5 — Artificial Analysis, 2026-09 (news) https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5 [73] DeepSeek V4 Flash vs DeepSeek V4 Pro — Artificial Analysis, 2026-09 (news) https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-deepseek-v4-pro [74] gpt-oss-120B vs Llama 4 Maverick — Artificial Analysis, 2026-09 (news) https://artificialanalysis.ai/models/comparisons/gpt-oss-120b-vs-llama-4-maverick [75] Comparison of AI models across intelligence, performance and price — Artificial Analysis, 2026-09 (news) https://artificialanalysis.ai/models [76] Enterprise AI costs hit 2026 low driven by price wars and Chinese open-source models — South China Morning Post, 2026-08 (news) https://www.scmp.com/tech/tech-trends/article/3363549/enterprise-ai-costs-hit-2026-low-driven-price-wars-chinese-open-source-models-research [77] OpenAI discounts GPT-5.6 Luna and Terra — Axios, 2026-07-30 (news) https://www.axios.com/2026/07/30/openai-cuts-prices-gpt-terra-luna5 [78] The Llama 4 herd: the beginning of a new era of natively multimodal AI innovation — Meta AI, 2025-04-05 (blog) https://ai.meta.com/blog/llama-4-multimodal-intelligence/ ================================================================================================ 9. METHOD LIBRARY — THE DISTILLATION METHOD LIBRARY ================================================================================================ Page: https://global-distillation.com/library Data: https://global-distillation.com/data/library.json Updated: 2026-09-04 SUMMARY -------- Knowledge distillation is not one technique but a family of at least two dozen distinct methods, separated by what signal crosses from teacher to student (logits, hidden features, pairwise relations, sampled text, preferences, or denoising trajectories) and by whether you can open the teacher at all. This library catalogues 27 methods with their loss functions, data and access requirements, tooling, and honest trade-offs. The central split is white-box versus black-box: white-box methods (soft targets, feature hints, reverse-KL, on-policy GKD) need teacher logits or activations and give the densest supervision per token, while black-box methods (SFT on API outputs, chain-of-thought and reasoning-trace distillation) need only sampled text and are what most practitioners actually run — and what triggered the 2025-2026 legal fights. The frontier has moved from static logit matching to on-policy methods where the student generates and the teacher grades: the Qwen3 technical report (Table 21) gives 74.4% on AIME'24 for a Qwen3-8B student at 1,800 GPU-hours of on-policy distillation versus 67.6% for RL at 17,920 GPU-hours, and Thinking Machines reproduced the method on the Tinker API. Method choice is mostly determined by three constraints — teacher access, label availability, and whether the student's architecture or tokenizer matches the teacher's — and this library is organised so you can walk those three constraints to a shortlist. KEY FIGURES ----------- - Methods catalogued here: 27 methods (spanning 2014-2025) Counted in extras.methods; families extend the response / feature / relation trichotomy of Gou et al. (arXiv 2006.05525) with six additional families specific to generative and compression-composed KD; eleven families in total. Source: https://arxiv.org/abs/2006.05525 - DistilBERT GLUE retention: 97 % of BERT-base (40% fewer params, 60% faster) The canonical compression datapoint for response+feature distillation on encoders. Source: https://arxiv.org/abs/1910.01108 - TinyBERT-4L GLUE retention: 96.8 % of BERT-base (7.5x smaller, 9.4x faster) Adds attention-matrix and hidden-state losses on top of logit matching. Source: https://arxiv.org/abs/1909.10351 - DeepSeek-R1-Distill-Qwen-32B, AIME 2024: 72.6 % pass@1 (pure SFT on 800k R1 traces, no RL stage) Shows sequence-level reasoning-trace distillation alone can transplant frontier reasoning into a 32B open-weight student. Source: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B - Alpaca teacher-data generation cost: 500 USD (upper bound) (52k instructions; <$100 more to finetune) Stanford CRFM reports under $500 of OpenAI API calls plus 3 hours on 8x A100 80GB. Source: https://crfm.stanford.edu/2023/03/13/alpaca.html - s1K reasoning-distillation dataset size: 1,000 examples (beats o1-preview on AIME24/MATH by up to 27%) 1,000 curated questions with traces distilled from Gemini Thinking Experimental, plus budget forcing at inference. Source: https://arxiv.org/abs/2501.19393 - On-policy distillation vs RL, GPU-hours: 1,800 GPU-hours (vs 17,920 for RL at a lower score) Qwen3-8B student: 74.4% AIME'24 via on-policy distillation at 1,800 GPU-hours vs 67.6% via RL at 17,920. Source: https://thinkingmachines.ai/blog/on-policy-distillation/ - Latent Consistency Model training cost: 32 A100 GPU-hours (50 steps -> 2-4 steps at 768x768) Diffusion distillation is the cheapest high-leverage distillation in this library. Source: https://arxiv.org/abs/2310.04378 KEY FINDINGS ------------ 1. The first question is not 'which loss' but 'what does the teacher let you see' Every method in this library sits on one side of a hard line. White-box methods need per-token logits or internal activations and are only available if you host the teacher's weights. Black-box methods need only sampled text and work against any API. That single constraint eliminates roughly half the catalogue before you compare anything else. If you are distilling from a commercial API you are restricted to sequence-level, chain-of-thought, preference and dataset methods — no soft targets, no feature hints, no reverse-KL. Sources: https://arxiv.org/abs/2402.13116 https://arxiv.org/abs/1606.07947 2. Soft targets carry more information per example than labels — that is the whole original insight Hinton, Vinyals and Dean's 2015 argument is that a teacher's full probability vector over wrong classes (the 'dark knowledge': a 2 that looks slightly like a 7) has high entropy and therefore more bits per training case and lower gradient variance than a one-hot label. Raising the softmax temperature exposes those relative probabilities. Because gradients through the softened softmax scale as 1/T^2, the soft-target term must be multiplied by T^2 to keep the two objectives balanced when T is tuned. Sources: https://www.cs.toronto.edu/~hinton/absps/distillation.pdf https://arxiv.org/abs/1503.02531 3. On-policy methods fixed the exposure-bias problem and are now the default frontier recipe Off-policy distillation trains the student on the teacher's own sequences, so at inference the student meets its own distribution for the first time and compounds errors. GKD (2023) and MiniLLM (2023) fix this by sampling from the student and having the teacher grade those tokens, with a divergence (reverse KL or generalized JSD) that tolerates a student too small to cover the teacher's modes. The Qwen3 technical report (Table 21) reports a Qwen3-8B student at 74.4 on AIME'24 for 1,800 GPU-hours of on-policy distillation versus 67.6 for RL at 17,920, up from 55.0 off-policy; Thinking Machines reproduced the method with the Tinker API and reached roughly 70% AIME'24 from a Qwen3-8B-Base SFT-400K checkpoint. Sources: https://arxiv.org/abs/2306.13649 https://arxiv.org/abs/2306.08543 https://arxiv.org/html/2505.09388v1 https://thinkingmachines.ai/blog/on-policy-distillation/ 4. Rationales beat labels: distilling the reasoning is worth more than distilling the answer Distilling Step-by-Step extracts teacher rationales as a second supervised task alongside the label, and reports a 770M T5 outperforming a few-shot-prompted 540B PaLM using only 80% of the available data — roughly a 700x parameter reduction. DeepSeek pushed the same idea to its limit by fine-tuning Qwen and Llama students on 800k long chain-of-thought traces from R1, with no RL stage, reaching 72.6% pass@1 on AIME 2024 at 32B. Sources: https://arxiv.org/abs/2305.02301 https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B 5. Distillation is not free: scaling laws say it only wins when the teacher cost is amortised Apple's 2025 distillation scaling law finds that when a teacher already exists, or when one teacher serves many students, distillation beats supervised learning up to a compute level that scales predictably with student size. But if you must train a teacher to produce exactly one student, plain supervised learning is generally preferable. This is the single most useful economic result in the field and it contradicts the folk assumption that distillation is always the cheaper path. Sources: https://arxiv.org/abs/2502.08606 6. Prune first, then distill — compression now routinely stacks The strongest small-model recipes are compositions, not single methods. NVIDIA's Minitron prunes Llama 3.1 8B along depth or width and re-trains with distillation on under 1% of the original token budget (94B tokens vs 15T, ~160x fewer), yielding 1.8x-2.7x speedups. Sheared LLaMA combines targeted structured pruning with dynamic batch loading to beat same-size models trained from scratch. Quantization-aware distillation then recovers the accuracy that low-bit quantization costs. Sources: https://arxiv.org/abs/2408.11796 https://arxiv.org/abs/2407.14679 https://arxiv.org/abs/2310.06694 https://arxiv.org/abs/2305.17888 7. Tokenizer mismatch is the quiet blocker on white-box distillation Logit-matching requires teacher and student to share a vocabulary, which rules out most cross-family pairs (Llama teacher, Qwen student). Two 2024 lines of work attack this: Universal Logit Distillation uses an optimal-transport cost between sorted probability vectors so no token alignment is needed, and Dual-Space KD projects both models into a shared output space with a cross-model attention mechanism. Both are still less reliable than same-tokenizer KD, and torchtune's own roadmap lists cross-tokenizer support as future work. Sources: https://arxiv.org/abs/2402.12030 https://arxiv.org/abs/2406.17328 https://pytorch.org/blog/llama-into-torchtune/ 8. Distillation is now infrastructure, not research — and its abuse is now a public legal fight OpenAI shipped Model Distillation into its API in late 2024 (stored completions plus evals plus fine-tuning), and Amazon Bedrock Model Distillation reached general availability on 1 May 2025 with claims of up to 500% faster and up to 75% cheaper distilled models at under 2% accuracy loss for RAG. The flip side arrived on 23 February 2026, when Anthropic disclosed what it called industrial-scale distillation attacks: roughly 24,000 fraudulent accounts and over 16 million exchanges attributed to DeepSeek, Moonshot AI and MiniMax, targeting agentic reasoning, tool use and coding. Sources: https://openai.com/index/api-model-distillation/ https://aws.amazon.com/blogs/machine-learning/amazon-bedrock-model-distillation-boost-function-calling-accuracy-while-reducing-cost-and-latency/ https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks 9. A fine-tuned teacher distills better than a raw one The PyTorch torchtune case study distilling Llama 3.1 8B into Llama 3.2 1B on Alpaca found that KD loss stayed nearly flat when the teacher was not first adapted to the transfer set, and that using a LoRA-finetuned teacher gave the best hellaswag and commonsense results. NVIDIA independently reports the same thing as 'teacher correction' — a light fine-tune of the teacher on the distillation corpus before pruning-and-distilling. The lesson: the teacher's distribution must match the transfer data or the distillation signal is nearly uninformative. Sources: https://pytorch.org/blog/llama-into-torchtune/ https://arxiv.org/abs/2408.11796 10. Distillation has escaped model compression entirely Four of the methods here do not shrink a model at all. Self-distillation (Born-Again Networks) trains a student identical to its teacher and still beats it. Dataset distillation compresses the training set rather than the network. Draft-model distillation (EAGLE, Medusa) trains a tiny head purely to make speculative decoding accept more tokens, leaving the target model's output distribution unchanged. Diffusion distillation compresses the number of sampling steps, not the parameter count — LCM reaches 2-4 step generation for 32 A100-hours. Sources: https://arxiv.org/abs/1805.04770 https://arxiv.org/abs/1811.10959 https://arxiv.org/abs/2401.15077 https://arxiv.org/abs/2310.04378 TABLES -------- TABLE: The eleven families of distillation, and what actually crosses from teacher to student [id: method-families, 11 rows] Every method in this library belongs to one of these families. The family determines what you must be able to read out of the teacher, which is almost always the binding constraint. Family | Signal transferred | Origin paper | Year | Teacher access | Where it dominates | Methods in this library | Source URL ------ | ------------------ | ------------ | ---- | -------------- | ------------------ | ----------------------- | ---------- Response-based (logit) | Softened output distribution over classes/tokens | Hinton, Vinyals & Dean | 2015 | white-box | Classifiers, encoders, same-tokenizer LLM pairs | 6 | https://arxiv.org/abs/1503.02531 Feature / intermediate | Hidden activations at chosen layers, via a projector | FitNets (Romero et al.) | 2014 | white-box | CNNs, BERT-family encoders | 2 | https://arxiv.org/abs/1412.6550 Attention-based | Spatial or head attention maps | Zagoruyko & Komodakis | 2016 | white-box | CNNs; transformer attention matrices | 1 | https://arxiv.org/abs/1612.03928 Relation-based | Pairwise/triplet structure of the embedding space | RKD (Park et al.) | 2019 | white-box | Metric learning, retrieval, re-ID | 3 | https://arxiv.org/abs/1904.05068 Sequence-level / data | Sampled teacher output sequences used as hard targets | Kim & Rush | 2016 | black-box | MT, instruction tuning, any API teacher | 2 | https://arxiv.org/abs/1606.07947 Rationale / reasoning | Chain-of-thought traces as an extra supervised target | Distilling Step-by-Step (Hsieh et al.) | 2023 | black-box | Math, code, multi-step QA, agents | 2 | https://arxiv.org/abs/2305.02301 On-policy / divergence-choice | Teacher scores the student's own samples | GKD (Agarwal et al.) and MiniLLM (Gu et al.) | 2023 | white-box | Generative LLMs where exposure bias dominates | 3 | https://arxiv.org/abs/2306.13649 Preference | Teacher- or AI-generated preference pairs | Zephyr dDPO (Tunstall et al.) | 2023 | black-box | Chat alignment, style, helpfulness | 2 | https://arxiv.org/abs/2310.16944 Compression-composed | KD used to repair a pruned or quantized network | Minitron / LLM-QAT | 2023 | white-box | Edge deployment, model-family shrinking | 2 | https://arxiv.org/abs/2305.17888 Trajectory / sampler | Denoising trajectory or next-feature prediction | Progressive Distillation (Salimans & Ho) | 2022 | white-box | Diffusion samplers, speculative decoding drafts | 2 | https://arxiv.org/abs/2202.00512 Dataset / context | The training set or the prompt itself, not the weights | Dataset Distillation (Wang et al.); Askell et al. | 2018 | either | Data-efficiency research; prompt internalisation | 2 | https://arxiv.org/abs/1811.10959 Notes: 'Methods in this library' counts entries in extras.methods; some methods legitimately span two families (e.g. TinyBERT is feature + attention + response) and are counted once under their dominant family. Sources: https://arxiv.org/abs/2503.12067 https://arxiv.org/abs/2402.13116 TABLE: White-box versus black-box distillation, dimension by dimension [id: white-box-vs-black-box, 12 rows] The practical decision table. If you cannot host the teacher's weights, the right-hand column is your entire option space. Dimension | White-box (weights in hand) | Black-box (API only) | Source URL --------- | ------------------------ | -------------------- | ---------- What you read from the teacher | Full logit vector per token, hidden states, attention matrices, gradients | Sampled text; sometimes top-k logprobs; nothing internal | https://arxiv.org/abs/2402.13116 Supervision density | Dense: one target distribution per token position | Sparse: one sampled sequence per prompt | https://thinkingmachines.ai/blog/on-policy-distillation/ Representative methods | Soft targets, FitNets, attention transfer, RKD, CRD, MiniLLM, GKD, SKD, Minitron, QAD | Seq-level KD, Alpaca/Vicuna-style SFT, Distilling Step-by-Step, R1-style trace distillation, dDPO, RLAIF | https://arxiv.org/abs/1606.07947 Tokenizer constraint | Must match, unless you use ULD or DSKD | None — text is tokenizer-agnostic | https://arxiv.org/abs/2402.12030 Architecture constraint | Feature methods need layer-mapping and a projector; relation methods need comparable embedding spaces | None | https://arxiv.org/abs/1412.6550 Marginal cost per training step | One extra teacher forward pass (or cached logits) | Zero after the dataset is generated; generation is a one-time API bill | https://crfm.stanford.edu/2023/03/13/alpaca.html Typical wall-clock to first result | Hours to days; needs a GPU large enough for teacher + student | Hours; dataset generation can be parallelised across API keys | https://pytorch.org/blog/llama-into-torchtune/ Handles a very small student | Better — reverse-KL and JSD variants let a low-capacity student pick modes instead of smearing | Worse — the student must imitate full sequences it cannot represent | https://arxiv.org/abs/2306.08543 Legal exposure | Governed by the teacher weights' licence (e.g. Llama, Apache-2.0, MIT) | Governed by API terms; most frontier vendors forbid using outputs to build competing models | https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks Detectability by the teacher's owner | Not applicable | High — Anthropic reports classifiers and behavioural fingerprinting that flag repeated narrow-capability prompt patterns across coordinated accounts | https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks Managed-service support | torchtune, TRL, DistillKit, NVIDIA Model Optimizer | OpenAI Model Distillation, Amazon Bedrock Model Distillation, Vertex AI | https://openai.com/index/api-model-distillation/ Where it is state of the art | Frontier-lab small models: Gemma 3 (all sizes distilled), Qwen3 strong-to-weak, Minitron | Open-community reasoning models: DeepSeek-R1-Distill, s1, Orca, Zephyr | https://arxiv.org/abs/2503.19786 Notes: 'Grey-box' is a real third case: some APIs return top-k logprobs, which supports a truncated logit-matching objective but not full-vocabulary KL. Sources: https://arxiv.org/abs/2402.13116 https://arxiv.org/abs/2306.08543 https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks TABLE: Every method: teacher access, data needed, compute profile, difficulty [id: method-compute-data-needs, 27 rows] The full 27-method matrix. 'Teacher compute' is the extra cost incurred by the teacher during student training (not the cost of training the teacher). Difficulty is a 1-5 implementation-effort rating, editorial, calibrated against the reference implementations linked in each method entry. Method | Teacher access | Data needed | Ground-truth labels? | Teacher compute during training | Difficulty (1-5) | Source URL ------ | -------------- | ----------- | -------------------- | ------------------------ | ---------------- | ---------- Response-based KD (soft targets) | white-box | logits | Optional (blended term) | One forward pass per batch, cacheable | 1 | https://arxiv.org/abs/1503.02531 Feature / hint KD (FitNets) | white-box | features | Yes, in stage two | One forward pass; plus a trained projector | 3 | https://arxiv.org/abs/1412.6550 Attention transfer | white-box | features | Yes | One forward pass | 2 | https://arxiv.org/abs/1612.03928 Relational KD (RKD) | white-box | features | Optional | One forward pass; O(batch^2) or O(batch^3) relation terms | 3 | https://arxiv.org/abs/1904.05068 Contrastive representation distillation (CRD) | white-box | features | Optional | Forward pass plus a negatives memory bank | 4 | https://arxiv.org/abs/1910.10699 Transformer-layer KD (TinyBERT / DistilBERT) | white-box | features | Yes for the task stage | Forward pass at both pretraining and task stages | 4 | https://arxiv.org/abs/1909.10351 Sequence-level KD | black-box | outputs | No | One-time beam-search generation over the corpus | 1 | https://arxiv.org/abs/1606.07947 Black-box output distillation (Alpaca/Vicuna style) | black-box | outputs | No | One-time API generation bill | 1 | https://crfm.stanford.edu/2023/03/13/alpaca.html Chain-of-thought / rationale distillation | black-box | outputs | Optional (multi-task with labels) | One-time generation, longer outputs so higher token bill | 2 | https://arxiv.org/abs/2305.02301 Reasoning-trace distillation (R1-Distill, s1) | black-box | outputs | Answers used for filtering | Very high one-time cost: long traces, rejection sampling | 2 | https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B Reverse-KL distillation (MiniLLM) | white-box | logits | No | Student generation plus teacher scoring each step | 5 | https://arxiv.org/abs/2306.08543 Generalized KD (GKD, on-policy) | white-box | logits | No | Student generation plus teacher scoring; dominates step cost | 4 | https://arxiv.org/abs/2306.13649 Speculative KD (SKD) | white-box | logits | No | Interleaved student proposal + teacher verification | 5 | https://arxiv.org/abs/2410.11325 Self-distillation (Born-Again Networks) | white-box | logits | Yes | One full training generation per round | 2 | https://arxiv.org/abs/1805.04770 Teacher-assistant KD (TAKD) | white-box | logits | Yes | Extra full training run for each intermediate model | 2 | https://arxiv.org/abs/1902.03393 Multi-teacher KD | white-box | logits | Optional | K forward passes per batch, or K cached logit sets | 3 | https://www.isca-archive.org/interspeech_2017/fukuda17_interspeech.html Deep mutual learning | none (peers) | logits | Yes | No teacher; K students trained simultaneously | 2 | https://arxiv.org/abs/1706.00384 Prune-then-distill (Minitron, Sheared LLaMA) | white-box | logits | No | Teacher forward pass over a re-training corpus (<1% of original tokens; 94B vs 15T) | 5 | https://arxiv.org/abs/2408.11796 Quantization-aware distillation | white-box | logits | No (LLM-QAT is data-free) | Full-precision teacher forward pass each step | 4 | https://arxiv.org/abs/2305.17888 Dataset distillation / condensation | none | none (synthesises data) | Yes, for the real set | No teacher; bi-level optimisation over inner training steps | 5 | https://arxiv.org/abs/1811.10959 Preference distillation / RLAIF | black-box | outputs | No | Teacher ranks candidate completions | 4 | https://arxiv.org/abs/2310.16944 DPO on teacher preferences (dDPO) | black-box | outputs | No | One-time preference labelling; no reward model, no PPO loop | 2 | https://arxiv.org/abs/2305.18290 Cross-vocabulary logit KD (ULD, DSKD) | white-box | logits | Optional | Teacher forward pass plus an optimal-transport or projection step | 5 | https://arxiv.org/abs/2402.12030 Draft-model distillation (EAGLE, Medusa) | white-box | features | No | Target-model features over a modest corpus; target weights frozen | 4 | https://arxiv.org/abs/2401.15077 Retriever distillation (cross-encoder to bi-encoder) | white-box or grey | outputs (scores) | No, teacher scores replace labels | Cross-encoder scoring over query-document pairs; usually precomputed | 3 | https://arxiv.org/abs/2010.02666 Diffusion distillation (progressive, consistency, LCM, ADD) | white-box | features | No | Teacher sampler queries; LCM reports 32 A100-hours end to end | 4 | https://arxiv.org/abs/2310.04378 Context / prompt distillation | white-box (self) | logits | No | The same model prompted with the context acts as teacher | 2 | https://arxiv.org/abs/2209.15189 Notes: Difficulty is an editorial 1-5 implementation-effort rating, not a measured quantity. 'Data needed' uses the SCHEMA vocabulary: logits | outputs | features | none. The 'difficulty' column is an editorial 1-5 rating; each row's _source is the method's own paper, not a source for that rating. Sources: https://arxiv.org/abs/2402.13116 https://huggingface.co/docs/trl/gkd_trainer https://pytorch.org/blog/llama-into-torchtune/ TABLE: Loss-function cheat sheet with the hyperparameters that actually matter [id: loss-cheatsheet, 18 rows] What you tune, and what a sensible starting value looks like according to the reference implementation. Method | Core objective | Hyperparameters that matter | Reference default | Reference implementation | Source URL ------ | -------------- | ------------------------ | ----------------- | ------------------------ | ---------- Response-based KD | a * T^2 * KL(teacher_T // student_T) + (1-a) * CE(labels) | Temperature T, mixing weight alpha | T in 2-10; alpha 0.5-0.9 | torchtune ForwardKLWithChunkedOutputLoss | https://pytorch.org/blog/llama-into-torchtune/ torchtune KD recipe | CE + forward-KL on logits | kd_ratio, learning rate | kd_ratio 0.5, lr 3e-4; the blog finds kd_ratio 0.75-1.0 slightly better | knowledge_distillation_single_device | https://pytorch.org/blog/llama-into-torchtune/ FitNets hint | MSE(projector(student_feat), teacher_feat) | Which layer pair; projector width | Middle-layer hint, 1x1 conv regressor | RepDistiller | https://arxiv.org/abs/1412.6550 Attention transfer | L_p distance between L2-normalised attention maps | p in the map power sum; layer group set | p = 2, per residual-block group | szagoruyko/attention-transfer | https://arxiv.org/abs/1612.03928 RKD | Huber loss on distance ratios + angle cosines | Weighting of distance vs angle terms | Both terms, distance normalised by batch mean | RepDistiller | https://arxiv.org/abs/1904.05068 CRD | InfoNCE-style contrastive bound on mutual information | Number of negatives N, embedding dim | N in the thousands via a memory buffer | HobbitLong/RepDistiller | https://arxiv.org/abs/1910.10699 TinyBERT | Embedding MSE + attention MSE + hidden MSE + prediction CE | Layer mapping g(m); term weights; two-stage schedule | 4-layer student, 312 hidden, uniform layer mapping | huawei-noah/TinyBERT | https://arxiv.org/abs/1909.10351 Sequence-level KD | NLL on the teacher's argmax (beam) output | Beam width for generation; whether to keep gold data | Beam 5; often mixed 50/50 with gold | OpenNMT | https://arxiv.org/abs/1606.07947 MiniLLM | Reverse KL, optimised on-policy with a policy-gradient estimator | rkl_advantage, single_step_decomposition, gamma, kd_temperature | TRL MiniLLMConfig: rkl_advantage True, gamma 0.0, temperature 1.0 | trl.experimental.minillm.MiniLLMTrainer | https://huggingface.co/docs/trl/main/minillm GKD | Generalized JSD, mixed on-policy/off-policy | lmbda (student-data fraction), beta (JSD interpolation), temperature, seq_kd | TRL GKDConfig: lmbda 0.5, beta 0.5, temperature 0.9 | trl.experimental.gkd.GKDTrainer | https://huggingface.co/docs/trl/gkd_trainer On-policy distillation (Tinker form) | Per-token reverse KL on student rollouts | Rollout length, teacher size | Qwen3-8B teacher, Qwen3-8B-Base student in the Thinking Machines run; the Qwen3 report's own run used a Qwen3-32B teacher. | tinker-cookbook train_on_policy.py | https://thinkingmachines.ai/blog/on-policy-distillation/ dDPO | Logistic loss on the implicit reward margin between chosen and rejected | beta (KL strength), reference model | Zephyr-7B used AI-ranked pairs and no human annotation | trl DPOTrainer | https://arxiv.org/abs/2310.16944 Margin-MSE (retrieval) | MSE between teacher and student score margins on (q, d+, d-) | Teacher ensemble size; negative sampling | 3 BERT-cat teachers on MS MARCO passage | sebastian-hofstaetter/neural-ranking-kd | https://arxiv.org/abs/2010.02666 Progressive distillation (diffusion) | Student one step matches teacher's two DDIM steps | Number of halving rounds N | Halve step count repeatedly, e.g. 1024 -> 4 | diffusers | https://arxiv.org/abs/2202.00512 Consistency distillation | Distance between f(x_{t+1}) and the EMA target at f(x_t) | EMA rate for the target network, discretisation schedule | OpenAI consistency_models repo settings | openai/consistency_models | https://arxiv.org/abs/2303.01469 LLM-QAT | KD on data generated by the model itself, with straight-through quantizers | Bit widths for W/A/KV-cache | Weights, activations and KV cache all quantized | facebookresearch/LLM-QAT | https://arxiv.org/abs/2305.17888 EAGLE | Regression on second-to-top-layer features + CE on tokens | Draft tree shape/depth; feature layer choice | One autoregressive draft head; EAGLE-3 fuses multiple layers | SafeAILab/EAGLE, vLLM, SGLang | https://arxiv.org/abs/2401.15077 ULD | Optimal-transport (Wasserstein) cost between sorted probability vectors | Truncation of the sorted vectors | No token alignment required | Nicolas-BZRD/llm-recipes | https://arxiv.org/abs/2402.12030 Notes: Formulas are written informally here; the exact LaTeX for each is in extras.methods[].lossFormula. TRL defaults quoted were read from the TRL v1.12.0 documented GKDConfig and MiniLLMConfig. Sources: https://huggingface.co/docs/trl/gkd_trainer https://huggingface.co/docs/trl/main/minillm https://pytorch.org/blog/llama-into-torchtune/ CHART DATA ---------- CHART: Six representative methods, six axes [id: radar-six-methods, type: radar, unit: score] Dimension | Response-based KD (soft targets) (score) Quality retention | 3 Data efficiency | 3 Implementation simplicity | 5 Works without teacher internals | 1 Low training compute | 4 Breadth of applicability | 4 Dimension | Black-box output distillation (score) Quality retention | 3 Data efficiency | 2 Implementation simplicity | 5 Works without teacher internals | 5 Low training compute | 4 Breadth of applicability | 5 Dimension | Reasoning-trace distillation (score) Quality retention | 5 Data efficiency | 4 Implementation simplicity | 4 Works without teacher internals | 5 Low training compute | 2 Breadth of applicability | 3 Dimension | Generalized KD (on-policy) (score) Quality retention | 5 Data efficiency | 5 Implementation simplicity | 2 Works without teacher internals | 1 Low training compute | 2 Breadth of applicability | 3 Dimension | Prune-then-distill (score) Quality retention | 4 Data efficiency | 5 Implementation simplicity | 1 Works without teacher internals | 1 Low training compute | 3 Breadth of applicability | 3 Dimension | Diffusion distillation (score) Quality retention | 4 Data efficiency | 5 Implementation simplicity | 2 Works without teacher internals | 1 Low training compute | 5 Breadth of applicability | 2 Notes: Editorial scores; the URLs listed under sources are the papers whose reported trade-offs these scores summarise, not sources for the numbers themselves. Editorial 1-5 ratings, not measured quantities. They summarise the trade-offs argued in each method's cited paper: 'Works without teacher internals' is 5 for black-box methods and 1 for white-box; 'Low training compute' is scored on the marginal cost of the student run, so reasoning-trace distillation scores low because generating long traces from a frontier teacher dominates the bill. Sources: https://arxiv.org/abs/1503.02531 https://arxiv.org/abs/2306.13649 https://arxiv.org/abs/2408.11796 https://arxiv.org/abs/2310.04378 https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B CHART: Distillation methods in this library, by year introduced [id: methods-per-year, type: bar, unit: methods] Year of the originating paper | Methods catalogued (methods) 2014 | 1 2015 | 1 2016 | 2 2017 | 2 2018 | 2 2019 | 4 2020 | 1 2021 | 1 2022 | 1 2023 | 8 2024 | 3 2025 | 1 Notes: Counts the 27 entries in extras.methods by the `year` field, which is the year of the originating arXiv preprint (not the conference year — FitNets is 2014 on arXiv, ICLR 2015). The 2023 spike is the instruction-tuning and reasoning-distillation wave; 2024-2025 counts are lower partly because recent work refines these families rather than founding new ones. Sources: https://arxiv.org/abs/2402.13116 https://arxiv.org/abs/2503.12067 CHART: Reasoning-trace distillation: AIME 2024 pass@1 by student size [id: r1-distill-aime, type: bar, unit: %] DeepSeek-R1-Distill student | AIME 2024 pass@1 (%) Qwen-1.5B | 28.9 Llama-8B | 50.4 Qwen-7B | 55.5 Qwen-14B | 69.7 Llama-70B | 70 Qwen-32B | 72.6 DeepSeek-R1-Distill student | MATH-500 pass@1 (%) Qwen-1.5B | 83.9 Llama-8B | 89.1 Qwen-7B | 92.8 Qwen-14B | 93.9 Llama-70B | 94.5 Qwen-32B | 94.3 Notes: All students are plain supervised fine-tunes on 800k reasoning samples curated with DeepSeek-R1 — no reinforcement-learning stage. Figures from the official model card. Sources: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B CHART: Off-policy distillation vs RL vs on-policy distillation (Qwen3-8B student) [id: onpolicy-vs-rl, type: bar, unit: %] Training method | AIME'24 (%) Off-policy distillation | 55 Reinforcement learning | 67.6 On-policy distillation | 74.4 Notes: The compute is the real story: the RL result cost 17,920 GPU-hours and the on-policy distillation result cost 1,800 GPU-hours, a roughly 10x reduction for a 6.8-point gain. Both figures come from the Qwen3 technical report (Table 21) and are quoted by Thinking Machines; Thinking Machines' own Tinker reproduction started from a different checkpoint and reached roughly 70% AIME'24. Sources: https://thinkingmachines.ai/blog/on-policy-distillation/ https://arxiv.org/html/2505.09388v1 CHART: Reported inference speedup of distilled students over their teachers [id: reported-speedups, type: bar, unit: x] Student (method) | Speedup (x) DistilBERT (response+cosine) | 1.6 Llama-3.1-Minitron-4B width | 1.8 Llama-3.1-Minitron-4B depth | 2.7 EAGLE draft on Llama2-Chat-70B | 3.1 TinyBERT-4L | 9.4 Kim & Rush NMT student | 10 Notes: DistilBERT is quoted as '60% faster', rendered here as 1.6x. EAGLE is reported as a 2.7x-3.5x latency speedup ratio on LLaMA2-Chat 70B; 3.1 is the midpoint and is the least precise bar here. These numbers come from different hardware, batch sizes and tasks and are not directly comparable to each other — read them as within-paper claims. Sources: https://arxiv.org/abs/1910.01108 https://arxiv.org/abs/1909.10351 https://arxiv.org/abs/2408.11796 https://arxiv.org/abs/2401.15077 https://arxiv.org/abs/1606.07947 THE METHODS (27) ---------------- One paragraph per method, in the order they appear in the library. Each entry gives the family, the year and paper it comes from, what the teacher must expose, what data it needs, its loss function in LaTeX, when to use it, and its honest trade-offs. The full multi-paragraph explanation for each method is on https://global-distillation.com/library and in data/library.json $.extras.methods[].howItWorks. ------------------------------------------------------------------------------------------------ Response-based KD (soft targets with temperature) [id: response-kd] Family: Response-based | Year: 2015 | Difficulty: 1/5 | Teacher access: white-box | Data needed: logits Paper: Distilling the Knowledge in a Neural Network — https://arxiv.org/abs/1503.02531 The founding method. Train the student to match the teacher's temperature-softened output distribution, optionally blended with ordinary cross-entropy against ground-truth labels. Loss (LaTeX): \mathcal{L} = (1-\alpha)\,\mathcal{H}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\,T^{2}\,\mathrm{KL}\big(\sigma(z_t/T)\ \|\ \sigma(z_s/T)\big),\qquad \sigma(z)_i = \frac{\exp(z_i/T)}{\sum_j \exp(z_j/T)} When to use: Use it first whenever you host both models and they share a tokenizer or label set. It is the right baseline for classifiers, encoders, and same-family LLM pairs (Llama 3.1 8B into Llama 3.2 1B), and the right first thing to try before reaching for on-policy methods. Pros: Simplest possible implementation: one extra forward pass and two lines of loss code; Teacher logits can be precomputed and cached, making the marginal training cost near zero; Architecture-agnostic — the student need share nothing with the teacher but the output space; Works with or without ground-truth labels, so it turns unlabelled data into training signal; Well-understood regularisation effect: soft targets reduce gradient variance and act like label smoothing with structure Cons: Requires teacher logits, so it is unavailable against any closed API; Requires an identical vocabulary or class set; Forward KL is mass-covering: a small student hedges across teacher modes instead of committing, which is exactly wrong for text generation; Off-policy — the student is never corrected on its own generated prefixes; Temperature and alpha need tuning per task and interact with the T² correction in ways that are easy to get wrong Tools: torchtune knowledge_distillation recipes; Hugging Face TRL; Arcee DistillKit (logit-based mode); NVIDIA TensorRT Model Optimizer; PyTorch (hand-rolled, ~10 lines) Examples: DistilBERT (with an added cosine-embedding term); Gemma 3 pretraining, sampling 256 logits per token weighted by teacher probability; torchtune's Llama 3.1 8B -> Llama 3.2 1B case study Related methods: feature-kd, reverse-kl-minillm, generalized-kd, self-distillation, cross-vocabulary-kd ------------------------------------------------------------------------------------------------ Feature / intermediate-layer KD (FitNets hints) [id: feature-kd] Family: Feature-based | Year: 2014 | Difficulty: 3/5 | Teacher access: white-box | Data needed: features Paper: FitNets: Hints for Thin Deep Nets — https://arxiv.org/abs/1412.6550 Supervise a student's hidden layer directly against a teacher's hidden layer, using a learned projector to bridge the width mismatch, then finish with ordinary response-based KD. Loss (LaTeX): \mathcal{L}_{\text{hint}} = \tfrac{1}{2}\Big\| \, r\big(F_s^{\,h};\,W_r\big) \;-\; F_t^{\,g} \Big\|_2^2 \qquad\text{then}\qquad \mathcal{L}_{\text{KD}} = (1-\alpha)\mathcal{H}(y,\sigma(z_s)) + \alpha T^2\mathrm{KL}(\sigma(z_t/T)\|\sigma(z_s/T)) When to use: Use when the student is much deeper or narrower than the teacher and logit-only KD is failing to converge, and when teacher and student are architecturally similar enough that a layer correspondence is meaningful. In transformer land, prefer the structured version (TinyBERT) over hand-picked hint layers. Pros: Gives dense supervision deep inside the student, which unlocks thin-and-deep architectures that pure logit KD cannot train; Substantially more information transferred than logits alone — features are high-dimensional; Composes cleanly with response-based KD as a two-stage schedule; Effective when teacher and student share an inductive bias (both CNNs, both transformers) Cons: Requires full white-box access to activations, not just logits; Layer-pairing is a hyperparameter with no good default and large effect; The projector adds parameters and its own optimisation dynamics; Poor across architecture families — forcing a transformer student to match CNN features rarely helps; Raw-feature MSE over-constrains: it demands the student reproduce the teacher's coordinate system, not just its information Tools: RepDistiller; Arcee DistillKit (hidden-states mode); torchvision + custom hooks; NVIDIA Model Optimizer (intermediate-state distillation) Examples: FitNets on CIFAR-10/100 and SVHN; TinyBERT's hidden-state loss; Minitron's intermediate-state distillation during pruning recovery Related methods: response-kd, attention-transfer, relation-kd, transformer-layer-kd, contrastive-rep-kd ------------------------------------------------------------------------------------------------ Attention transfer [id: attention-transfer] Family: Attention-based | Year: 2016 | Difficulty: 2/5 | Teacher access: white-box | Data needed: features Paper: Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer — https://arxiv.org/abs/1612.03928 Instead of matching full activation tensors, match a cheap 2-D summary of where each layer is 'looking' — the channel-collapsed activation energy map — after L2 normalisation. Loss (LaTeX): \mathcal{L}_{AT} = \sum_{j\in\mathcal{I}} \left\| \frac{Q_s^{\,j}}{\|Q_s^{\,j}\|_2} - \frac{Q_t^{\,j}}{\|Q_t^{\,j}\|_2} \right\|_p ,\qquad Q^{\,j} = \mathrm{vec}\Big(\sum_{i=1}^{C} |A_i^{\,j}|^{\,p}\Big) When to use: Use as an add-on term when teacher and student are both convolutional or both transformer but differ in width, and when FitNets-style feature matching is proving brittle. In NLP, use it as one component of a TinyBERT-style multi-loss recipe rather than on its own. Pros: Channel-count agnostic, so no projector is needed and width mismatch is free; Cheap: the summary map is H x W, orders of magnitude smaller than the activation tensor; More robust to architecture differences than raw-feature matching; Consistent gains reported across several datasets and CNN families; Transfers an interpretable quantity — you can visualise exactly what is being copied Cons: Lossy: discards all channel-identity information; Weaker standalone effect than logit or feature KD; usually needs to be combined; Still white-box, still needs a layer correspondence (though a coarser one); The power p and the layer-group set are extra hyperparameters; For transformers, per-head attention matching can force spurious head alignment when head counts differ Tools: szagoruyko/attention-transfer; RepDistiller; TinyBERT (attention-matrix MSE); custom forward hooks Examples: Wide ResNet teacher to thin ResNet student on CIFAR/ImageNet; TinyBERT attention distillation; Cross-modal attention transfer in video models Related methods: feature-kd, transformer-layer-kd, relation-kd, response-kd ------------------------------------------------------------------------------------------------ Relational KD (RKD) [id: relation-kd] Family: Relation-based | Year: 2019 | Difficulty: 3/5 | Teacher access: white-box | Data needed: features Paper: Relational Knowledge Distillation — https://arxiv.org/abs/1904.05068 Transfer the geometry *between* examples — pairwise distances and triplet angles in embedding space — rather than the representation of any single example. Loss (LaTeX): \mathcal{L}_{\text{RKD-D}} = \sum_{(i,j)} l_{\delta}\Big(\psi_D(t_i,t_j),\ \psi_D(s_i,s_j)\Big),\quad \psi_D(a,b) = \frac{\|a-b\|_2}{\mu} \qquad\text{and}\qquad \mathcal{L}_{\text{RKD-A}} = \sum_{(i,j,k)} l_{\delta}\Big(\psi_A(t_i,t_j,t_k),\ \psi_A(s_i,s_j,s_k)\Big),\quad \psi_A = \cos\angle\, t_i t_j t_k When to use: Use for embedding models: image retrieval, face and person re-identification, recommendation towers, and sentence encoders — anywhere the deployed operation is a nearest-neighbour search rather than an argmax over classes. Pros: Invariant to rotation, reflection and scaling of the embedding space — no coordinate-frame tax; Works when teacher and student embedding dimensions differ, with no projector; Students can exceed their teachers in metric-learning settings; Directly optimises the property that retrieval and clustering actually consume; Composable with logit KD as an extra term Cons: O(B²) and O(B³) terms make large batches costly and small batches statistically noisy; Needs a meaningful embedding layer, so it does not apply to bare classifiers without one; Two loss weights to balance against each other and against the task loss; Less effective than logit KD for plain classification, where absolute class scores are the target; Still white-box Tools: RepDistiller; sentence-transformers (custom losses); PyTorch Metric Learning Examples: Metric learning on CUB-200, Cars-196 and Stanford Online Products; Face-recognition backbone compression; Embedding-tower distillation in retrieval stacks Related methods: contrastive-rep-kd, feature-kd, retriever-distillation, attention-transfer ------------------------------------------------------------------------------------------------ Contrastive representation distillation (CRD) [id: contrastive-rep-kd] Family: Relation-based | Year: 2019 | Difficulty: 4/5 | Teacher access: white-box | Data needed: features Paper: Contrastive Representation Distillation — https://arxiv.org/abs/1910.10699 Reframe distillation as maximising the mutual information between teacher and student representations, optimised with a contrastive (InfoNCE-style) objective over positive and negative pairs. Loss (LaTeX): \mathcal{L}_{\text{CRD}} = -\,\mathbb{E}_{q(T,S\mid C=1)}\big[\log h(T,S)\big] \;-\; N\,\mathbb{E}_{q(T,S\mid C=0)}\big[\log\!\big(1-h(T,S)\big)\big],\qquad h(T,S)=\frac{e^{\,g_T(T)^\top g_S(S)/\tau}}{e^{\,g_T(T)^\top g_S(S)/\tau} + \tfrac{N}{M}} When to use: Use when you have exhausted logit and simple feature methods and the representation quality itself is the deliverable — cross-modal transfer, ensemble compression, or a backbone that will feed many downstream heads. Pros: Captures higher-order structure and correlations that per-dimension KL discards; Principled: an explicit lower bound on mutual information; Works cross-modally and for ensemble distillation, not just same-modality compression; Strong empirical results; frequently the best single feature-family method in benchmark sweeps; Stacks with response-based KD for further gains Cons: Heaviest implementation in the feature family: memory bank, critic, projection heads; Sensitive to the number of negatives and the temperature of the critic; Extra memory footprint during training; Gains over simpler methods are modest relative to the added complexity for straightforward compression tasks; White-box only Tools: HobbitLong/RepDistiller; PyTorch (custom InfoNCE + memory bank) Examples: CIFAR-100 and ImageNet compression benchmarks; RGB-to-depth cross-modal transfer; Ensemble-into-single-model distillation Related methods: relation-kd, feature-kd, response-kd, retriever-distillation ------------------------------------------------------------------------------------------------ Transformer-layer KD (TinyBERT / DistilBERT) [id: transformer-layer-kd] Family: Feature-based | Year: 2019 | Difficulty: 4/5 | Teacher access: white-box | Data needed: features Paper: TinyBERT: Distilling BERT for Natural Language Understanding — https://arxiv.org/abs/1909.10351 The standard recipe for compressing encoder transformers: simultaneously match embeddings, attention matrices, hidden states and prediction logits, applied at both the pretraining and task-specific stages. Loss (LaTeX): \mathcal{L} = \sum_{m} \lambda_m \Big[ \underbrace{\mathrm{MSE}\big(A_s^{m},\,A_t^{\,g(m)}\big)}_{\text{attention}} + \underbrace{\mathrm{MSE}\big(H_s^{m}W_h,\,H_t^{\,g(m)}\big)}_{\text{hidden}} \Big] \;+\; \underbrace{\mathrm{MSE}\big(E_sW_e,\,E_t\big)}_{\text{embedding}} \;+\; \underbrace{\mathrm{CE}\big(\sigma(z_t/T),\,\sigma(z_s/T)\big)}_{\text{prediction}} When to use: Use for encoder models you will deploy at scale — classification, NER, reranking, embeddings — where latency and cost matter and you can afford a two-stage training pipeline. For a quicker win, take DistilBERT's simpler triple loss instead. Pros: Best-validated recipe for encoder compression, with a decade of production use behind it; Very large speedups at small accuracy cost: 9.4x faster at 96.8% of GLUE for TinyBERT-4L; Attention matching transfers linguistic structure that logits alone do not; Two-stage schedule yields a reusable general student plus a task-tuned one; Reference implementations and pretrained students are widely available Cons: Many interacting hyperparameters: layer mapping, four loss weights, two projectors, two stages; Requires a large unlabelled corpus for the general stage; Attention matching assumes comparable head structure between teacher and student; Less used for modern decoder-only LLMs, where on-policy logit methods dominate; Expensive: two full distillation passes plus data augmentation Tools: huawei-noah/TinyBERT; Hugging Face transformers (DistilBERT); Arcee DistillKit (hidden-states mode); Sentence Transformers Examples: TinyBERT-4L and -6L on GLUE; DistilBERT, DistilRoBERTa, DistilGPT-2; MobileBERT and MiniLM as descendants of the same recipe Related methods: feature-kd, attention-transfer, response-kd, retriever-distillation, prune-then-distill ------------------------------------------------------------------------------------------------ Sequence-level KD (Kim & Rush) [id: sequence-level-kd] Family: Sequence-level / data | Year: 2016 | Difficulty: 1/5 | Teacher access: black-box | Data needed: outputs Paper: Sequence-Level Knowledge Distillation — https://arxiv.org/abs/1606.07947 Replace token-level distribution matching with plain maximum-likelihood training on complete sequences generated by the teacher, typically its beam-search output. Loss (LaTeX): \mathcal{L}_{\text{SeqKD}} = -\sum_{y\in\mathcal{Y}} q(y\mid x)\,\log p_\theta(y\mid x) \;\approx\; -\log p_\theta(\hat{y}\mid x), \qquad \hat{y} = \arg\max_{y}\, q(y\mid x) \ \ \text{(teacher beam search)} When to use: Use it as the default whenever the teacher is behind an API, or whenever you want a reusable, auditable training set. It is also the correct first step before any on-policy method, since it gives a strong initialisation cheaply. Pros: Black-box: needs only sampled teacher text, so it works against any API; Trivial to implement — after generation it is ordinary supervised fine-tuning; Teacher outputs are self-consistent, which removes reference ambiguity and helps small students disproportionately; Often removes the need for beam search at inference, compounding the speedup; Generation is a one-time cost and the resulting dataset is reusable and inspectable Cons: Off-policy: the student never sees its own generated prefixes, so exposure bias remains; Approximating the teacher distribution by its mode discards all diversity and uncertainty; Inherits and can amplify teacher errors, since there is no label to correct them; Generation cost scales with corpus size and output length; Against a commercial API this is the exact activity most terms of service prohibit for competing-model training Tools: OpenNMT; Hugging Face TRL (GKDConfig seq_kd=True); vLLM / SGLang for bulk generation; OpenAI Model Distillation (Stored Completions); Amazon Bedrock Model Distillation Examples: WMT English-German NMT students 10x faster than the teacher; Every Alpaca-lineage instruction dataset; Bedrock and OpenAI managed distillation pipelines Related methods: blackbox-output-distillation, cot-rationale-distillation, reasoning-trace-distillation, generalized-kd, response-kd ------------------------------------------------------------------------------------------------ Black-box output distillation (Alpaca / Vicuna / Orca style) [id: blackbox-output-distillation] Family: Sequence-level / data | Year: 2023 | Difficulty: 1/5 | Teacher access: black-box | Data needed: outputs Paper: Alpaca: A Strong, Replicable Instruction-Following Model; Orca: Progressive Learning from Complex Explanation Traces of GPT-4 — https://crfm.stanford.edu/2023/03/13/alpaca.html Generate a synthetic instruction-following dataset by prompting a strong API teacher, then supervised-fine-tune an open base model on it. The dominant form of distillation actually practised. Loss (LaTeX): \mathcal{L} = -\,\mathbb{E}_{(x,y)\sim\mathcal{D}_{\text{teacher}}}\ \sum_{t=1}^{|y|} \log p_\theta\big(y_t \mid x,\ y_{ mid -> small); Orca's use of ChatGPT as teacher assistant alongside GPT-4 Related methods: response-kd, reverse-kl-minillm, multi-teacher-kd, self-distillation, prune-then-distill ------------------------------------------------------------------------------------------------ Multi-teacher KD [id: multi-teacher-kd] Family: Response-based | Year: 2017 | Difficulty: 3/5 | Teacher access: white-box | Data needed: logits Paper: Efficient Knowledge Distillation from an Ensemble of Teachers (Fukuda et al., Interspeech 2017); weighting variants surveyed in Knowledge Distillation: A Survey (Gou et al., 2020, arXiv 2006.05525) — https://www.isca-archive.org/interspeech_2017/fukuda17_interspeech.html Distil from several teachers at once, combining their output distributions (or features) into a single target for the student. Loss (LaTeX): \mathcal{L} = T^2\,\mathrm{KL}\Big(\textstyle\sum_{k=1}^{K} w_k\,\sigma\!\big(z_{t_k}/T\big)\ \Big\|\ \sigma\!\big(z_s/T\big)\Big) + (1-\alpha)\mathcal{H}(y,\sigma(z_s)),\qquad \textstyle\sum_k w_k = 1,\ w_k \ \text{uniform or adaptive} When to use: Use when you have several strong models with complementary strengths and want one deployable student. Prefer the black-box best-of-K data variant with a verifier or judge unless you specifically need per-token logit blending. Pros: Captures ensemble-level accuracy at single-model inference cost; Averaging reduces the variance and idiosyncratic bias of any single teacher; Lets you combine specialists — a code teacher, a maths teacher, a safety teacher; The black-box data variant needs no tokenizer alignment and works with API teachers; Teacher logits can be precomputed once and reused across student runs Cons: Uniform averaging destroys the diversity that motivated the ensemble; Conflicting teachers produce compromise targets that can be worse than any teacher; K forward passes or K cached logit sets: K times the compute or storage; Adaptive weighting adds real complexity and its own hyperparameters; White-box form requires a shared vocabulary across all teachers Tools: Custom PyTorch loops; Arcee DistillKit; vLLM for multi-model bulk generation; LLM-as-judge pipelines for candidate selection Examples: Acoustic-model ensembles distilled into one recogniser; Margin-MSE's 3-teacher BERT-cat ensemble for retrieval; Open instruction datasets mixing outputs from several frontier models Related methods: response-kd, deep-mutual-learning, teacher-assistant-kd, retriever-distillation, blackbox-output-distillation ------------------------------------------------------------------------------------------------ Deep mutual learning (online distillation) [id: deep-mutual-learning] Family: Response-based | Year: 2017 | Difficulty: 2/5 | Teacher access: none | Data needed: logits Paper: Deep Mutual Learning — https://arxiv.org/abs/1706.00384 Train a cohort of untrained peer networks simultaneously, each learning from the labels and from the others' predictions. No pretrained teacher exists at any point — every peer is white-box to every other. Loss (LaTeX): \mathcal{L}_{\theta_1} = \mathcal{H}\big(y,\sigma(z_1)\big) \;+\; \frac{1}{K-1}\sum_{k\neq 1} \mathrm{KL}\big(\sigma(z_k)\ \|\ \sigma(z_1)\big) \qquad\text{(symmetrically for every peer } k) When to use: Use when you must train from scratch in a domain with no strong pretrained teacher — specialised scientific, industrial or medical models — and you can afford to train a small cohort. Also useful when you want a large and a small model co-trained in one run. Pros: No pretrained teacher required — usable when none exists for your domain; Outperforms distillation from a stronger static teacher in the original experiments; Peers may have different architectures, so you can co-train a deployable small model with a large one; Produces an ensemble as a by-product if you want to keep all peers; Empirically finds flatter, better-generalising minima Cons: Trains K models to deploy one — K times the compute and memory; Cohorts can collectively converge on a shared error with no external correction; Cohort size and mimicry weight are extra hyperparameters; Harder to reason about and debug than a fixed teacher-student pipeline; Rarely used at LLM scale, where training two frontier models in lockstep is impractical Tools: Custom PyTorch training loops; timm (multi-model harness); MMPretrain / mmrazor Examples: CIFAR-100 and Market-1501 person re-identification in the original paper; Online distillation variants in production vision stacks Related methods: self-distillation, multi-teacher-kd, response-kd, teacher-assistant-kd ------------------------------------------------------------------------------------------------ Pruning + distillation (Minitron, Sheared LLaMA) [id: prune-then-distill] Family: Compression-composed | Year: 2023 | Difficulty: 5/5 | Teacher access: white-box | Data needed: logits Paper: Compact Language Models via Pruning and Knowledge Distillation; LLM Pruning and Distillation in Practice: The Minitron Approach; Sheared LLaMA — https://arxiv.org/abs/2407.14679 Structurally prune a large model down to a target shape, then use distillation from the unpruned original to recover the lost accuracy on a small fraction of the original token budget. Loss (LaTeX): \hat{M} = \mathrm{Prune}\big(M;\ \mathcal{I}(\text{layers, heads, neurons, channels})\big)\ \ \text{then}\ \ \mathcal{L} = \mathrm{KL}\big(p_{M}\,\|\,p_{\hat{M}}\big) + \sum_m \gamma_m\,\mathrm{MSE}\big(H^{m}_{\hat{M}},\,H^{m}_{M}\big) When to use: Use when you own or can license the teacher weights and need a specific smaller architecture for a hard latency or memory budget. This is how you build a model family from a single large model rather than training each size independently. Pros: Retains the teacher's learned weights instead of starting over — the reason it needs under 1% of the original tokens; Real deployment wins: 1.8x-2.7x measured speedups for Llama-3.1-Minitron-4B; Lets you hit an exact target architecture and latency budget; Sheared LLaMA students beat same-size models trained from scratch; Open weights and reference recipes exist for both lineages Cons: The most complex pipeline here: importance scoring, pruning, teacher correction, distillation retraining; Requires full teacher weights and a large re-training corpus; Aggressive pruning can remove narrow capabilities that aggregate benchmarks will not reveal; Depth versus width is a genuine trade (latency versus quality) with no universal answer; Needs substantial GPU memory: teacher and student resident together over billions of tokens Tools: NVIDIA TensorRT Model Optimizer; NVIDIA NeMo (pruning + distillation recipes); princeton-nlp/LLM-Shearing; torch.nn.utils.prune for the basics Examples: Llama-3.1-Minitron-4B (width and depth variants); Mistral-NeMo-Minitron-8B from Mistral NeMo 12B; Sheared-LLaMA-1.3B and -2.7B from LLaMA2-7B; Llama 3.2 1B/3B, which used logits from Llama 3.1 8B and 70B to recover after pruning Related methods: quantization-aware-distillation, response-kd, feature-kd, teacher-assistant-kd, generalized-kd ------------------------------------------------------------------------------------------------ Quantization-aware distillation (QAD / LLM-QAT) [id: quantization-aware-distillation] Family: Compression-composed | Year: 2023 | Difficulty: 4/5 | Teacher access: white-box | Data needed: logits Paper: LLM-QAT: Data-Free Quantization Aware Training for Large Language Models — https://arxiv.org/abs/2305.17888 Train a low-bit student with quantization simulated in the loop while distilling from the full-precision original, recovering the accuracy that post-training quantization loses. Loss (LaTeX): \mathcal{L} = \mathrm{KL}\big(p_{\mathrm{fp}}(\cdot\mid x)\ \big\|\ p_{Q(\theta)}(\cdot\mid x)\big),\qquad Q(w) = s\cdot\mathrm{clip}\Big(\Big\lfloor \tfrac{w}{s} \Big\rceil,\ -2^{\,b-1},\ 2^{\,b-1}-1\Big),\qquad \frac{\partial Q}{\partial w}\ \approx\ \mathbb{1}_{|w|\le \tau}\ \ \text{(STE)} When to use: Use as the last step before deployment when you are quantizing below 8 bits and post-training quantization has cost you more accuracy than you can accept, particularly on edge hardware or long-context serving where the KV cache dominates. Pros: Recovers most of the accuracy lost by post-training quantization, especially below 8 bits; Data-free variant needs no access to the original training corpus; Quantizing the KV cache unlocks long-context throughput that weight-only quantization does not; Applies to any generative model independent of its training data; Stacks directly on top of pruning-and-distillation as the final compression step Cons: Training with fake-quantize ops is slower per step than normal training; Straight-through gradients are biased, so optimisation is noisier; 2-3 bit regimes are still an open problem despite steady progress; The realised speedup depends on kernel support for the chosen format, not on the bit-width alone; Recovers accuracy but teaches nothing new — it cannot fix a weak model Tools: NVIDIA TensorRT Model Optimizer; facebookresearch/LLM-QAT; torchao / torch.ao.quantization; torchtune QAT recipes; Intel Neural Compressor Examples: LLM-QAT 4-bit weight/activation/KV-cache LLaMA models; NVIDIA NVFP4 accuracy-recovery pipelines; On-device deployment of Minitron- and Gemma-class students Related methods: prune-then-distill, response-kd, self-distillation, draft-model-distillation ------------------------------------------------------------------------------------------------ Dataset distillation / condensation [id: dataset-distillation] Family: Dataset / context | Year: 2018 | Difficulty: 5/5 | Teacher access: none | Data needed: none Paper: Dataset Distillation — https://arxiv.org/abs/1811.10959 Hold the model fixed and compress the *training set* into a tiny synthetic set that trains a network to comparable accuracy. The data, not the network, is the thing distilled. Loss (LaTeX): \tilde{\mathcal{D}}^{*} = \arg\min_{\tilde{\mathcal{D}}}\ \mathcal{L}\big(\mathcal{D};\ \theta_1(\tilde{\mathcal{D}})\big) \qquad\text{s.t.}\qquad \theta_1(\tilde{\mathcal{D}}) = \theta_0 - \eta\,\nabla_{\theta}\,\mathcal{L}\big(\tilde{\mathcal{D}};\,\theta_0\big),\quad \theta_0\sim p(\theta_0) When to use: Use for neural architecture search proxies, continual-learning replay buffers, and federated or data-constrained settings — not as a route to a better production model. Treat it as data-axis compression, orthogonal to every other method here. Pros: Extreme compression of the data axis — orders of magnitude fewer examples; Makes architecture search and hyperparameter sweeps dramatically cheaper via proxy datasets; Natural fit for continual-learning replay buffers and federated settings; Synthetic examples are not literal records, which helps (informally) with data-sharing constraints; Reusable: once condensed, the set trains many models Cons: Bi-level optimisation is expensive and memory-hungry; differentiating through training steps does not scale naively; Distilled sets are often architecture-specific and transfer poorly; Scaling beyond small image benchmarks remains an open problem; Synthetic examples are uninterpretable, so you cannot audit what the set actually encodes; Privacy protection is empirical, not a formal guarantee Tools: VICO-UoE/DatasetCondensation; DC-BENCH; torchvision + custom bi-level loops Examples: MNIST and CIFAR-10 condensed to a handful of images per class; Proxy datasets for architecture search; Condensed replay buffers in continual learning Related methods: context-distillation, self-distillation, blackbox-output-distillation ------------------------------------------------------------------------------------------------ Preference distillation / RLAIF [id: preference-distillation-rlaif] Family: Preference | Year: 2023 | Difficulty: 4/5 | Teacher access: black-box | Data needed: outputs Paper: Constitutional AI: Harmlessness from AI Feedback; Zephyr: Direct Distillation of LM Alignment — https://arxiv.org/abs/2212.08073 Distil judgement rather than answers: a teacher model ranks candidate completions, those rankings train a reward model, and the student is optimised against it — human annotation replaced by AI feedback. Loss (LaTeX): r_\phi \leftarrow \arg\max_\phi \sum \log\sigma\big(r_\phi(x,y_w) - r_\phi(x,y_l)\big),\ \ (y_w,y_l)\ \text{ranked by the teacher};\qquad \max_\theta\ \mathbb{E}_{y\sim\pi_\theta}\big[r_\phi(x,y)\big] - \beta\,\mathrm{KL}\big(\pi_\theta\,\|\,\pi_{\text{ref}}\big) When to use: Use for alignment, tone, helpfulness and safety behaviours where 'better' is a judgement call rather than a verifiable fact. If you do not need an explicit reward model, go straight to dDPO — it is far simpler for most of the benefit. Pros: Removes the human-annotation bottleneck, so preference data scales with compute; Ranking is an easier task than generation, so teacher preferences are often more reliable than teacher completions; The governing principles are explicit and auditable, unlike implicit annotator preferences; Black-box: only the teacher's judgements are needed; Zephyr-7B beat Llama2-Chat-70B on MT-Bench with no human labels Cons: The student inherits the teacher's values, biases and blind spots without correction; Reward models are exploitable; policies find degenerate high-reward behaviours; Full RLAIF with PPO is operationally heavy — reward model, policy, reference and value models in memory; Judge models exhibit known artefacts such as position and verbosity bias; Requires a KL anchor to a reference model or the policy drifts off distribution Tools: Hugging Face TRL (RewardTrainer, PPOTrainer, GRPOTrainer); OpenRLHF; NVIDIA NeMo Aligner; alignment-handbook; LLM-as-judge harnesses Examples: Anthropic Constitutional AI / RLAIF; Zephyr-7B's AI-ranked preference stage; UltraFeedback and the open AI-feedback dataset ecosystem Related methods: dpo-teacher-preferences, blackbox-output-distillation, reasoning-trace-distillation, reverse-kl-minillm ------------------------------------------------------------------------------------------------ DPO on teacher preferences (dDPO) [id: dpo-teacher-preferences] Family: Preference | Year: 2023 | Difficulty: 2/5 | Teacher access: black-box | Data needed: outputs Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model; Zephyr: Direct Distillation of LM Alignment — https://arxiv.org/abs/2305.18290 Skip the reward model and the RL loop: optimise the student directly on teacher-ranked preference pairs with a simple classification-style loss. Loss (LaTeX): \mathcal{L}_{\text{DPO}} = -\,\mathbb{E}_{(x,y_w,y_l)\sim\mathcal{D}_{\text{teacher}}}\left[\log \sigma\!\left(\beta\log\frac{\pi_\theta(y_w\mid x)}{\pi_{\text{ref}}(y_w\mid x)} \;-\; \beta\log\frac{\pi_\theta(y_l\mid x)}{\pi_{\text{ref}}(y_l\mid x)}\right)\right] When to use: Use as the default alignment stage after any SFT-based distillation. If you cannot afford RLAIF's infrastructure — which is almost everyone — this captures most of the benefit. Escalate to online preference methods only if you can show DPO is plateauing. Pros: No reward model, no value model, no rollouts — a plain supervised loss over a static dataset; Trains in hours on hardware where PPO would not fit; Zephyr-7B beat Llama2-Chat-70B on MT-Bench using only AI-ranked preferences; Black-box: needs teacher rankings, not teacher internals; First-class support in TRL and the alignment-handbook, with well-trodden recipes Cons: Off-policy — cannot correct behaviours that emerge during training but are absent from the pair set; Sensitive to beta and to the choice of reference model; Can lower the chosen response's probability as long as the rejected one falls faster, degrading both; Inherits every bias of the ranking teacher; Needs a good SFT checkpoint first; DPO on a weak base is unreliable Tools: Hugging Face TRL DPOTrainer; alignment-handbook; Axolotl; LLaMA-Factory; OpenRLHF Examples: Zephyr-7B-beta (dDPO on GPT-4-ranked UltraFeedback); Tulu and Nous-family preference-tuned models; Most open 7B-70B chat models released since late 2023 Related methods: preference-distillation-rlaif, blackbox-output-distillation, reasoning-trace-distillation, sequence-level-kd ------------------------------------------------------------------------------------------------ Cross-vocabulary logit KD (ULD, DSKD) [id: cross-vocabulary-kd] Family: Response-based | Year: 2024 | Difficulty: 5/5 | Teacher access: white-box | Data needed: logits Paper: Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs; Dual-Space Knowledge Distillation for Large Language Models — https://arxiv.org/abs/2402.12030 Make white-box logit distillation work when teacher and student have different tokenizers, either by comparing sorted probability vectors with an optimal-transport cost or by projecting both models into a shared output space. Loss (LaTeX): \mathcal{L}_{\text{ULD}} = \sum_{t=1}^{T} \mathcal{W}_1\Big(\mathrm{sort}\big(p_t^{\text{teacher}}\big),\ \mathrm{sort}\big(q_t^{\text{student}}\big)\Big) \qquad\text{vs.}\qquad \mathcal{L}_{\text{DSKD}} = \mathrm{KL}\Big(\mathcal{P}_{\text{shared}}\big(p_t\big)\ \Big\|\ \mathcal{P}_{\text{shared}}\big(q_t\big)\Big) When to use: Use when the best available teacher is in a different model family from your required student and you host both sets of weights. If you do not strictly need per-token signal, black-box sequence-level distillation is simpler and more predictable. Pros: Unlocks white-box distillation across model families — a Llama teacher into a Qwen student; ULD needs no token alignment at all and tolerates different vocabulary sizes; DSKD is a general framework covering same-tokenizer and cross-tokenizer cases with one interface; Much denser signal than falling back to black-box sequence-level KD; Lets you pick the best available teacher rather than the best same-family teacher Cons: ULD discards token identity, so some information is provably lost; DSKD's cross-model attention adds trainable parameters and optimisation complexity; Both are consistently reported as less reliable than same-tokenizer KD; Little first-class support in mainstream training libraries; expect to work from the reference repos; Optimal-transport computation adds per-step cost Tools: Nicolas-BZRD/llm-recipes (ULD); songmzhang/DSKD and DSKDv2; Custom TRL trainers Examples: Cross-family distillation experiments in the ULD paper; DSKD's same- and cross-vocabulary LLM benchmarks; Multi-Level Optimal Transport and later cross-tokenizer variants Related methods: response-kd, generalized-kd, sequence-level-kd, multi-teacher-kd ------------------------------------------------------------------------------------------------ Draft-model distillation for speculative decoding (EAGLE, Medusa) [id: draft-model-distillation] Family: Trajectory / sampler | Year: 2024 | Difficulty: 4/5 | Teacher access: white-box | Data needed: features Paper: EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty; Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — https://arxiv.org/abs/2401.15077 Distil a tiny draft head from a frozen target model so that speculative decoding accepts more proposed tokens. The target model's output distribution is provably unchanged; only latency falls. Loss (LaTeX): \mathcal{L} = \underbrace{\mathrm{SmoothL1}\big(f_{\text{draft}}(h_{