A public compendium · updated daily
How frontier intelligence is compressed, priced, and contested.
Distillation moves capability from a large teacher model into a small, cheap student. It is the quiet engine behind most models people actually pay for, the subject of an open dispute between the largest labs, and a live regulatory question in three jurisdictions. This compendium tracks all of it.
Financial · compiled 3 September 2026 · 63 sources
The economics of AI distillation: prices, training costs, and the market shocks
Distillation is, at bottom, an arbitrage: the capability embedded in a $40M-$500M frontier training run can be harvested through an API for a four- or five-figure query bill and re-trained into a small model for hundreds of dollars. That asymmetry is now visible in published price lists, where the gap between a vendor's flagship and its small tier runs 10x to 42x on output tokens, and in the research record, where Sky-T1-32B was distilled for under $450 and s1-32B for a reported ~$50. It became a macro event on 27 January 2025, when DeepSeek-R1 and its six open-weight distilled students wiped $589B off Nvidia's market capitalisation in a single session - the largest one-day loss in stock-market history - and it became a policy event in February 2026 when OpenAI told the House Select Committee on China that DeepSeek was running obfuscated distillation pipelines against its models. The counter-trend matters too: Epoch AI measures inference prices for a fixed capability level falling 9x-900x per year, which compresses the payback on self-hosting a distilled model to the point where, on our own TCO model, a single-GPU deployment only beats Claude Haiku 4.5 above ~658M output tokens a month and never beats the cheapest serverless open-model endpoints. And in August 2026 DeepSeek reversed the race to zero, raising V4 API prices by as much as 1,100% and introducing peak/off-peak rates - the first major signal that ultra-cheap distilled inference was being priced against capacity, not against marginal cost.
Key figures · 8 figures
Nvidia single-day market-cap loss
589 USD billions
-17% in one session
27 Jan 2025, after DeepSeek-R1 and its distilled students shipped. Largest one-day loss in US market history.
Cost to train Sky-T1-32B by distillation
450 USD
8x H100 for 19 hours
Matches o1-preview on Math500 and AIME24; teacher was QwQ-32B-Preview, base was Qwen2.5-32B-Instruct.
DeepSeek-R1 reinforcement-learning training cost
294,000 USD
512 H800s x 80 hours
Disclosed in the peer-reviewed Nature paper, Sept 2025. Excludes the ~$5.6M base-model run and all R&D, data and infrastructure.
Frontier-vs-small output price spread, OpenAI
42 x
$50 vs $1.20 per MTok
gpt-6-astra output $50/MTok against gpt-5.6-luna at $1.20/MTok on the same list.
Inference price decline for fixed capability
900 x per year (upper bound)
range 9x-900x
Epoch AI, measuring the cheapest model clearing a fixed benchmark threshold. GPT-4-level fell from $37.50/MTok (Mar 2023) to $0.18/MTok (Feb 2025).
Frontier training-run cost growth
2.4 x per year
95% CI 2.0x-3.1x, since 2016
Epoch AI, amortised hardware + energy across 45 frontier models; cloud-rental method gives 2.6x/yr.
Teacher-query bill to rebuild an 800k-trace reasoning corpus
40,000 USD (upper end)
$1,920 at the cheap end
Author calculation: 800k traces x 2,000 output tokens = 1.6B output tokens, priced at 2026 list rates from Claude Opus 5 ($25/MTok) down to gpt-5.6-luna ($1.20/MTok).
DeepSeek V4 API price increase
1,100 % (maximum)
effective 16 Aug 2026
Cache-hit input tokens rose up to 1,100%; output tokens 127%-371%. Peak/off-peak schedule introduced for the first time.
Key findings · 10 findings
The distillation arbitrage is a five-order-of-magnitude gap, and it is widening
Epoch AI puts the final training run of GPT-4 at roughly $40M on an amortised-hardware basis (the Stanford AI Index puts it at $78M on cloud-rental accounting), and frontier run costs have grown 2.4x per year since 2016. Against that, Sky-T1-32B was distilled for under $450 of GPU time and TinyZero reproduced R1-Zero-style behaviour for under $30. Even the teacher-query cost is small: at 2026 list prices, regenerating an 800k-trace reasoning corpus like the one behind DeepSeek's R1-Distill family costs $1,920 (gpt-5.6-luna) to $40,000 (Claude Opus 5). The ratio between building the capability and copying it is roughly 1,000:1 to 100,000:1.
Every major vendor now sells its own distillation discount as a product tier
The frontier-to-small spread on output tokens is 10x at Anthropic (Claude Fable 5.1 at $50/MTok vs Haiku 4.5 at $5/MTok), 42x at OpenAI (gpt-6-astra $50 vs gpt-5.6-luna $1.20), 30x at Google (Gemini 3.1 Pro $12 vs Gemini 2.5 Flash-Lite $0.40) and 75x at Mistral (Medium 3.5 at $7.50 vs Ministral 3 3B at $0.10). Vendors capture the distillation margin internally rather than losing it to third parties. The commercial logic is that a customer who would otherwise self-distill can be retained at a price point the vendor still profits at.
Self-hosting a distilled model is now the expensive option for most buyers
On our TCO model (1x H100 SXM at $2.40/GPU-hour on demand, 24/7, plus 0.25 FTE of MLOps at $200k/yr fully loaded = $5,919/month), a self-hosted distilled model only undercuts Claude Opus 5 above ~132M output tokens a month, Claude Sonnet 5 above ~329M, and Claude Haiku 4.5 above ~658M. It never undercuts gpt-5.6-luna or DeepSeek V4-Flash within the throughput capacity of a single GPU. Serverless open-model endpoints have eaten the economic case for owning inference except at genuine scale or where data residency forces it.
The January 2025 shock repriced compute, not software
DeepSeek-R1 shipped on 20 January 2025 with six open-weight distilled students under MIT licence. Seven days later Nvidia fell 17% and lost $589B of market capitalisation, the Nasdaq 100 fell 3%, and the semiconductor index had its worst day since March 2020. The market read distillation as a claim that frontier capability could be reproduced without frontier capital expenditure. It was wrong on the timescale: Nvidia crossed $5T in October 2025 - the first company ever to do so - and stood at roughly $5.43T on 2 September 2026.
DeepSeek's $5.6M figure was a marginal-cost number, and the dispute is about accounting
The DeepSeek-V3 technical report states 2.788M H800 GPU-hours for full training; the widely circulated $5.576M is that figure multiplied by an assumed $2/GPU-hour rental rate. SemiAnalysis countered on 31 January 2025 that DeepSeek's server capex is around $1.6B across roughly 50,000 Hopper GPUs with about $944M of operating cost, and that the published number excludes R&D, data, failed runs and hardware total cost of ownership. The Nature paper on R1 (Sept 2025) is narrower still: $294,000 covers only the reinforcement-learning stage on top of an already-trained base.
The race to zero reversed in August 2026
DeepSeek warned on 6 August 2026 of a significant price increase and implemented it on 16 August: V4-Flash output went from $0.28/MTok to $0.66 off-peak and $1.32 at peak; V4-Pro output from $0.87 to $1.98/$3.96. Cache-hit input tokens rose by as much as 1,100%. Seventeen of twenty-four hours remain at the half-price off-peak rate, and peak hours are set on Beijing business time, so the increase falls hardest on domestic users and lightest on Western buyers. The signal is that ultra-cheap distilled inference was capacity-constrained, not structurally free.
Cloud vendors monetise distillation through fine-tuning and provisioned throughput, not through the distillation itself
OpenAI's Model Distillation ships Stored Completions free and charges standard fine-tuning rates ($25/MTok training for gpt-4.1, $1.50/MTok for gpt-4.1-nano). Amazon Bedrock Model Distillation (GA 1 May 2025) charges for the teacher inference calls used to synthesise data, then bills the resulting custom model at $1.95/month storage plus Provisioned Throughput - there is no on-demand tier for a distilled model on Bedrock at any volume. AWS markets distilled models as up to 500% faster and up to 75% cheaper to run with under 2% accuracy loss on RAG. The workflow is free; the lock-in is in where the student runs.
Distillation allegations moved from a commercial dispute to a congressional one
Microsoft security researchers observed large-scale data exfiltration through OpenAI developer accounts they linked to DeepSeek in late 2024; the probe became public on 29 January 2025. The House Select Committee on the CCP concluded in its April 2025 report that it is highly likely DeepSeek used unlawful model distillation techniques. In February 2026 OpenAI submitted a memo to the same committee describing sophisticated, multi-stage distillation pipelines using obfuscated third-party routers to conceal origin. The financial stake is that a distillation attack converts a multi-hundred-million-dollar capital asset into a commodity a competitor can rent.
Model extraction is cheap enough to be an operating expense, not a capital project
Carlini et al. recovered the exact hidden dimension of gpt-3.5-turbo and estimated the full embedding-projection matrix could be extracted for under $2,000 in API queries; a limited version of the attack cost under $200, and ada and babbage were fully extracted for under $20. Combined with the corpus-generation figures above, the total cash cost of a serious behavioural-cloning effort against a frontier model sits in the $10^3-$10^5 range against a $10^8 asset. No defensive spend scales down to that.
Capital followed the small-model thesis, but the exits were modest
Arcee AI raised a $24M Series A led by Emergence Capital for domain-specific small language models. Predibase, which sold fine-tuning tooling for small open models, raised over $28M and was acquired by Rubrik in June 2025 for a reported $100M-$500M. Together AI, the largest pure-play open-model inference platform, raised $305M at $3.3B in February 2025 and $800M at $8.3B in July 2026. Mistral raised a EUR 1.7B Series C at EUR 11.7B in September 2025 with ASML taking 11%. The value accrued to inference capacity and to European sovereignty plays, not to distillation tooling as a standalone category.
Charts · 7 charts
Output price per million tokens: frontier tier vs small tier
USD/MTok| Vendor | Frontier tier USD/MTok | Small / distilled tier USD/MTok |
|---|---|---|
| OpenAI (gpt-6-astra) | 50 | — |
| Anthropic (Fable 5.1) | 50 | — |
| Anthropic (Opus 5) | 25 | — |
| Google (Gemini 3.1 Pro) | 12 | — |
| Mistral (Medium 3.5) | 7.5 | — |
| xAI (grok-4.6) | 6 | — |
| Alibaba (qwen3.8-max) | 6 | — |
| DeepSeek (V4-Pro peak) | 3.96 | — |
| Meta (Llama 3.3 70B) | 1.04 | — |
| OpenAI (gpt-5.6-luna) | — | 1.2 |
| Anthropic (Haiku 4.5) | — | 5 |
| Google (Gemini 2.5 Flash-Lite) | — | 0.4 |
| Mistral (Ministral 3 3B) | — | 0.1 |
| xAI (grok-build-0.1) | — | 2 |
| Alibaba (qwen-turbo) | — | 0.2 |
| DeepSeek (V4-Flash off-peak) | — | 0.66 |
| Meta (Llama 3 8B Lite) | — | 0.14 |
Standard list prices as of 3 September 2026. Anthropic appears twice because Opus 5 and Fable 5.1 sit at different points on the same line. Meta prices are Together AI's serverless rates, since Meta does not sell a first-party API for these models.
Sources: developers.openai.com · platform.claude.com · ai.google.dev · mistral.ai · docs.x.ai · alibabacloud.com · api-docs.deepseek.com · together.ai
Price of a fixed capability level, 2021-2026
USD/MTok| Date | GPT-3 level (MMLU 42) - a16z USD/MTok | GPT-3.5 level (MMLU >= 64.8) - Epoch USD/MTok | GPT-4 level (MMLU >= 86) - Epoch USD/MTok | GPQA Diamond >= 50 - Epoch USD/MTok |
|---|---|---|---|---|
| 2021-11 | 60 | — | — | — |
| 2024-11 | 0.06 | — | — | — |
| 2022-11 | — | 20 | — | — |
| 2024-10 | — | 0.07 | — | — |
| 2023-03 | — | — | 37.5 | — |
| 2025-02 | — | — | 0.18 | — |
| 2023-11 | — | — | — | 15 |
| 2024-12 | — | — | — | 0.12 |
Plot on a log y-axis. Each series has only the two endpoints published by the source; the intermediate path was not disclosed as a series. Epoch's aggregate finding across six benchmarks is a 9x-900x annual decline with a median near 50x.
Monthly bill vs monthly volume: when does self-hosting a distilled model win?
USD| Output tokens per month (millions) | Claude Opus 5 (teacher) USD | Claude Sonnet 5 USD | Claude Haiku 4.5 (vendor small tier) USD | gpt-5.6-luna (vendor nano tier) USD | DeepSeek V4-Flash off-peak USD | Self-hosted distilled 8B, 1x H100 + 0.25 FTE USD |
|---|---|---|---|---|---|---|
| 1 | 45 | 18 | 9 | 2 | 2 | 5,919 |
| 10 | 450 | 180 | 90 | 20 | 15 | 5,919 |
| 50 | 2,250 | 900 | 450 | 100 | 77 | 5,919 |
| 100 | 4,500 | 1,800 | 900 | 200 | 154 | 5,919 |
| 200 | 9,000 | 3,600 | 1,800 | 400 | 308 | 5,919 |
| 500 | 22,500 | 9,000 | 4,500 | 1,000 | 770 | 5,919 |
| 1,000 | 45,000 | 18,000 | 9,000 | 2,000 | 1,540 | 5,919 |
Author-computed. Assumes 4 input tokens per output token. Self-hosted line is flat at $5,919/month ($1,752 GPU + $4,167 staffing) up to ~1,051M output tokens/month capacity. Crossings: 132M vs Opus 5, 329M vs Sonnet 5, 658M vs Haiku 4.5. The gpt-5.6-luna and DeepSeek lines never cross within capacity.
Sources: platform.claude.com · developers.openai.com · api-docs.deepseek.com · gmicloud.ai
Training cost: frontier runs vs distilled students
USD| Model / run | USD (log scale) USD |
|---|---|
| DeepSeek infrastructure (SemiAnalysis est.) | 1,600,000,000 |
| Gemini Ultra 1.0 (AI Index) | 191,000,000 |
| GPT-4 (AI Index) | 78,000,000 |
| GPT-4 (Epoch amortised) | 40,000,000 |
| DeepSeek-V3 final run (claimed) | 5,576,000 |
| DeepSeek-R1 RL stage (Nature) | 294,000 |
| Sky-T1-32B | 450 |
| s1-32B | 50 |
| TinyZero | 30 |
Log scale spans eight orders of magnitude. The bars are not accounted on a common basis - see the training-cost table notes. The point of the chart is the shape of the ladder, not a like-for-like comparison.
Sources: epoch.ai · semianalysis.com · arxiv.org · novasky-ai.github.io · cnn.com
Teacher-query bill to assemble a distillation corpus, 2026 list prices
USD| Teacher model | 17k traces (Sky-T1 / Bespoke-Stratos scale) USD | 800k traces (DeepSeek R1-Distill scale) USD |
|---|---|---|
| Claude Opus 5 | 850 | 40,000 |
| gpt-5.6-sol | 680 | 32,000 |
| Gemini 3.5 Flash | 306 | 14,400 |
| grok-4.6 | 204 | 9,600 |
| deepseek-v4-pro (off-peak) | 67 | 3,168 |
| gpt-5.6-luna | 41 | 1,920 |
| gpt-oss-120b (Groq) | 20 | 960 |
Author calculation: 2,000 output tokens per trace, input cost ignored. Halve every bar again if the Batch API 50% discount applies.
Sources: platform.claude.com · developers.openai.com · ai.google.dev · api-docs.deepseek.com · console.groq.com
Nvidia market capitalisation around the distillation shocks
USD billions| Event | Market-cap change | Market-cap level |
|---|---|---|
| 27 Jan 2025: DeepSeek-R1 shock (one session) | -589 | — |
| 14 May - 8 Jul 2026: drawdown from peak | -1,000 | — |
| Oct 2025: first $5T company | — | 5,060 |
| 2 Sep 2026 | — | 5,430 |
The 27 January 2025 move was a 17% single-session fall and the largest one-day market-cap loss in US stock-market history; the Nasdaq 100 fell 3% and the S&P 500 1.5% the same day. The 2026 drawdown was attributed to rotation into memory and storage semiconductors rather than to distillation news.
Sources: cnbc.com · finance.yahoo.com · stockanalysis.com
DeepSeek V4 pricing before and after 16 August 2026
USD/MTok| Model and token type | Before 16 Aug 2026 USD/MTok | After, off-peak USD/MTok | After, peak USD/MTok |
|---|---|---|---|
| V4-Flash input (cache miss) | 0.14 | 0.22 | 0.44 |
| V4-Flash output | 0.28 | 0.66 | 1.32 |
| V4-Pro input (cache miss) | 0.435 | 0.66 | 1.32 |
| V4-Pro output | 0.87 | 1.98 | 3.96 |
Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday - Beijing business hours - so 17 of 24 hours stay at the off-peak rate and Western buyers are largely insulated. Cache-hit input tokens, not shown here, rose by as much as 1,100%.
Sources: api-docs.deepseek.com · infoworld.com
Tables · 8 tables
Frontier vs small-tier list prices, September 2026
11 rows| Vendor | Frontier model | Frontier in USD/MTok | Frontier out USD/MTok | Small / distilled tier | Small in USD/MTok | Small out USD/MTok | Output price ratio x |
|---|---|---|---|---|---|---|---|
| OpenAI | gpt-6-astra | 10 | 50 | gpt-5.6-luna | 0.2 | 1.2 | 41.7 developers.openai.com |
| OpenAI | gpt-5.6-sol | 4 | 20 | gpt-5.4-nano | 0.2 | 1.25 | 16 developers.openai.com |
| Anthropic | Claude Fable 5.1 | 10 | 50 | Claude Haiku 4.5 | 1 | 5 | 10 platform.claude.com |
| Anthropic | Claude Opus 5 | 5 | 25 | Claude Haiku 4.5 | 1 | 5 | 5 platform.claude.com |
| Gemini 3.1 Pro Preview | 2 | 12 | Gemini 2.5 Flash-Lite | 0.1 | 0.4 | 30 ai.google.dev | |
| xAI | grok-4.6 (<200k) | 2 | 6 | grok-build-0.1 | 1 | 2 | 3 docs.x.ai |
| DeepSeek | deepseek-v4-pro (peak) | 1.32 | 3.96 | deepseek-v4-flash (off-peak) | 0.22 | 0.66 | 6 api-docs.deepseek.com |
| Mistral | Mistral Medium 3.5 | 1.5 | 7.5 | Ministral 3 (3B) | 0.1 | 0.1 | 75 mistral.ai |
| Alibaba (Qwen) | qwen3.8-max | 2 | 6 | qwen-turbo | 0.05 | 0.2 | 30 alibabacloud.com |
| Meta (via Together AI) | Llama 3.3 70B | 1.04 | 1.04 | Llama 3 8B Instruct Lite | 0.14 | 0.14 | 7.4 together.ai |
| OpenAI open weights (via Groq) | gpt-oss-120b | 0.15 | 0.6 | gpt-oss-20b | 0.075 | 0.3 | 2 console.groq.com |
Ratios are output-token price ratios, computed from the listed figures. DeepSeek peak hours are 01:00-04:00 and 06:00-10:00 UTC Monday-Friday; the other 17 hours are half price. xAI and Anthropic rates shown are for prompts below the long-context threshold.
Sources: developers.openai.com · platform.claude.com · ai.google.dev · docs.x.ai · api-docs.deepseek.com · mistral.ai · alibabacloud.com · together.ai · console.groq.com
What it costs to build a model vs what it costs to copy one
11 rows| Model / run | Organisation | Reported cost USD | Compute | Date | What the number covers |
|---|---|---|---|---|---|
| Gemini Ultra 1.0 | Google DeepMind | 191,000,000 | TPU v4, ~35 MW | 2023-12 | Stanford AI Index cloud-rental accounting; Epoch's amortised-hardware method gives ~$30M epoch.ai |
| GPT-4 | OpenAI | 78,000,000 | A100 cluster | 2023-03 | Stanford AI Index cloud-rental accounting; Epoch amortised-hardware estimate ~$40M epoch.ai |
| GPT-4 (amortised) | OpenAI | 40,000,000 | A100 cluster | 2023-03 | Epoch AI: depreciated hardware + energy for the final run only epoch.ai |
| DeepSeek infrastructure (all-in estimate) | DeepSeek / High-Flyer | 1,600,000,000 | ~50,000 Hopper GPUs | 2025-01 | SemiAnalysis estimate of total server capex; ~$944M of operating cost on top semianalysis.com |
| DeepSeek-V3 (final run, claimed) | DeepSeek | 5,576,000 | 2.788M H800 GPU-hours | 2024-12 | GPU-hours are in the technical report; the dollar figure applies an assumed $2/GPU-hour rental rate arxiv.org |
| DeepSeek-R1 (RL stage) | DeepSeek | 294,000 | 512 H800s x 80 hours (~41k GPU-hours) | 2025-09 | Disclosed in the Nature paper; excludes the V3 base model, data, energy, infrastructure and staff cnn.com |
| DeepSeek-R1-Distill-Qwen-32B | DeepSeek | — | SFT on 800k R1-generated samples | 2025-01 | Cost undisclosed; the 800k-sample corpus is the expensive input, not the fine-tune arxiv.org |
| Bespoke-Stratos-32B | Bespoke Labs | — | 8x H100 for 27 hours (216 GPU-hours) | 2025-01 | Cost undisclosed. 17k traces distilled from DeepSeek-R1 in 1.5 hours of teacher inference; 47x fewer examples than R1-Distill-Qwen-32B bespokelabs.ai |
| Sky-T1-32B-Preview | NovaSky, UC Berkeley | 450 | 8x H100 for 19 hours (152 GPU-hours) | 2025-01 | Full disclosed cost of the fine-tune; 17k traces from QwQ-32B-Preview reformatted with GPT-4o-mini novasky-ai.github.io |
| s1-32B | Stanford / University of Washington | 50 | 16x H100 for 26 minutes (~6.9 GPU-hours) | 2025-02 | GPU time is stated in the paper; the ~$50 figure comes from press coverage, not the paper itself. 1,000 traces distilled from Gemini Thinking Experimental arxiv.org |
| TinyZero | UC Berkeley (Jiayi Pan et al.) | 30 | 3B Qwen base, RL on Countdown task | 2025-01 | Server cost for the experiments only; reproduces R1-Zero-style self-verification on a narrow task dailycal.org |
The costs in this table are NOT comparable on a like-for-like basis. Frontier rows are whole pre-training runs; distillation rows are fine-tunes that assume a free base model and a paid-for teacher. The honest comparison is the ratio between the top of the table and the bottom, which is roughly 10^5 to 10^6.
Sources: epoch.ai · semianalysis.com · arxiv.org · novasky-ai.github.io · arxiv.org
What it costs to buy a teacher's reasoning traces at 2026 list prices
7 rows| Teacher model | Output price USD/MTok | 17k traces (34M out-tok) USD | 800k traces (1.6B out-tok) USD |
|---|---|---|---|
| Claude Opus 5 | 25 | 850 | 40,000 platform.claude.com |
| gpt-5.6-sol | 20 | 680 | 32,000 developers.openai.com |
| Gemini 3.5 Flash | 9 | 306 | 14,400 ai.google.dev |
| grok-4.6 | 6 | 204 | 9,600 docs.x.ai |
| deepseek-v4-pro (off-peak) | 1.98 | 67 | 3,168 api-docs.deepseek.com |
| gpt-5.6-luna | 1.2 | 41 | 1,920 developers.openai.com |
| gpt-oss-120b (via Groq) | 0.6 | 20 | 960 console.groq.com |
Applying the Batch API discount (50% at Anthropic and OpenAI) halves every figure again. These are list prices for legitimate API use; a distillation programme that violates terms of service would face the same compute bill plus the cost of evading detection, which OpenAI's February 2026 memo to Congress describes as obfuscated third-party routers.
Sources: platform.claude.com · developers.openai.com · api-docs.deepseek.com
TCO: teacher API vs distilled API vs self-hosted, by monthly volume
7 rows| Output tokens/month millions | Claude Opus 5 ($5/$25) USD/mo | Claude Sonnet 5 ($2/$10) USD/mo | Claude Haiku 4.5 ($1/$5) USD/mo | gpt-5.6-luna ($0.20/$1.20) USD/mo | DeepSeek V4-Flash off-peak ($0.22/$0.66) USD/mo | Self-hosted distilled 8B, 1x H100 USD/mo |
|---|---|---|---|---|---|---|
| 1 | 45 | 18 | 9 | 2 | 2 | 5,919 platform.claude.com |
| 10 | 450 | 180 | 90 | 20 | 15 | 5,919 platform.claude.com |
| 50 | 2,250 | 900 | 450 | 100 | 77 | 5,919 platform.claude.com |
| 100 | 4,500 | 1,800 | 900 | 200 | 154 | 5,919 platform.claude.com |
| 200 | 9,000 | 3,600 | 1,800 | 400 | 308 | 5,919 platform.claude.com |
| 500 | 22,500 | 9,000 | 4,500 | 1,000 | 770 | 5,919 platform.claude.com |
| 1,000 | 45,000 | 18,000 | 9,000 | 2,000 | 1,540 | 5,919 platform.claude.com |
Breakevens for the self-hosted option: 132M output tokens/month vs Opus 5, 329M vs Sonnet 5, 658M vs Haiku 4.5, 2,959M vs gpt-5.6-luna and 3,843M vs DeepSeek V4-Flash off-peak. The last two exceed single-GPU capacity, so at these prices a one-GPU deployment never pays back against the cheapest hosted endpoints. Excludes the one-off distillation cost ($450-$40,000, see other tables), redundancy, and the fact that measured self-hosted cost per output MTok on identical H100 hardware spans $0.21 to $15.25 depending purely on request concurrency.
Sources: platform.claude.com · developers.openai.com · api-docs.deepseek.com · gmicloud.ai · arxiv.org
Cloud vendor distillation products and how they charge
7 rows| Vendor | Product | Launched | How it charges | Stated benefit |
|---|---|---|---|---|
| OpenAI | Model Distillation (Stored Completions + Evals) | 2024-10 | Stored Completions free; standard fine-tuning rates apply ($25/MTok training for gpt-4.1, $5 for gpt-4.1-mini, $1.50 for gpt-4.1-nano, $100/hr for o4-mini) | Train smaller cost-efficient models on frontier outputs for a specific task developers.openai.com |
| Amazon Web Services | Amazon Bedrock Model Distillation | 2024-12 preview, GA 2025-05-01 | Teacher inference charged at on-demand rates when Bedrock synthesises data; custom model storage $1.95/model/month; inference only via Provisioned Throughput | Up to 500% faster and up to 75% less expensive to run, with under 2% accuracy loss on RAG docs.aws.amazon.com |
| Google Cloud | Vertex AI supervised tuning / distillation | 2024 | Per training token; tuned model endpoints billed at 1.5x the base model rate | Task-specific tuned Gemini Flash and Flash-Lite students cloud.google.com |
| Microsoft | Azure OpenAI / Microsoft Foundry stored completions and distillation | 2024-10 | Standard Azure fine-tuning and inference rates; distillation requires a minimum of 10 stored completions | Turn production traffic against a large model into a fine-tuning set for a small one learn.microsoft.com |
| Anthropic | No first-party distillation product | n/a | Tiered model line (Fable/Opus/Sonnet/Haiku) plus Batch API 50% discount and prompt caching down to 0.025x input price | Vendor captures the cost-reduction margin internally rather than selling a distillation pipeline platform.claude.com |
| Together AI | Serverless open-model inference and fine-tuning | 2023 | Per token for serverless; dedicated H100 at $3.99/hr promotional ($5.49 regular), B200 at $8.99/hr | Host the distilled student without owning hardware together.ai |
| Fireworks AI | Serverless and on-demand GPU deployments | 2023 | H100 80GB and H200 141GB at $8.00/hr from 1 Sep 2026 (previously $7.00); B200 $13.00/hr; GB300 $20.00/hr; 1.5x premium for region-restricted deployments | Dedicated capacity for custom and distilled models fireworks.ai |
Bedrock's constraint is the commercially interesting one: a distilled model cannot be served on demand, so the customer trades a per-token bill for an hourly Provisioned Throughput commitment, which reverses the economics for low-volume users.
Sources: developers.openai.com · docs.aws.amazon.com · aws.amazon.com · together.ai · fireworks.ai
Price of a fixed capability level over time
5 rows| Capability level | First available | Price then USD/MTok | Cheapest as of | Price then USD/MTok | Decline x |
|---|---|---|---|---|---|
| GPT-3 level (MMLU 42) | 2021-11 | 60 | 2024-11 | 0.06 | 1,000 a16z.com |
| GPT-3.5 level (MMLU >= 64.8) | 2022-11 | 20 | 2024-10 | 0.07 | 286 epoch.ai |
| GPT-4 level (MMLU >= 86.0) | 2023-03 | 37.5 | 2025-02 | 0.18 | 208 epoch.ai |
| PhD-level science (GPQA Diamond >= 50) | 2023-11 | 15 | 2024-12 | 0.12 | 125 epoch.ai |
| MMLU 83 (GPT-4 launch level) | 2023-03 | 30 | 2024-11 | 0.48 | 62 a16z.com |
Epoch's headline range across six benchmarks is a 9x-900x annual decline, with a median around 50x and roughly 200x for models released since 2024. a16z's 'LLMflation' framing is a 10x decline per year and 1,000x over three years. The MMLU 83 start price is a16z's stated ~62x reduction applied to the GPT-4 launch price; treat it as derived rather than directly quoted.
Capital raised against the small-model / distillation thesis
8 rows| Company | Event | Amount USD millions | Valuation | Date | Thesis |
|---|---|---|---|---|---|
| Arcee AI | Seed | 5.5 | undisclosed | 2023-12 | Domain-specific small language models arcee.ai |
| Arcee AI | Series A (Emergence Capital) | 24 | undisclosed | 2024-07 | Model merging and Spectrum training to cut SLM training cost venturebeat.com |
| Predibase | Total VC raised before exit | 28 | undisclosed | 2024 | Fine-tuning tooling for small open models (Llama, Mistral) techtarget.com |
| Predibase | Acquired by Rubrik | 300 | reported $100M-$500M range | 2025-06-25 | Agentic AI needs cheap task-specific models techcrunch.com |
| Together AI | Series B (General Catalyst, Prosperity7) | 305 | $3.3B | 2025-02-20 | End-to-end platform for building with 200+ open-source models siliconangle.com |
| Together AI | Series C | 800 | $8.3B | 2026-07-01 | Neocloud inference capacity for open and distilled models techcrunch.com |
| Mistral AI | Series C (ASML lead, EUR 1.3B of EUR 1.7B) | 1,900 | EUR 11.7B post-money (~$14B) | 2025-09-09 | European open-weight model family from 3B Ministral up to Large mistral.ai |
| Mistral AI | Reported raise in progress | 3,200 | reported EUR 20B target | 2026-06 | Rumoured EUR 3B round; unconfirmed techcrunch.com |
Amounts converted to USD millions where the original is in euros, at approximately 1.13 USD/EUR for the September 2025 round (CNBC reported the valuation as ~$14B). The Predibase acquisition amount is the midpoint of a reported $100M-$500M range and should be treated as an estimate, not a disclosed figure. The Mistral 2026 round is rumoured and unconfirmed.
Sources: arcee.ai · techcrunch.com · techcrunch.com · cnbc.com
The cost-of-theft asymmetry, line by line
14 rows| Line item | Side | Cost USD | Basis |
|---|---|---|---|
| Gemini Ultra 1.0 final training run | Build | 191,000,000 | Stanford AI Index cloud-rental accounting epoch.ai |
| GPT-4 final training run | Build | 78,000,000 | Stanford AI Index cloud-rental accounting epoch.ai |
| GPT-4 final training run (amortised) | Build | 40,000,000 | Epoch AI hardware depreciation + energy epoch.ai |
| DeepSeek-V3 pre-training (claimed) | Build | 5,576,000 | 2.788M H800 GPU-hours at an assumed $2/hr arxiv.org |
| DeepSeek-R1 RL stage | Build | 294,000 | Nature paper: 512 H800s x 80 hours cnn.com |
| 800k-trace reasoning corpus from Claude Opus 5 | Copy | 40,000 | Author calculation at $25/MTok output, 2,000 tokens/trace platform.claude.com |
| 800k-trace reasoning corpus from gpt-5.6-luna | Copy | 1,920 | Author calculation at $1.20/MTok output developers.openai.com |
| Full projection-matrix extraction of gpt-3.5-turbo (estimated) | Copy | 2,000 | Carlini et al., estimated query cost arxiv.org |
| Partial model-stealing attack on gpt-3.5 | Copy | 200 | Carlini et al., executed attack arxiv.org |
| Bespoke-Stratos 17k-trace corpus generation | Copy | — | 1.5 hours of DeepSeek-R1 inference; dollar cost undisclosed bespokelabs.ai |
| Sky-T1-32B student fine-tune | Copy | 450 | 8x H100 for 19 hours novasky-ai.github.io |
| s1-32B student fine-tune | Copy | 50 | 16x H100 for 26 minutes; dollar figure from press coverage arxiv.org |
| Full projection-matrix extraction of ada and babbage | Copy | 20 | Carlini et al., executed attack arxiv.org |
| TinyZero R1-Zero-style reproduction | Copy | 30 | Server cost for the experiments, narrow task only dailycal.org |
The build side and copy side are not substitutes: a distilled student inherits behaviour on the distribution it was distilled over, not the teacher's full capability surface. But for the specific task a buyer cares about, the copy side is 3-6 orders of magnitude cheaper, and no legal or technical defence currently scales down to that price point.
Sources: epoch.ai · arxiv.org · novasky-ai.github.io
Timeline · 30 events
OpenAI ships Model Distillation in the API
Stored Completions (free) plus Evals give developers a managed pipeline to capture frontier outputs and fine-tune a smaller student. Distillation becomes a supported product rather than a research technique.
Source: openai.coma16z publishes 'LLMflation'
Argues inference cost for a fixed capability falls ~10x per year and has fallen 1,000x in three years, from $60/MTok for GPT-3 in Nov 2021 to $0.06/MTok for Llama 3.2 3B on Together.ai.
Source: a16z.comAWS announces Amazon Bedrock Model Distillation at re:Invent
Claims distilled models can be up to 500% faster and up to 75% less expensive to run, with under 2% accuracy loss for RAG use cases.
Source: press.aboutamazon.comDeepSeek-V3 technical report discloses 2.788M H800 GPU-hours
The report gives GPU-hours, not dollars; the widely quoted $5.576M figure is that number multiplied by an assumed $2/GPU-hour rental rate.
Source: arxiv.orgSky-T1-32B-Preview trained for under $450
NovaSky at UC Berkeley distils QwQ-32B-Preview traces into Qwen2.5-32B-Instruct using 17k examples and 19 hours on 8 H100s, matching o1-preview on Math500 (82.4 vs 81.4) and AIME24 (43.3 vs 40.0).
Source: novasky-ai.github.ioDeepSeek releases R1 plus six open-weight distilled students under MIT licence
R1-Distill-Qwen-32B scores 72.6 on AIME 2024 and 94.3 on MATH-500; R1-Distill-Llama-70B scores 86.7 on AIME 2024. Frontier-class reasoning becomes free to download.
Source: huggingface.coBespoke-Stratos-32B distilled from DeepSeek-R1 on 17k traces
Trained on 8xH100 for 27 hours using 47x fewer examples than R1-Distill-Qwen-32B; the R1 traces took 1.5 hours to generate. GPT-4o-mini filtering raised retained-correct-solution rate from 25% to 73%.
Source: bespokelabs.aiTinyZero reproduces R1-Zero-style behaviour for under $30
UC Berkeley graduate researchers use RL on a 3B Qwen base for Countdown and multiplication tasks; the $30 is server cost for the experiments.
Source: dailycal.orgNvidia loses $589B of market cap in one session
Shares fall ~17% from an open of $142.02 to close at $118.50. Largest single-day market-cap loss in US stock-market history. Nasdaq 100 -3%, S&P 500 -1.5%, semiconductor index worst day since March 2020.
Source: cnbc.comMicrosoft and OpenAI investigate DeepSeek-linked accounts
Microsoft security researchers had observed large-scale data exfiltration through OpenAI developer accounts in late 2024. White House AI adviser David Sacks says there is substantial evidence DeepSeek used OpenAI model outputs.
Source: bloomberg.comSemiAnalysis disputes the $5.6M figure
Estimates DeepSeek's server capex at ~$1.6B across roughly 50,000 Hopper GPUs with ~$944M of operating cost, arguing the published figure covers only GPU time for the pre-training run.
Source: semianalysis.coms1-32B: 1,000 traces, 26 minutes, reported ~$50
Distilled from Gemini Thinking Experimental into Qwen2.5-32B-Instruct on 16 H100s. Exceeds o1-preview on competition maths by up to 27% with budget forcing at inference time.
Source: arxiv.orgTogether AI raises $305M Series B at $3.3B
General Catalyst and Prosperity7 lead; Nvidia, Salesforce Ventures, Kleiner Perkins participate. The capital funds Blackwell capacity for serving open and distilled models.
Source: siliconangle.comHouse Select Committee publishes 'DeepSeek Unmasked'
Concludes it is highly likely DeepSeek used unlawful model distillation techniques against US models, and that the scale was such that V3 often self-identifies as ChatGPT.
Source: techpolicy.pressAmazon Bedrock Model Distillation reaches general availability
Distilled models can only be served on Provisioned Throughput, not on demand, which changes the economics for low-volume buyers.
Source: docs.aws.amazon.comOpenAI cuts o3 pricing by ~80%
From $10/$40 to $2/$8 per million input/output tokens, attributed to inference-stack optimisation on the same model. o3-pro launches alongside.
Source: venturebeat.comRubrik acquires Predibase
Reported at $100M-$500M. Predibase sold fine-tuning tooling for small open models and had raised over $28M from Felicis, Greylock and Sancus Ventures.
Source: techcrunch.comOpenAI releases gpt-oss-120b and gpt-oss-20b under Apache 2.0
First open-weight OpenAI models since GPT-2. gpt-oss-120b has 116.8B total / 5.1B active parameters and matches or exceeds o4-mini on competition coding. Now served at $0.15/$0.60 and $0.075/$0.30 per MTok on Groq.
Source: openai.comMistral raises EUR 1.7B Series C at EUR 11.7B, ASML leads
ASML puts in EUR 1.3B for ~11% fully diluted. Valuation roughly doubles from EUR 5.8B. Mistral's line spans Ministral 3B at $0.10/$0.10 up to Medium 3.5 at $1.50/$7.50.
Source: mistral.aiNature publishes the DeepSeek-R1 paper with a $294,000 cost figure
512 H800s for 80 hours for the reinforcement-learning stage. The figure excludes the ~$6M base model, data, energy, infrastructure and staff; commentators including The Register noted the number is not a total cost.
Source: cnn.comNvidia becomes the first $5 trillion company
Closes at a $5.06T market capitalisation. DeepSeek launches V4-Pro the same day; Nvidia's valuation is unaffected.
Source: thenationalnews.comOne year on, DeepSeek no longer moves markets
CNBC reports that the companies hit by the January 2025 selloff have not just recovered but grown, and that subsequent DeepSeek releases produced no comparable investor reaction.
Source: cnbc.comOpenAI memo to the House Select Committee on China
Accuses DeepSeek of sophisticated multi-stage distillation pipelines using obfuscated third-party routers and unauthorised resellers to conceal origin and evade access restrictions.
Source: cdn.openai.comFrontier Model Forum publishes an issue brief on adversarial distillation
Frames the risk in terms of how many teacher outputs an attacker can obtain and how much compute they have, but publishes no cost figures.
Source: frontiermodelforum.orgTogether AI raises $800M at $8.3B
The largest pure-play open-model inference platform roughly 2.5x its valuation in 17 months, on demand for serving open-weight and distilled models.
Source: techcrunch.comNvidia sheds roughly $1T from its 14 May 2026 peak
Down ~16% from the peak on rotation into memory and storage semiconductors, not on distillation news. Nvidia still holds ~97% of the server GPU market and trades at 18x forward earnings.
Source: finance.yahoo.comDeepSeek V4-Flash launches at $0.14/$0.28 per MTok
Described in the press as accelerating the AI industry's race to zero. The price lasts sixteen days.
Source: axios.comDeepSeek warns of a significant API price increase
Bloomberg and SCMP report the warning ahead of the change, citing surging demand for low-cost models and strained capacity.
Source: scmp.comDeepSeek raises V4 prices by up to 1,100% and introduces peak/off-peak rates
V4-Flash output goes from $0.28 to $0.66 off-peak / $1.32 peak; V4-Pro output from $0.87 to $1.98/$3.96. Cache-hit input tokens rise most. Peak hours are 01:00-04:00 and 06:00-10:00 UTC.
Source: infoworld.comNvidia market capitalisation stands at ~$5.43T
Up ~28% year on year, and about 9x the value erased on 27 January 2025. The distillation shock repriced sentiment, not the compute build-out.
Source: stockanalysis.com
Glossary · 16 terms
- MTok
- One million tokens. The standard unit for LLM API pricing. Input and output tokens are billed at different rates, with output typically 4-5x more expensive.
- Blended cost per MTok
- A single price figure for a workload with a known input:output ratio. For a 4:1 ratio, blended cost per million output tokens = 4 x input price + output price. Used throughout the TCO tables here.
- Teacher API cost
- The money paid to a frontier vendor for the inference calls that generate a distillation corpus. Usually the dominant cash cost of a distillation project, larger than the student fine-tune itself.
- Amortised training cost
- Epoch AI's accounting method: hardware depreciation plus energy over the final training run. Produces figures roughly half those of cloud-rental accounting, which charges the full market rate for every chip-hour.
- Cloud-rental accounting
- Costing a training run at the price of renting equivalent GPUs on the open market. Used by the Stanford AI Index. Gives $78M for GPT-4 where Epoch's amortised method gives ~$40M.
- Marginal training cost
- The GPU-hour bill for a single successful run, excluding R&D, failed runs, data acquisition, staff and infrastructure. DeepSeek's $5.6M and $294,000 figures are both marginal costs.
- Provisioned Throughput
- AWS Bedrock's committed-capacity billing mode, charged hourly per Model Unit. Custom and distilled models on Bedrock can only be served this way, which imposes a fixed monthly floor regardless of volume.
- Prompt caching
- Charging a reduced rate for repeated prompt prefixes. Anthropic charges 1.25x base input to write a 5-minute cache and 0.1x to read it (0.025x on Fable 5.1 and Mythos 5.1), which can cut effective input cost by up to 90%.
- Batch API discount
- A 50% reduction on both input and output tokens for asynchronous processing, offered by both OpenAI and Anthropic. Applies to distillation-corpus generation, halving the teacher-query bill.
- Peak/off-peak pricing
- Time-of-day API pricing, introduced by DeepSeek on 16 August 2026. Peak hours (01:00-04:00 and 06:00-10:00 UTC) cost double the off-peak rate; 17 of 24 hours remain off-peak.
- LLMflation
- a16z's term for the ~10x-per-year fall in inference cost at a fixed capability level. Epoch AI measures the same phenomenon at 9x-900x per year across six benchmarks, median near 50x.
- Breakeven volume
- The monthly token volume at which a fixed-cost self-hosted deployment becomes cheaper than a per-token API. Rises as API prices fall, which is why the case for owning inference has weakened even as distillation has got easier.
- Cost-of-theft asymmetry
- The gap between what it costs to create a capability and what it costs to copy it. For frontier LLMs the ratio is roughly 10^3 to 10^5, because behavioural cloning requires only API access and modest fine-tuning compute.
- Model extraction attack
- Recovering structural parameters of a black-box model through API queries alone. Carlini et al. recovered the full projection matrix of ada and babbage for under $20 and estimated under $2,000 for gpt-3.5-turbo.
- Neocloud
- A GPU-specialist cloud provider (Together AI, Fireworks, CoreWeave, RunPod, Vast.ai) that undercuts hyperscalers on GPU-hour pricing. H100 on-demand rates run ~$1.73-$4.00/hr at neoclouds vs $4.00-$8.00 at AWS, GCP and Azure.
- Student / teacher
- In distillation, the small model being trained (student) and the large model whose outputs it learns from (teacher). Financially: the student is the asset you own, the teacher is the line item on your API bill.
Sources · 63 sources
Every figure on this page comes from one of these primary sources. Compiled 3 September 2026.
- OpenAI API Pricing
- Anthropic Claude Pricing
- Gemini API Pricing
- DeepSeek API Pricing
- Mistral API Pricing
- xAI Models and Pricing
- Alibaba Cloud Model Studio model pricing
- Together AI Pricing
- Fireworks AI Pricing
- Groq supported models
- Amazon Bedrock Pricing
- Customize a model with distillation in Amazon Bedrock
- AWS Strengthens Amazon Bedrock with Industry-First AI Safeguard, New Agent Capability and Model Customization
- Model Distillation in the API
- How to use stored completions and distillation in Azure OpenAI
- Vertex AI pricing
- Sky-T1: Train your own O1 preview model within $450
- Researchers open source Sky-T1, a reasoning AI model that can be trained for less than $450
- s1: Simple test-time scaling
- Bespoke-Stratos: The unreasonable effectiveness of reasoning distillation
- Campus researchers replicate disruptive Chinese AI for $30
- DeepSeek-V3 Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-R1-Distill-Llama-70B model card
- China's DeepSeek shook the tech world. Its developer just revealed the cost
- DeepSeek didn't really train its flagship model for $294,000
- DeepSeek Debates: Chinese Leadership On Cost, True Training Cost
- Nvidia sheds almost $600 billion in market cap, biggest drop ever
- Nvidia stock plummets, loses record $589 billion as DeepSeek prompts questions over AI spending
- Nvidia market cap falls by $1 trillion as stock swoons
- NVIDIA (NVDA) Market Cap & Net Worth
- Will DeepSeek's new AI model crash Nvidia's $5tn party?
- Why DeepSeek didn't cause an investor frenzy again in 2025
- Microsoft Probing If DeepSeek-Linked Group Improperly Obtained OpenAI Data
- US House Select Committee Report Accuses DeepSeek of Spying and Circumventing Export Controls on Chips
- OpenAI update to the US House Select Committee on the CCP
- Issue Brief: Adversarial Distillation
- LLM inference prices have fallen rapidly but unequally across tasks
- How much does it cost to train frontier AI models?
- Welcome to LLMflation - LLM inference cost is going down fast
- Stealing Part of a Production Language Model
- Stealing part of a production language model (ICML 2024)
- OpenAI announces 80% price drop for o3
- gpt-oss-120b & gpt-oss-20b Model Card
- DeepSeek raises some V4 prices by more than 10x as AI demand strains capacity
- Tech Brief: DeepSeek Launches V4-Pro and Raises API Prices by as Much as 1,100%
- DeepSeek signals significant price hike amid surge in demand for low-cost AI models
- DeepSeek's new bargain model accelerates AI's race to zero
- Arcee AI secures $24M Series A to transform the landscape of small language models
- Small language models rising as Arcee AI lands $24M Series A
- Rubrik acquires Predibase to accelerate adoption of AI agents
- Rubrik pivots to generative AI with Predibase acquisition
- Together AI raises $305M for its AI-optimized public cloud
- Neocloud Together AI raises $800M, leaps to $8.3B valuation
- Mistral AI raises EUR 1.7B to accelerate technological progress with AI
- AI firm Mistral valued at $14 billion as chip giant ASML takes major stake
- Mistral is rumored to be raising EUR 3B at EUR 20B valuation
- NVIDIA H100 GPU Pricing: 2026 Rent vs. Buy Cost Analysis
- H100 Cloud Pricing: Compare 53+ Providers (2026)
- Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Serving Cost
- State of AI: An Empirical 100 Trillion Token Study with OpenRouter
- OpenRouter State of AI
- Small Language Model Market Report 2026