Customer · compiled 4 September 2026 · 78 sources

The buyer’s view of AI distillation: which small model to actually deploy

For an enterprise buyer in September 2026, the distillation question is no longer "is the small model good enough" but "which small model, and what does the licence let me do with it". Within a single generation, distilled and small-sibling tiers retain 88–98% of their larger sibling's benchmark score for 3–25% of the price: GPT-5.6 Luna holds about 94% of Sol's GPQA Diamond at 5% of the input rate, DeepSeek-V4-Flash scores 98% of V4-Pro's at a third of the price, and Gemini 2.5 Flash — which Google explicitly documents as a k-sparse logit distillation of 2.5 Pro — holds 96% at 24%; across a large parameter gap, retention falls to 47–76%, as DeepSeek's own R1 students and Google's Gemma E-series show. The gap that remains is not knowledge but agentic reliability: on SWE-bench Verified, Claude Sonnet 5 retains 89% of Opus 5 and Haiku 4.5 only 76%, so long-horizon coding and tool-use workloads are the last place a frontier tier still pays for itself. Open weights have become the buyer's real leverage — DeepSeek V4 (MIT), Gemma 4 (Apache 2.0), Qwen3.8-27B (Apache 2.0) and Mistral Small 4 (Apache 2.0) all clear 71.2–90.1 GPQA Diamond with no per-token fee and no anti-distillation clause — which is why the Silicon Data enterprise inference index fell from $2.04/MTok on 31 May 2026 to $1.16–1.18 in early August. The three things that should drive the decision, in order, are: whether your traffic needs frontier-grade agentic reliability on more than 15% of calls; whether the vendor's terms let you distil its outputs into your own model (OpenAI and Anthropic say no, DeepSeek's MIT licence says yes); and whether you can absorb the migration cost when a tier is repriced or retired — which, on 2026 evidence, happens roughly every quarter.

Key figures · 8 figures

Enterprise inference price index

1.17 USD / M tokens

-43% since 31 May 2026

Silicon Data index cited by Jefferies; hit a 2026 low of $1.16–$1.18 on 6–8 Aug 2026, down from $2.04 on 31 May and $1.45 in late July.

scmp.com

Cheapest hosted model per 1M tokens (in)

0.035 USD

Amazon Nova Micro

Amazon's cheapest Nova tier. AWS supports it as a distillation student in Bedrock Model Distillation, but does not state that the shipped weights were distilled from Premier or Pro. Text-only, 128K context.

pricepertoken.com

GPQA retained by GPT-5.6 Luna vs Sol

94.2 %

at 5% of the input price (94.2–97.3% across providers)

87.0 vs 92.4 GPQA Diamond; $0.20 vs $4.00 per M input tokens. Both figures are third-party (OpenRouter auto-routing); provider-specific Luna scores on the same page run to 89.9, and DataLearner reports Sol at 93.5 at max thinking, so retention is best read as a 94.2–97.3% range pending the GPT-5.6 system card.

openrouter.ai

GPQA of DeepSeek-V4-Flash vs V4-Pro

97.8 %

at 33% of the price

88.1 vs 90.1 GPQA Diamond, both MIT-licensed open weights. Same-family comparison: DeepSeek documents V4-Flash as a consolidation of V4 domain experts via on-policy distillation, not as a distillation of V4-Pro.

huggingface.co

SWE-bench Verified retained by Claude Haiku 4.5 vs Opus 5

76.4 %

at 20% of the price

73.3 vs 96.0. Agentic coding is where the distilled tiers still lose most.

datanorth.ai

Cost per 1M requests, GPT-5.6 Sol vs Luna

18 x cheaper

$3,600 → $200

At 400 input + 100 output tokens per request, list price, no caching.

developers.openai.com

Best open-weight GPQA Diamond in this dataset

90.1 %

DeepSeek-V4-Pro, MIT licence (open-weight range 71.2–90.1)

1.6T-parameter MoE with 49B active; within 4.2 points of Gemini 3.1 Pro.

huggingface.co

Bedrock Model Distillation claim

75 % cheaper

and up to 500% faster

AWS states distilled models are up to 500% faster and up to 75% less expensive than the original, with under 2% accuracy loss on RAG-style use cases.

aws.amazon.com

Key findings · 10 findings

  1. The cheap tier is now good enough for roughly 85% of enterprise traffic

    Within a single current generation the small sibling retains 88–98% of its teacher's knowledge benchmark. GPT-5.6 Luna scores 87.0 GPQA Diamond against Sol's 92.4; GPT-5.4 mini scores 88.0 against 93.0; Gemini 2.5 Flash scores 82.8 against 2.5 Pro's 86.4; DeepSeek-V4-Flash scores 88.1 against V4-Pro's 90.1. Across a large parameter gap retention falls to 47–76% — DeepSeek's own R1 students range from 91.2% at 70B down to 47.3% at 1.5B, and Gemma 4 31B to E4B retains 69.5%. The practical rule that falls out of the numbers is to route the hardest 5–15% of traffic to a frontier tier and everything else down to a same-generation small tier, not to the smallest model available.

    Sources openrouter.ai · openrouter.ai · arxiv.org · huggingface.co

  2. Agentic coding is the one place the frontier tier still earns its price

    Knowledge benchmarks compress; long-horizon agent benchmarks do not. Claude Opus 5 scores 96.0 on SWE-bench Verified, Sonnet 5 85.2 (89% retention) and Haiku 4.5 73.3 (76% retention). Llama 4 Scout retains 92% of Maverick’s MMLU-Pro but only 76% of its LiveCodeBench. If your workload is a coding agent that must finish a multi-step task unsupervised, the retention curve is much steeper than the GPQA curve suggests and a cheap tier will show up as retries, not as wrong answers.

    Sources datanorth.ai · morphllm.com · anthropic.com · github.com

  3. Open weights are the buyer’s only real negotiating position

    Four open-weight families now clear 71.2–90.1 GPQA Diamond with no per-token fee and no anti-distillation clause: DeepSeek V4-Pro/Flash (MIT, 90.1/88.1), Qwen3.8-27B (Apache 2.0, 89.2), Gemma 4 31B (Apache 2.0, 84.3 GPQA / 85.2 MMLU-Pro) and Mistral Small 4 (Apache 2.0, 71.2 GPQA / 78.0 MMLU-Pro). Having a credible self-host fallback is what makes an API price cut stick — and 2026 has been a year of price cuts.

    Sources huggingface.co · huggingface.co · ai.google.dev · openrouter.ai

  4. You may not distil the models you buy — but you may distil the ones you download

    Anthropic’s Commercial Terms section D.4 bars customers from accessing the Services "to build a competing product or service, including to train competing AI models". OpenAI’s Services Agreement carries an equivalent restriction. Both apply to enterprise accounts. By contrast DeepSeek R1 and V4 ship under MIT, and the R1 release explicitly encourages distillation; Qwen3.8, Gemma 4, Mistral Small 4 and Ministral 3 are Apache 2.0. If your roadmap includes training a small in-house model on model outputs, the licence question decides your teacher before any benchmark does.

    Sources anthropic.com · huggingface.co · ai.google.dev

  5. Distillation is now a purchasable service, not just a research technique

    Amazon Bedrock Model Distillation takes only your prompts: it generates synthetic teacher responses and fine-tunes the student. Nova Premier, Claude 3.5 Sonnet v2 and Llama 3.3 70B are supported teachers; Amazon Nova Pro and Llama 3.2 1B/3B are supported students. AWS claims up to 500% faster and up to 75% less expensive inference than the original models, with less than 2% accuracy loss for use cases like RAG. For most buyers this is a cheaper path to a task-specific small model than running a distillation pipeline in-house.

    Sources aws.amazon.com · docs.aws.amazon.com

  6. Distillation transfers narrow skills far better than broad knowledge

    The DeepSeek R1 student series is the cleanest natural experiment available. R1-Distill-Qwen-1.5B keeps 86% of R1’s MATH-500 (83.9 vs 97.3-class teacher performance) but only 47% of its GPQA Diamond (33.8 vs 71.5). At 7B the split is still stark: MATH-500 92.8 but GPQA 49.1. Buyers evaluating a tiny distilled model on a maths or format-following eval will systematically overestimate how it behaves on open-domain knowledge work.

    Sources huggingface.co · huggingface.co

  7. Advertised latency and measured latency have diverged because of adaptive thinking

    Vendor tables still say "fastest", but the number a buyer feels now depends on the reasoning effort setting. Artificial Analysis measures GPT-5.6 Luna at 1.70s time-to-first-token at low effort and 19.87s at high; Claude Haiku 4.5 with reasoning on measures 19.92s; Claude Sonnet 5 at max effort measures 177.77s. Gemini 2.5 Flash-Lite in non-reasoning mode measures 0.30s. Any latency SLA written against a reasoning model has to pin the effort level or it is not a specification.

    Sources artificialanalysis.ai · artificialanalysis.ai · artificialanalysis.ai · artificialanalysis.ai

  8. Legacy tiers are the biggest silent line item in most 2026 AI bills

    GPT-4o still lists at $2.50/$10 — the same rate as in 2024 — while GPT-5.6 Luna lists at $0.20/$1.20 with a far larger context window. Claude Sonnet 4.6 costs 50% more than the newer, better Sonnet 5. Gemini 3.5 Flash at $1.50/$9.00 costs twice Gemini 3.8 Flash's $0.75/$3.75 — but that Flash rate is promotional through 2026-12-31 and lists at $1.50/$7.50 from 2027-01-01, i.e. the same input rate as 3.5 Flash. Nothing forces a migration, so pinned model IDs from 2024–25 quietly bill at up to 12x the current market rate for the same job.

    Sources developers.openai.com · platform.claude.com · ai.google.dev

  9. Price stability is not something you can assume any more

    In 2026 alone: OpenAI cut GPT-5.6 Luna 80% and Terra 20% on 30 July, then cut Sol over 20% on 21 August for a three-month window; Anthropic cancelled a scheduled Sonnet 5 increase from $2/$10 to $3/$15 and made the lower price permanent; DeepSeek raised V4 standard rates roughly 3–4.7x on 16 August — and cache-hit input rates by up to 11x — as demand strained capacity. Contract for the workload, not for the price — and keep a second vendor wired up.

    Sources axios.com · platform.claude.com · infoworld.com

  10. Context window is no longer a reason to pay frontier prices

    Every GPT-5.6 tier including the $0.20 Luna carries a 1.05M-token window. Claude Sonnet 5 and Opus 5 both carry 1M. Gemini 3.5 Flash-Lite carries 1M at $0.30/MTok. Llama 4 Scout carries 10M with open weights. Two caveats matter for whole-repository or whole-contract passes: Claude Haiku 4.5 is still capped at 200K, and long prompts are repriced above 200K input tokens — Gemini 3.1 Pro and 2.5 Pro double the input rate above that threshold (and GPT-5.4-class models apply 2x input / 1.5x output), so the cheap headline rate is not the rate you pay on a million-token prompt.

    Sources developers.openai.com · platform.claude.com · ai.google.dev · github.com

Charts · 6 charts

Blended price vs MMLU-Pro

%
The values plotted in “Blended price vs MMLU-Pro”, in %.
PointSeriesBlended price (USD per million tokens, 3:1 input:output)MMLU-Pro (%) %
Gemini 3.1 Pro (Preview)Frontier teachers4.592.6
DeepSeek-V4-ProFrontier teachers1.9887.5
Gemini 3.1 Flash-LiteSmall siblings0.56383
Mistral Small 4Small siblings0.26378
DeepSeek-V4-FlashDocumented distillations0.6686.2
Llama 3.3 70B InstructOpen weights (hosted price)1.0468.9

Only models that publish MMLU-Pro AND have a sourced token price appear here. The frontier is almost flat between $0.26 and $2.00: Mistral Small 4 buys 78.0 MMLU-Pro for $0.26 blended, DeepSeek-V4-Flash 86.2 for $0.66, DeepSeek-V4-Pro 87.5 for $1.98, and Gemini 3.1 Pro 92.6 for $4.50 — so the last 6 MMLU-Pro points cost roughly 7x. Gemma 4 31B (85.2) and Gemma 4 26B A4B (82.6) sit off this chart because they are self-host-only and therefore have no list token price.

Sources: ai.google.dev · deepmind.google · layerlens.ai · api-docs.deepseek.com · huggingface.co · huggingface.co · mistral.ai · openrouter.ai · together.ai · github.com

Blended price vs GPQA Diamond — the whole shortlist on one plot

%
The values plotted in “Blended price vs GPQA Diamond — the whole shortlist on one plot”, in %.
PointSeriesBlended price (USD per million tokens, 3:1 input:output)GPQA Diamond (%) %
GPT-5.6 SolFrontier teachers892.4
GPT-5.4Frontier teachers5.6393
o3Frontier teachers3.583.3
Gemini 3.1 Pro (Preview)Frontier teachers4.594.3
Gemini 2.5 ProFrontier teachers3.4486.4
DeepSeek-V4-ProFrontier teachers1.9890.1
Amazon Nova ProFrontier teachers1.446.9
GPT-5.6 TerraSmall siblings (method undisclosed)4.588.4
GPT-5.6 LunaSmall siblings (method undisclosed)0.4587
GPT-5.4 miniSmall siblings (method undisclosed)1.6988
GPT-5.4 nanoSmall siblings (method undisclosed)0.46382.8
GPT-5 miniSmall siblings (method undisclosed)0.68880.3
GPT-5 nanoSmall siblings (method undisclosed)0.13870.9
o4-miniSmall siblings (method undisclosed)1.9381.4
Gemini 3.1 Flash-LiteSmall siblings (method undisclosed)0.56372.2
Mistral Small 4Small siblings (method undisclosed)0.26371.2
Amazon Nova LiteSmall siblings (method undisclosed)0.10542
Amazon Nova MicroSmall siblings (method undisclosed)0.06140
Gemini 2.5 FlashDocumented distillations0.8582.8
DeepSeek-V4-FlashDocumented distillations0.6688.1
gpt-oss-20bOpen weights0.13158.6
Llama 3.3 70B InstructOpen weights1.0450.5
Qwen3.8-27BOpen weights1.1389.2
Ministral 3 14B InstructOpen weights0.271.2

The upper-left corner is the whole story: GPT-5.6 Luna at $0.45 blended / 87.0 GPQA and DeepSeek-V4-Flash at $0.66 / 88.1 sit close to the frontier for a fraction of the price. Gemini 3.1 Flash-Lite is cheaper still at $0.56 blended but scores 72.2 GPQA Diamond, so it trades roughly 15 points of knowledge benchmark for the saving. Paying 8–18x more than Luna or V4-Flash moves you 5–7 GPQA points. Amazon Nova rows use the 2024 Nova technical report’s GPQA methodology and are not comparable with the reasoning-era scores.

Sources: developers.openai.com · openrouter.ai · openrouter.ai · openrouter.ai · the-decoder.com · datacamp.com · ai.google.dev · arxiv.org · layerlens.ai · api-docs.deepseek.com · huggingface.co · huggingface.co · mistral.ai · huggingface.co · console.groq.com · huggingface.co · assets.amazon.science

Larger vs smaller tier on GPQA Diamond (and MMLU for Nova)

%
The values plotted in “Larger vs smaller tier on GPQA Diamond (and MMLU for Nova)”, in %.
Pair (documented distillations and same-family comparisons)Larger / teacher tier %Smaller tier (distilled or same-family) %
GPT-5.6 Sol → Luna92.487
GPT-5.4 → GPT-5.4 mini9388
o3 → o4-mini83.381.4
Gemini 2.5 Pro → 2.5 Flash86.482.8
DeepSeek V4-Pro vs V4-Flash90.188.1
DeepSeek-R1 → R1-Distill-Llama-70B71.565.2
DeepSeek-R1 → R1-Distill-Qwen-32B71.562.1
DeepSeek-R1 → R1-Distill-Qwen-7B71.549.1
DeepSeek-R1 → R1-Distill-Qwen-1.5B71.533.8
Gemma 4 31B vs Gemma 4 E4B84.358.6
Nova Pro vs Nova Lite85.980.5

Inside a single generation the bars are nearly the same height. Across a large parameter gap they are not: R1 → R1-Distill-Qwen-1.5B loses more than half the teacher's GPQA, and Gemma 4 31B vs E4B loses 30%. Only the DeepSeek-R1 and Gemini 2.5 pairs are vendor-documented distillations; DeepSeek V4-Pro vs V4-Flash, Gemma 4 and Nova are same-family comparisons — V4-Flash is documented as a consolidation of V4 domain experts via on-policy distillation, not a compression of V4-Pro. Nova uses MMLU rather than GPQA because that is what the Amazon technical report publishes.

Sources: openrouter.ai · the-decoder.com · datacamp.com · arxiv.org · huggingface.co · huggingface.co · ai.google.dev · assets.amazon.science

Cost per 1,000,000 requests (400 input + 100 output tokens each)

USD
The values plotted in “Cost per 1,000,000 requests (400 input + 100 output tokens each)”, in USD.
ModelUSD per 1M requests USD
Amazon Nova Micro28
qwen-turbo40
Amazon Nova Lite48
GPT-5 nano60
gpt-oss-20b60
Ministral 3 14B Instruct100
qwen3.8-flash107
GPT-4o mini120
Mistral Small 4120
gpt-oss-120b120
GPT-5.6 Luna200
GPT-5 mini300
DeepSeek-V4-Flash308
Gemini 2.5 Flash370
Gemini 3.5 Flash-Lite370
Amazon Nova 2 Lite370
Llama 3.3 70B Instruct520
Gemini 3.8 Flash675
GPT-5.4 mini750
o4-mini880
Claude Haiku 4.5900
DeepSeek-V4-Pro924
Claude Sonnet 51,800
Gemini 3.1 Pro (Preview)2,000
GPT-5.6 Terra2,000
GPT-5.6 Sol3,600
Claude Opus 54,500
Claude Fable 5.19,000
GPT-6 Astra9,000

A classification or extraction workload at a million requests a month costs $28 on Amazon Nova Micro, $200 on GPT-5.6 Luna, $900 on Claude Haiku 4.5, $1,800 on Claude Sonnet 5 and $9,000 on the two most expensive tiers, Claude Fable 5.1 and GPT-6 Astra (both $10/$50) — a 320x spread across the same shortlist. Batch APIs cut these by 50% at both Anthropic and OpenAI, and prompt caching cuts the input component by up to 90% (97.5% on Claude Fable 5.1), which is usually a bigger lever than moving one tier down.

Sources: developers.openai.com · platform.claude.com · ai.google.dev · api-docs.deepseek.com · mistral.ai · alibabacloud.com · console.groq.com · together.ai · pricepertoken.com · pricepertoken.com · pricepertoken.com

Quality retention: smaller tier score as a percentage of the larger tier

%
The values plotted in “Quality retention: smaller tier score as a percentage of the larger tier”, in %.
Pair (documented distillations and same-family comparisons)Knowledge benchmarks (GPQA-D / MMLU / MMLU-Pro) %Coding / agentic benchmarks %
GPT-5.6 Sol → Terra95.7
GPT-5.6 Sol → Luna94.2
GPT-5.4 → GPT-5.4 mini94.6
GPT-5.4 → GPT-5.4 nano89
o3 → o4-mini97.7
Gemini 2.5 Pro → 2.5 Flash95.8
DeepSeek V4-Pro vs V4-Flash97.8
DeepSeek-R1 → R1-Distill-Llama-70B91.2
DeepSeek-R1 → R1-Distill-Qwen-32B86.9
DeepSeek-R1 → R1-Distill-Qwen-14B82.7
DeepSeek-R1 → R1-Distill-Llama-8B68.5
DeepSeek-R1 → R1-Distill-Qwen-7B68.7
DeepSeek-R1 → R1-Distill-Qwen-1.5B47.3
Gemma 4 31B vs Gemma 4 E4B69.5
Llama 4 Maverick → Scout92.3
Nova Pro vs Nova Lite93.7
Nova Pro vs Nova Micro90.3
Claude Opus 5 → Sonnet 588.8
Claude Opus 5 → Haiku 4.576.4
Llama 4 Maverick → Scout (coding)75.6
DeepSeek V4-Pro vs V4-Flash (coding)98
Gemini 2.5 Pro → 2.5 Flash (coding)89.7

Read the two series against each other. Within one generation, knowledge retention clusters at 94–98%. Coding and agentic retention is bimodal: DeepSeek V4-Flash keeps 98% of V4-Pro on SWE-bench Verified, but Claude Haiku 4.5 keeps only 76% of Opus 5 and Llama 4 Scout only 76% of Maverick on LiveCodeBench. Only the DeepSeek-R1, Gemini 2.5 and Llama 4 rows are vendor-documented distillations; DeepSeek V4-Pro vs V4-Flash, Gemma 4 and Nova are same-family size comparisons (V4-Flash is documented as a consolidation of V4 domain experts, not a compression of V4-Pro). That spread, not the price sheet, is what should decide whether a workload can be moved down a tier.

Sources: huggingface.co · huggingface.co · arxiv.org · morphllm.com · datanorth.ai · github.com · ai.google.dev · assets.amazon.science · openrouter.ai · the-decoder.com

Enterprise inference price index, 2026

USD
The values plotted in “Enterprise inference price index, 2026”, in USD.
DateUSD per million tokens USD
2026-05-312.04
2026-07-251.45
2026-08-071.17

A 43% fall in ten weeks. Jefferies attributes it to OpenAI cutting GPT-5.6 rates by up to 80%, Anthropic matching prior frontier performance at half the price with Claude Opus 5, and Chinese open-weight models such as DeepSeek V4-Flash-0731 competing at roughly $0.03 per task. The 7 August figure is the midpoint of the reported $1.16–$1.18 low.

Sources: scmp.com · axios.com

Tables · 5 tables

Master comparison: 74 models a buyer could shortlist in September 2026

74 rows
Master comparison: 74 models a buyer could shortlist in September 2026 — Price, quality, context and licence for frontier teachers, vendor small siblings and openly-licensed distilled students. "Quality (GPQA-D)" is GPQA Diamond; MMLU-Pro is shown where the vendor publishes it. Blank cells mean the figure is not published — nothing here is estimated. — Units: Input in USD/MTok; Output in USD/MTok; GPQA-D in %; MMLU-Pro in %; Context in K tokens.
ModelVendorRoleDistilled?TeacherInput USD/MTokOutput USD/MTokGPQA-D %MMLU-Pro %CodingContext K tokensLicence
GPT-6 AstraOpenAIFrontier teacherNon/a10501,050Proprietary API developers.openai.com
GPT-5.6 SolOpenAIFrontier teacherNon/a42092.41,050Proprietary API openrouter.ai
GPT-5.6 TerraOpenAISmall siblingUndisclosedundisclosed21288.41,050Proprietary API openrouter.ai
GPT-5.6 LunaOpenAISmall siblingUndisclosedundisclosed0.21.2871,050Proprietary API openrouter.ai
GPT-5.5OpenAIFrontier teacherNon/a530Proprietary API developers.openai.com
GPT-5.4OpenAIFrontier teacherNon/a2.5159357.7 (SWE-bench Pro)1,050Proprietary API the-decoder.com
GPT-5.4 miniOpenAISmall siblingUndisclosedundisclosed0.754.58854.4 (SWE-bench Pro)400Proprietary API the-decoder.com
GPT-5.4 nanoOpenAISmall siblingUndisclosedundisclosed0.21.2582.852.4 (SWE-bench Pro)400Proprietary API the-decoder.com
GPT-5OpenAIFrontier teacherNon/a1.251074.9 (SWE-bench Verified)400Proprietary API arxiv.org
GPT-5 miniOpenAISmall siblingUndisclosedundisclosed0.25280.345.7 (SWE-bench Pro)400Proprietary API openrouter.ai
GPT-5 nanoOpenAISmall siblingUndisclosedundisclosed0.050.470.9400Proprietary API openrouter.ai
GPT-4oOpenAIFrontier teacherNon/a2.510128Proprietary API openrouter.ai
GPT-4o miniOpenAISmall siblingUndisclosedundisclosed0.150.687.2 (HumanEval)128Proprietary API openai.com
o3OpenAIFrontier teacherNon/a2883.369.1 (SWE-bench Verified)200Proprietary API datacamp.com
o4-miniOpenAISmall siblingUndisclosedundisclosed1.14.481.468.1 (SWE-bench Verified)200Proprietary API datacamp.com
gpt-oss-120bOpenAI (open weights)Open weightsUndisclosedundisclosed0.150.6128Apache 2.0 huggingface.co
gpt-oss-20bOpenAI (open weights)Open weightsUndisclosedundisclosed0.0750.358.653.2 (SWE-bench Verified)128Apache 2.0 huggingface.co
Claude Fable 5.1AnthropicFrontier teacherNon/a10501,000Proprietary API platform.claude.com
Claude Opus 5AnthropicFrontier teacherNon/a52596 (SWE-bench Verified)1,000Proprietary API datanorth.ai
Claude Sonnet 5AnthropicSmall siblingUndisclosedundisclosed21085.2 (SWE-bench Verified)1,000Proprietary API morphllm.com
Claude Sonnet 4.6AnthropicSmall siblingUndisclosedundisclosed3151,000Proprietary API platform.claude.com
Claude Haiku 4.5AnthropicSmall siblingUndisclosedundisclosed1573.3 (SWE-bench Verified)200Proprietary API anthropic.com
Claude Haiku 3.5AnthropicSmall siblingUndisclosedundisclosed0.84200Proprietary API platform.claude.com
Gemini 3.1 Pro (Preview)GoogleFrontier teacherNon/a21294.392.680.6 (SWE-bench Verified)1,000Proprietary API deepmind.google
Gemini 3.8 FlashGoogleSmall siblingUndisclosedundisclosed0.753.7573.7 (DeepSWE v1.1)1,000Proprietary API deepmind.google
Gemini 3.5 FlashGoogleSmall siblingUndisclosedundisclosed1.591,000Proprietary API ai.google.dev
Gemini 3.5 Flash-LiteGoogleSmall siblingUndisclosedundisclosed0.32.51,000Proprietary API ai.google.dev
Gemini 3.1 Flash-LiteGoogleSmall siblingUndisclosedundisclosed0.251.572.2831,000Proprietary API layerlens.ai
Gemini 2.5 ProGoogleFrontier teacherNon/a1.251086.467.2 (SWE-bench Verified)1,000Proprietary API arxiv.org
Gemini 2.5 FlashGoogleDocumented distillationYes (documented)Gemini 2.5 Pro (k-sparse logit distillation)0.32.582.860.3 (SWE-bench Verified)1,000Proprietary API arxiv.org
Gemini 2.5 Flash-LiteGoogleDocumented distillationYes (documented)Gemini 2.5 Pro (k-sparse logit distillation)0.10.41,000Proprietary API arxiv.org
Gemma 4 31BGoogleOpen weightsUndisclosedundisclosed84.385.280 (LiveCodeBench v6)256Apache 2.0 ai.google.dev
Gemma 4 26B A4B (MoE)GoogleOpen weightsUndisclosedundisclosed82.382.677.1 (LiveCodeBench v6)256Apache 2.0 ai.google.dev
Gemma 4 12B UnifiedGoogleOpen weightsUndisclosedundisclosed78.877.272 (LiveCodeBench v6)256Apache 2.0 ai.google.dev
Gemma 4 E4BGoogleOpen weightsUndisclosedundisclosed58.669.452 (LiveCodeBench v6)128Apache 2.0 ai.google.dev
Gemma 4 E2BGoogleOpen weightsUndisclosedundisclosed43.46044 (LiveCodeBench v6)128Apache 2.0 ai.google.dev
Gemma 3 27B ITGoogleOpen weightsUndisclosedundisclosed24.348.8 (HumanEval)128Gemma Terms of Use (custom) huggingface.co
Gemma 3 4B ITGoogleOpen weightsUndisclosedundisclosed1536 (HumanEval)128Gemma Terms of Use (custom) huggingface.co
Llama 4 MaverickMetaDocumented distillationYes (documented)Llama 4 Behemoth (codistillation)69.880.543.4 (LiveCodeBench)1,000Llama 4 Community License (custom commercial) github.com
Llama 4 ScoutMetaDocumented distillationYes (documented)Llama 4 Behemoth (codistillation)57.274.332.8 (LiveCodeBench)10,000Llama 4 Community License (custom commercial) github.com
Llama 3.3 70B InstructMetaOpen weightsUndisclosedundisclosed1.041.0450.568.988.4 (HumanEval)128Llama 3.3 Community License (custom commercial) github.com
Llama 3.2 3B InstructMetaDocumented distillationYes (documented)Llama 3.1 8B and 70B (token-level logit distillation after pruning)32.8128Llama 3.2 Community License (custom commercial) github.com
Llama 3.2 1B InstructMetaDocumented distillationYes (documented)Llama 3.1 8B and 70B (token-level logit distillation after pruning)27.2128Llama 3.2 Community License (custom commercial) github.com
DeepSeek-V4-ProDeepSeekFrontier teacherNon/a1.323.9690.187.580.6 (SWE-bench Verified)1,000MIT huggingface.co
DeepSeek-V4-FlashDeepSeekDocumented distillationYes (documented)DeepSeek V4 domain experts (on-policy distillation consolidation)0.441.3288.186.279 (SWE-bench Verified)1,000MIT huggingface.co
DeepSeek-V3.2DeepSeekOpen weightsYes (documented)DeepSeek specialist models (specialist distillation into the generalist)82.48573.1 (SWE-bench Verified)128MIT arxiv.org
DeepSeek-R1 (0528)DeepSeekFrontier teacherNon/a818573.3 (LiveCodeBench)128MIT huggingface.co
DeepSeek-R1-Distill-Llama-70BDeepSeek / Meta baseDocumented distillationYes (documented)DeepSeek-R165.257.5 (LiveCodeBench)128MIT (weights) over Llama 3.3 Community License base huggingface.co
DeepSeek-R1-Distill-Qwen-32BDeepSeek / Qwen baseDocumented distillationYes (documented)DeepSeek-R162.157.2 (LiveCodeBench)128MIT (weights), Qwen2.5-32B base under Apache 2.0 huggingface.co
DeepSeek-R1-Distill-Qwen-14BDeepSeek / Qwen baseDocumented distillationYes (documented)DeepSeek-R159.153.1 (LiveCodeBench)128MIT (weights), Qwen2.5-14B base under Apache 2.0 huggingface.co
DeepSeek-R1-Distill-Llama-8BDeepSeek / Meta baseDocumented distillationYes (documented)DeepSeek-R14939.6 (LiveCodeBench)128MIT (weights) over Llama 3.1 Community License base huggingface.co
DeepSeek-R1-Distill-Qwen-7BDeepSeek / Qwen baseDocumented distillationYes (documented)DeepSeek-R149.137.6 (LiveCodeBench)128MIT (weights), Qwen2.5-Math-7B base under Apache 2.0 huggingface.co
DeepSeek-R1-Distill-Qwen-1.5BDeepSeek / Qwen baseDocumented distillationYes (documented)DeepSeek-R133.816.9 (LiveCodeBench)128MIT (weights), Qwen2.5-Math-1.5B base under Apache 2.0 huggingface.co
Qwen3.8-27BAlibabaOpen weightsUndisclosedundisclosed0.5389.261.7 (SWE-bench Pro)262Apache 2.0 huggingface.co
qwen3.8-maxAlibabaFrontier teacherNon/a26Proprietary API alibabacloud.com
qwen3.8-flashAlibabaSmall siblingUndisclosedundisclosed0.150.47Proprietary API alibabacloud.com
qwen-turboAlibabaSmall siblingUndisclosedundisclosed0.050.2Proprietary API alibabacloud.com
Qwen3-4B-Instruct-2507AlibabaDocumented distillationYes (documented)Qwen3-32B / Qwen3-235B-A22B (off-policy + on-policy strong-to-weak distillation)6269.635.1 (LiveCodeBench v6)262Apache 2.0 huggingface.co
qwen3-8b (hosted)AlibabaDocumented distillationYes (documented)Qwen3-32B / Qwen3-235B-A22B (strong-to-weak distillation)0.180.7128Apache 2.0 arxiv.org
Phi-4 (14B)MicrosoftOpen weightsUndisclosedundisclosed56.170.482.6 (HumanEval)16MIT huggingface.co
Phi-4-mini-instruct (3.8B)MicrosoftOpen weightsUndisclosedundisclosed25.252.8128MIT huggingface.co
Mistral Medium 3.5Mistral AIFrontier teacherNon/a1.57.5Modified MIT docs.mistral.ai
Mistral Large 3Mistral AIOpen weightsUndisclosedundisclosed0.51.5Apache 2.0 docs.mistral.ai
Mistral Small 4Mistral AISmall siblingUndisclosedundisclosed0.150.671.278Apache 2.0 openrouter.ai
Ministral 3 14B InstructMistral AIOpen weightsUndisclosedundisclosed0.20.271.264.6 (LiveCodeBench)256Apache 2.0 huggingface.co
Ministral 3 8BMistral AIOpen weightsUndisclosedundisclosed0.150.15256Apache 2.0 docs.mistral.ai
Ministral 3 3BMistral AIOpen weightsUndisclosedundisclosed0.10.1256Apache 2.0 docs.mistral.ai
Amazon Nova PremierAmazonFrontier teacherNon/a1,000Proprietary API docs.aws.amazon.com
Amazon Nova ProAmazonFrontier teacherUndisclosed (Bedrock distillation student)undisclosed0.83.246.9300Proprietary API assets.amazon.science
Amazon Nova LiteAmazonSmall tierUndisclosed (Bedrock distillation student)undisclosed0.060.2442300Proprietary API assets.amazon.science
Amazon Nova MicroAmazonSmall tierUndisclosed (Bedrock distillation student)undisclosed0.0350.1440128Proprietary API assets.amazon.science
Amazon Nova 2 LiteAmazonSmall siblingUndisclosedundisclosed0.32.51,000Proprietary API docs.aws.amazon.com
SmolLM3-3BHugging FaceOpen weightsUndisclosedn/a41.730.48 (HumanEval+)128Apache 2.0 huggingface.co
grok-4.6xAIFrontier teacherNon/a26200Proprietary API docs.x.ai

Prices are vendor list rates in USD per million tokens, before batch (typically -50%) or cache discounts. Open-weight rows with no price are self-host only in this dataset. DeepSeek prices are peak-hour; off-peak is half. Gemini 3.1 Pro and 2.5 Pro input prices double above 200K input tokens. Gemini 3.8 Flash's $0.75/$3.75 is promotional through 2026-12-31; it lists at $1.50/$7.50 from 2027-01-01. Qwen3.8-27B is shown at Alibaba Model Studio's list rate; Groq hosts the same weights at $0.80/$4.00. Gemini 3.7 Flash shipped in August 2026 but is not included in this shortlist.

Sources: developers.openai.com · platform.claude.com · ai.google.dev · api-docs.deepseek.com · mistral.ai · alibabacloud.com · together.ai · console.groq.com · ai.google.dev · github.com · huggingface.co · assets.amazon.science

Price vs quality: GPQA Diamond points per dollar of blended token price

24 rows
Price vs quality: GPQA Diamond points per dollar of blended token price — Blended price uses the 3:1 input:output weighting (0.75 x input + 0.25 x output). "GPQA per
quot; is the crude but decisive ranking metric for knowledge-heavy workloads. Cost per 1M requests assumes 400 input + 100 output tokens per request at list price with no caching. — Units: GPQA Diamond in %; Blended price in USD/MTok; GPQA points per $ in pts/USD; Cost / 1M requests in USD.
ModelVendorTierGPQA Diamond %Blended price USD/MTokGPQA points per $ pts/USDCost / 1M requests USD
Amazon Nova MicroAmazonSmall tier400.06165328 assets.amazon.science
GPT-5 nanoOpenAISmall sibling70.90.13851660 openrouter.ai
gpt-oss-20bOpenAI (open weights)Open weights58.60.13144760 huggingface.co
Amazon Nova LiteAmazonSmall tier420.10540048 assets.amazon.science
Ministral 3 14B InstructMistral AIOpen weights71.20.2356100 huggingface.co
Mistral Small 4Mistral AISmall sibling71.20.263271120 openrouter.ai
GPT-5.6 LunaOpenAISmall sibling870.45193200 openrouter.ai
GPT-5.4 nanoOpenAISmall sibling82.80.463179205 the-decoder.com
Gemini 3.1 Flash-LiteGoogleSmall sibling72.20.563128250 layerlens.ai
DeepSeek-V4-FlashDeepSeekDistilled88.10.66134308 huggingface.co
GPT-5 miniOpenAISmall sibling80.30.688117300 openrouter.ai
Gemini 2.5 FlashGoogleDistilled82.80.8597.4370 arxiv.org
Qwen3.8-27BAlibabaOpen weights89.21.1379.3500 huggingface.co
GPT-5.4 miniOpenAISmall sibling881.6952.1750 the-decoder.com
Llama 3.3 70B InstructMetaOpen weights50.51.0448.6520 github.com
DeepSeek-V4-ProDeepSeekTeacher90.11.9845.5924 huggingface.co
o4-miniOpenAISmall sibling81.41.9342.3880 openrouter.ai
Amazon Nova ProAmazonTeacher46.91.433.5640 assets.amazon.science
o3OpenAITeacher83.33.523.81,600 datacamp.com
Gemini 2.5 ProGoogleTeacher86.43.4425.11,500 arxiv.org
Gemini 3.1 Pro (Preview)GoogleTeacher94.34.5212,000 deepmind.google
GPT-5.6 TerraOpenAISmall sibling88.44.519.62,000 openrouter.ai
GPT-5.4OpenAITeacher935.6316.52,500 the-decoder.com
GPT-5.6 SolOpenAITeacher92.4811.63,600 openrouter.ai

GPT-5 nano tops this ranking on raw efficiency but scores only 70.9 GPQA; GPT-5.6 Luna is the highest-scoring model in the top three, which is why it dominates most 2026 routing configurations. Amazon Nova figures use MMLU-era GPQA methodology from the Nova technical report and are not directly comparable with the 2026 reasoning-model scores.

Sources: developers.openai.com · openrouter.ai · api-docs.deepseek.com · huggingface.co · ai.google.dev · arxiv.org · assets.amazon.science · console.groq.com

Measured latency and throughput (Artificial Analysis, September 2026)

12 rows
Measured latency and throughput (Artificial Analysis, September 2026) — Independently measured time-to-first-token and output speed. On adaptive-thinking models TTFT includes reasoning time, so the effort setting is stated for every row — without it these numbers are not comparable. — Units: Time to first token in s; Output speed in tokens/s; Blended price in USD/MTok.
Model (setting)Time to first token sOutput speed tokens/sAA Intelligence IndexBlended price USD/MTok
Gemini 2.5 Flash-Lite (non-reasoning)0.30.18 artificialanalysis.ai
gpt-oss-120b (high)0.85151240.2 artificialanalysis.ai
Llama 4 Maverick0.9282140.31 artificialanalysis.ai
DeepSeek V4-Flash 0731 (reasoning, max)1.19140520.23 artificialanalysis.ai
GPT-5.6 Luna (low)1.7109340.17 artificialanalysis.ai
DeepSeek V4-Pro 0813 (reasoning, max)1.960.2530.69 artificialanalysis.ai
Gemini 3.5 Flash-Lite6.48391370.33 artificialanalysis.ai
GPT-5.6 Luna (high)19.9123470.17 artificialanalysis.ai
Claude Haiku 4.5 (reasoning)19.990300.77 artificialanalysis.ai
Claude Opus 5 (adaptive, max effort)77.357633.85 artificialanalysis.ai
Claude Sonnet 5 (adaptive, max effort)17878551.54 artificialanalysis.ai
Amazon Nova 2 Lite (non-reasoning)149120.85 pricepertoken.com

Gemini 3.5 Flash-Lite is the throughput leader at 391 tokens/s. GPT-5.6 Luna is the only model here that spans both ends of the latency range purely through its effort parameter (1.70s to 19.87s), which makes it unusually easy to run interactive and batch traffic on one model ID. Claude Sonnet 5’s 177.77s figure is max-effort adaptive thinking, not a typical production setting.

Sources: artificialanalysis.ai · artificialanalysis.ai · artificialanalysis.ai · artificialanalysis.ai · artificialanalysis.ai · artificialanalysis.ai

Licence and deployment: what you are allowed to do with each family

14 rows
Licence and deployment: what you are allowed to do with each family — The question that decides procurement before any benchmark does. "Distil outputs?" means: may you legally train your own model on this model’s outputs.
FamilyLicenceWeightsSelf-host / air-gapDistil outputs?Data residency optionsBuyer note
OpenAI GPT-5.x / GPT-6 / o-seriesProprietary APIClosedNoNo — Services Agreement bars using Output to develop competing modelsAzure/Foundry regionsDeepest price ladder in the market and 1.05M context on every 5.6 tier. openai.com
OpenAI gpt-oss 20b / 120bApache 2.0OpenYes (120b on one 80GB GPU; 20b in 16GB)YesAnywhereThe escape hatch inside the OpenAI ecosystem. Served by Groq at ~500–1000 tok/s. huggingface.co
Anthropic Claude (Fable / Opus / Sonnet / Haiku)Proprietary APIClosedNoNo — Commercial Terms D.4 bars building a competing product or training competing AI modelsinference_geo:"us" at a 1.1x multiplier; Bedrock/Vertex regional endpoints at +10%Best published SWE-bench Verified (Opus 5, 96.0). 1M context on Fable/Opus/Sonnet, 200K on Haiku 4.5. anthropic.com
Google Gemini 2.5 / 3.xProprietary APIClosedNoNo — Gemini API Additional Terms restrict competitive model developmentVertex AI regionsOnly closed vendor that has publicly documented distilling its own small tier (Gemini 2.5 report). arxiv.org
Google Gemma 4Apache 2.0OpenYes (2.3B–31B)YesAnywhereMMLU-Pro 85.2 at 31B under Apache 2.0 — the strongest permissively-licensed model you can put on one node. ai.google.dev
Google Gemma 3Gemma Terms of Use (custom, use restrictions apply)OpenYesYes, subject to the Gemma prohibited-use policyAnywhereSuperseded by Gemma 4 on both quality and licence terms. huggingface.co
Meta Llama 3.x / 4Llama Community License (custom commercial; 700M MAU clause)OpenYesYes, with Llama attribution and naming obligations on derivativesAnywhereLlama 3.2 1B/3B are explicitly documented distillations; Meta's Llama 4 launch post describes Scout/Maverick as codistilled from the unreleased Behemoth (the model card itself does not mention it). ai.meta.com
DeepSeek R1 / V3.2 / V4MITOpenYes (V4-Pro is a 1.6T MoE — non-trivial)Yes — the R1 release shipped six distilled students itselfAnywhere; first-party API is PRC-hostedHighest open-weight GPQA Diamond here (90.1). Many enterprises self-host rather than use the PRC-hosted API. huggingface.co
Alibaba Qwen3 / Qwen3.8 (open weights)Apache 2.0OpenYesYesAnywhere; Model Studio API is PRC/SingaporeQwen3 technical report documents strong-to-weak distillation for the 0.6B–14B dense sizes and 30B-A3B. huggingface.co
Microsoft Phi-4 / Phi-4-miniMITOpenYesYesAnywhereLeast legally encumbered small models in this dataset. Phi-4’s 16K context is the catch. huggingface.co
Mistral Small 4 / Ministral 3 / Large 3Apache 2.0OpenYesYesEU-headquartered vendor; EU hosting availableThe default answer when the requirement is EU sovereignty plus a permissive licence. docs.mistral.ai
Mistral Medium 3.5Modified MITOpenYesYes, subject to the modified termsEU hosting availableFrontier-class tier of the Mistral line; read the modification before assuming MIT. docs.mistral.ai
Amazon Nova / Nova 2Proprietary APIClosedNoOnly through Bedrock Model Distillation, into another Amazon-supported studentAWS regions incl. GovCloud (US-West)The only vendor that documents its teacher/student graph in product docs: Premier → Pro/Lite/Micro, Pro → Lite/Micro. docs.aws.amazon.com
Hugging Face SmolLM3Apache 2.0Open (plus data mixture and training configs)YesYesAnywhereFully reproducible supply chain — relevant where an auditor asks what the model was trained on. huggingface.co

Anti-distillation clauses bind the enterprise account holder, not just individual developers, and survive termination in most of these agreements. If your roadmap includes training an in-house model on teacher outputs, pick an MIT or Apache 2.0 teacher up front rather than seeking a waiver later.

Sources: anthropic.com · openai.com · huggingface.co · ai.google.dev · github.com · huggingface.co · huggingface.co · huggingface.co · docs.mistral.ai · docs.aws.amazon.com · huggingface.co

Quality retention by pair (documented distillations and same-family size comparisons)

22 rows
Quality retention by pair (documented distillations and same-family size comparisons) — What fraction of the teacher’s score the cheaper model keeps, and what fraction of the teacher’s input price it costs. Rows where the student is documented as a distillation are the ones with a named teacher in the master table; the rest are same-family price ladders. — Units: Teacher in %; Student in %; Retention in %; Student price / teacher price (input) in x.
Teacher → studentRelationshipBenchmarkTeacher %Student %Retention %Student price / teacher price (input) x
GPT-5.6 Sol → TerraSame-generation tier comparison (method undisclosed)GPQA Diamond92.488.495.70.5 openrouter.ai
GPT-5.6 Sol → LunaSame-generation tier comparison (method undisclosed); third-party benchmark, 94.2–97.3% across providersGPQA Diamond92.48794.20.05 openrouter.ai
GPT-5.4 → GPT-5.4 miniSame-generation tier comparison (method undisclosed)GPQA Diamond938894.60.3 the-decoder.com
GPT-5.4 → GPT-5.4 nanoSame-generation tier comparison (method undisclosed)GPQA Diamond9382.8890.08 the-decoder.com
o3 → o4-miniSame-generation tier comparison (method undisclosed)GPQA Diamond83.381.497.70.55 datacamp.com
Gemini 2.5 Pro → 2.5 FlashDocumented distillationGPQA Diamond86.482.895.80.24 arxiv.org
DeepSeek V4-Pro vs V4-FlashSame-family price/quality comparison — V4-Flash is a consolidation of V4 domain experts via on-policy distillation, not a compression of V4-ProGPQA Diamond90.188.197.80.333 huggingface.co
DeepSeek-R1 → R1-Distill-Llama-70BDocumented distillationGPQA Diamond71.565.291.2huggingface.co
DeepSeek-R1 → R1-Distill-Qwen-32BDocumented distillationGPQA Diamond71.562.186.9huggingface.co
DeepSeek-R1 → R1-Distill-Qwen-14BDocumented distillationGPQA Diamond71.559.182.7huggingface.co
DeepSeek-R1 → R1-Distill-Llama-8BDocumented distillationGPQA Diamond71.54968.5huggingface.co
DeepSeek-R1 → R1-Distill-Qwen-7BDocumented distillationGPQA Diamond71.549.168.7huggingface.co
DeepSeek-R1 → R1-Distill-Qwen-1.5BDocumented distillationGPQA Diamond71.533.847.3huggingface.co
Gemma 4 31B vs Gemma 4 E4BSame-family size comparison — Google does not describe E4B as a distillation of 31BGPQA Diamond84.358.669.5ai.google.dev
Llama 4 Maverick → ScoutDocumented codistillation (both from Behemoth), compared here by sizeMMLU-Pro80.574.392.3github.com
Nova Pro vs Nova LiteSame-family size comparison — Bedrock supports both as distillation students; shipped provenance undisclosedMMLU85.980.593.70.075 assets.amazon.science
Nova Pro vs Nova MicroSame-family size comparison — Bedrock supports both as distillation students; shipped provenance undisclosedMMLU85.977.690.30.044 assets.amazon.science
Claude Opus 5 → Sonnet 5Same-generation tier comparison (method undisclosed)SWE-bench Verified9685.288.80.4 morphllm.com
Claude Opus 5 → Haiku 4.5Cross-generation tier comparison (method undisclosed)SWE-bench Verified9673.376.40.2 datanorth.ai
Llama 4 Maverick → Scout (coding)Documented codistillation (both from Behemoth), compared here by sizeLiveCodeBench43.432.875.6github.com
DeepSeek V4-Pro vs V4-Flash (coding)Same-family price/quality comparison — see row 6SWE-bench Verified80.679980.333 huggingface.co
Gemini 2.5 Pro → 2.5 Flash (coding)Documented distillationSWE-bench Verified67.260.389.70.24 arxiv.org

Two patterns stand out. First, knowledge retention above 94% is now routine inside a family, and DeepSeek V4-Flash reaches 97.8% for a third of the price. Second, retention falls off a cliff for very small students on broad-knowledge benchmarks — R1-Distill-Qwen-1.5B keeps only 47% of R1’s GPQA — while the same model keeps 86% of R1’s MATH-500. Distillation buys you narrow competence cheaply and broad competence expensively. Only rows marked "Documented distillation" carry a vendor statement that the student was trained from the teacher. The remaining rows are same-family or same-generation price/quality comparisons and should not be read as training provenance.

Sources: openrouter.ai · the-decoder.com · arxiv.org · huggingface.co · huggingface.co · ai.google.dev · github.com · assets.amazon.science · morphllm.com · datanorth.ai · datacamp.com

Timeline · 21 events

  1. product

    GPT-4o mini launches at $0.15/$0.60

    MMLU 82.0, HumanEval 87.2, MMMU 59.4 — over 60% cheaper than GPT-3.5 Turbo. Establishes the "cheap tier" as a permanent product line every vendor now copies.

    Source: openai.com
  2. product

    Llama 3.2 1B/3B ship as documented distillations

    Meta states that logits from Llama 3.1 8B and 70B were used as token-level targets in pretraining, with distillation applied after pruning to recover performance. The first mainstream model card to spell out the recipe.

    Source: github.com
  3. product

    Amazon Bedrock Model Distillation announced

    Turns distillation into a managed service: the customer supplies prompts, AWS generates teacher responses and fine-tunes the student. Later reaches GA with claims of up to 500% faster and 75% cheaper inference.

    Source: aws.amazon.com
  4. product

    Llama 3.3 70B released

    MMLU 86.0 CoT, MMLU-Pro 68.9, HumanEval 88.4, 128K context under the Llama 3.3 Community License. Becomes the default open student base for enterprise distillation projects.

    Source: github.com
  5. product

    DeepSeek ships R1 plus six distilled students under MIT

    R1-Distill-Qwen-32B beats o1-mini on AIME 2024 (72.6 vs 63.6), MATH-500 (94.3 vs 90.0) and GPQA Diamond (62.1 vs 60.0). For buyers this was the first time a free download credibly replaced a paid reasoning tier.

    Source: huggingface.co
  6. product

    Llama 4 Scout and Maverick ship as codistilled models

    Meta's launch post says both are codistilled from the ~2T-parameter Llama 4 Behemoth, which was never released; the Llama 4 model card carries the benchmarks but not the provenance claim. Maverick posts MMLU-Pro 80.5 / GPQA-D 69.8; Scout adds a 10M-token context window.

    Source: ai.meta.com
  7. research

    Gemini 2.5 report confirms the Flash line is distilled

    Google states the smaller 2.5 models are distilled and that the teacher’s next-token distribution is approximated with a k-sparse distribution to cut storage cost. Rare public confirmation from a closed-model vendor.

    Source: arxiv.org
  8. product

    GPT-5, GPT-5 mini and GPT-5 nano launch together

    A three-tier release at $1.25/$10, $0.25/$2 and $0.05/$0.40 with 400K context on all three. Tiered families become the default shape of a frontier launch.

    Source: developers.openai.com
  9. product

    Claude Haiku 4.5 released at $1/$5

    SWE-bench Verified 73.3. Anthropic positions it as matching Claude Sonnet 4 coding performance at a third of the cost and more than twice the speed.

    Source: anthropic.com
  10. product

    Ministral 3 (3B/8B/14B) ships Apache 2.0 with 256K context

    The 14B posts GPQA Diamond 71.2, AIME25 85.0 and MATH 90.4 — edge-class weights with no licence friction, which matters for EU buyers.

    Source: huggingface.co
  11. product

    Mistral Small 4 released under Apache 2.0

    MMLU-Pro 78.0, GPQA Diamond 71.2, priced at $0.15/$0.60 on the Mistral API. Becomes the reference EU-sovereign small model.

    Source: openrouter.ai
  12. product

    Gemma 4 ships Apache 2.0 with frontier-class small models

    The 31B dense posts MMLU-Pro 85.2 and GPQA Diamond 84.3; the 26B MoE reaches 82.3 GPQA with only 3.8B active parameters. Google moves the Gemma line from a custom licence to Apache 2.0.

    Source: deepmind.google
  13. product

    DeepSeek-V4-Pro released under MIT

    1.6T total / 49B active parameters, 1M context, MMLU-Pro 87.5, GPQA Diamond 90.1, SWE-bench Verified 80.6. The strongest openly-licensed model a buyer can self-host.

    Source: huggingface.co
  14. product

    Claude Sonnet 5 launches at $2/$10 with 1M context

    SWE-bench Verified 85.2, Terminal-Bench 2.1 80.4 — close to Opus 4.8 at a fraction of the price. Introductory pricing later made permanent.

    Source: anthropic.com
  15. product

    GPT-5.6 Sol, Terra and Luna launch as one price ladder

    All three carry a 1.05M-token context. GPQA Diamond 92.4 / 88.4 / 87.0 across a 20x price spread — the clearest published price-vs-quality ladder in the market.

    Source: openrouter.ai
  16. product

    Claude Opus 5 released at $5/$25

    SWE-bench Verified 96.0, SWE-bench Pro 79.2, OSWorld 2.0 70.6. Anthropic positions it as near-Fable frontier quality at half the price, holding the Opus price flat.

    Source: platform.claude.com
  17. market

    OpenAI cuts GPT-5.6 Luna by 80% and Terra by 20%

    Luna drops to $0.20/$1.20. This single move reset the floor for closed-model pricing and is a major driver of the mid-2026 fall in the enterprise inference index.

    Source: axios.com
  18. market

    Enterprise inference index hits a 2026 low of $1.16–$1.18/MTok

    Silicon Data index cited by Jefferies, down from $2.04 on 31 May and $1.45 in late July, driven by OpenAI’s cuts, Anthropic’s Opus 5 repricing and Chinese open-weight competition.

    Source: scmp.com
  19. market

    Anthropic makes Sonnet 5 $2/$10 permanent

    The scheduled 1 September increase to $3/$15 is cancelled. A rare case of a vendor withdrawing an announced price rise under competitive pressure.

    Source: platform.claude.com
  20. market

    DeepSeek raises V4 standard rates ~3–4.7x, and cache-hit input rates by up to 11x

    Standard rates: V4-Flash goes from $0.14/$0.28 to $0.44/$1.32 at peak (about 3.1x input, 4.7x output); V4-Pro from $0.435/$0.87 to $1.32/$3.96 (about 3.0x / 4.6x). InfoWorld's "more than 10x" headline refers specifically to cache-hit input tokens, which rose between 52% and 1,100%. Capacity, not competition, sets the floor for the cheapest tiers.

    Source: infoworld.com
  21. product

    Gemini 3.8 Flash ships at $0.75/$3.75

    HLE-Verified 54.9 and 1M context per 9to5Google; Terminal-Bench 2.1 89.4 and DeepSWE v1.1 73.7 per Google's Gemini Flash model page (https://deepmind.google/models/gemini/flash/). Google's third Flash update in three months — the previous model, Gemini 3.7 Flash, shipped three weeks earlier — evidence that the cheap tier is now the fastest-moving part of the market.

    Source: 9to5google.com

Glossary · 16 terms

Knowledge distillation
Training a small "student" model to reproduce the behaviour of a large "teacher" — either from its output text (black-box) or from its output probability distribution over tokens (white-box logit distillation).
Teacher / student
The large source model and the small target model in a distillation. AWS documents Nova Premier as a teacher to Nova Pro, Lite and Micro, and Nova Pro as a teacher to Lite and Micro.
Quality retention
The student’s benchmark score as a percentage of the teacher’s on the same benchmark. Useful for buyers because it is scale-free, but it varies enormously by benchmark: broad knowledge retains worse than narrow maths.
Small sibling
A cheaper tier released alongside a frontier model (mini, nano, Flash, Haiku, Luna) where the vendor has not publicly documented how it was built. Behaves like a distillation commercially whether or not it is one technically.
k-sparse distillation
Storing only the top-k entries of the teacher’s next-token probability distribution instead of the full vocabulary, to make logit distillation affordable at scale. Documented in the Gemini 2.5 technical report.
Codistillation
Training several student models jointly against a shared teacher during the teacher’s own training run. Meta used this for Llama 4 Scout and Maverick against Llama 4 Behemoth.
Strong-to-weak distillation
Alibaba’s two-phase Qwen3 pipeline: off-policy distillation on teacher outputs in both thinking and non-thinking modes, then on-policy distillation aligning student logits with a Qwen3-32B or 235B-A22B teacher by KL divergence.
GPQA Diamond
A 198-question set of graduate-level science problems written to be resistant to web search. The most commonly quoted knowledge benchmark for 2025–26 frontier models.
MMLU-Pro
A harder, ten-choice successor to MMLU with more reasoning-heavy questions. Scores are typically 10–20 points below MMLU for the same model, so the two are not interchangeable in a comparison table.
SWE-bench Verified
A 500-issue human-validated subset of SWE-bench measuring whether a model can resolve a real GitHub issue end to end. The benchmark where distilled tiers lose the most ground.
Time to first token (TTFT)
Latency from request to first streamed token. On reasoning models it now includes thinking time, so a "fast" model at high effort can measure slower than a "slow" model at low effort.
Blended price
A single price per million tokens computed as 0.75 x input + 0.25 x output, the 3:1 weighting Artificial Analysis uses. Handy for ranking, misleading for workloads with unusual input:output ratios.
Prompt caching
Charging a reduced rate for repeated prefix tokens. Cache reads cost 10% of base input on most Claude models and 2.5% on Fable 5.1; OpenAI caches at roughly 10% of input. Often a larger saving than switching model tiers.
Anti-distillation clause
Contract language barring customers from using a vendor’s outputs to train a competing model. Anthropic’s Commercial Terms D.4 and OpenAI’s Services Agreement both contain one; MIT and Apache 2.0 open-weight models do not.
Model tiering / routing
Sending most traffic to a cheap model and escalating only hard requests to a frontier tier. The dominant 2026 cost-control pattern and the main practical way distillation shows up on a buyer’s invoice.
Effective parameters
For MoE and Matformer-style models, the parameters actually activated per token (e.g. Gemma 4 26B A4B activates 3.8B of 25.2B; DeepSeek-V4-Flash activates 13B of 284B). Determines serving cost far more than total parameter count.

Sources · 78 sources

Every figure on this page comes from one of these primary sources. Compiled 4 September 2026.

  1. OpenAI API pricingOpenAI · September 2026 · pricing
  2. OpenAI API models referenceOpenAI · September 2026 · docs
  3. GPT-4o mini: advancing cost-efficient intelligenceOpenAI · 18 July 2024 · blog
  4. OpenAI GPT-5 System CardOpenAI / arXiv · January 2026 · paper
  5. OpenAI Services AgreementOpenAI · 1 January 2026 · law
  6. GPT-5.6 Sol model pageOpenRouter · 9 July 2026 · docs
  7. GPT-5.6 Terra model pageOpenRouter · 9 July 2026 · docs
  8. GPT-5.6 Luna model pageOpenRouter · 9 July 2026 · docs
  9. GPT-5 mini model pageOpenRouter · 7 August 2025 · docs
  10. GPT-5 nano model pageOpenRouter · 7 August 2025 · docs
  11. GPT-4o model pageOpenRouter · 13 May 2024 · docs
  12. o4-mini model pageOpenRouter · 16 April 2025 · docs
  13. OpenAI ships GPT-5.4 mini and nanoThe Decoder · February 2026 · news
  14. o4-mini: tests, features, o3 comparison, benchmarksDataCamp · April 2025 · news
  15. gpt-oss-20b model cardOpenAI / Hugging Face · August 2025 · docs
  16. Claude platform pricingAnthropic · September 2026 · pricing
  17. Claude models overviewAnthropic · September 2026 · docs
  18. Claude Opus 5 model pageAnthropic · 24 July 2026 · docs
  19. Introducing Claude Sonnet 5Anthropic · 30 June 2026 · blog
  20. Introducing Claude Haiku 4.5Anthropic · 15 October 2025 · blog
  21. Anthropic Commercial Terms of ServiceAnthropic · 17 June 2025 · law
  22. Claude Opus 5 by Anthropic: benchmarks and pricingDataNorth · July 2026 · news
  23. Claude benchmarks 2026Morph · September 2026 · news
  24. Anthropic launches Opus 5TechCrunch · 24 July 2026 · news
  25. Gemini API pricingGoogle · September 2026 · pricing
  26. Gemini API modelsGoogle · September 2026 · docs
  27. Gemini 3.1 Pro model pageGoogle DeepMind · February 2026 · docs
  28. Gemini Flash model pageGoogle DeepMind · 2 September 2026 · docs
  29. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context and Next Generation Agentic CapabilitiesGoogle DeepMind / arXiv · July 2025 · paper
  30. Gemma 4 model cardGoogle · 2 April 2026 · docs
  31. Gemma 4Google DeepMind · 2 April 2026 · blog
  32. gemma-3-27b-it model cardGoogle / Hugging Face · March 2025 · docs
  33. Gemini 3.8 Flash rolling out three weeks after last release9to5Google · 2 September 2026 · news
  34. Gemini 3.8 Flash model statsLLM-Stats · 2 September 2026 · news
  35. Gemini 3.1 Flash-Lite benchmark resultsLayerLens · March 2026 · news
  36. DeepSeek API pricingDeepSeek · August 2026 · pricing
  37. DeepSeek-R1 model card (with distilled model evaluations)DeepSeek / Hugging Face · 20 January 2025 · docs
  38. DeepSeek-R1-0528 model cardDeepSeek / Hugging Face · 28 May 2025 · docs
  39. DeepSeek-R1-Distill-Qwen-32B model cardDeepSeek / Hugging Face · 20 January 2025 · docs
  40. DeepSeek-V4-Pro model cardDeepSeek / Hugging Face · 26 April 2026 · docs
  41. DeepSeek-V4-Flash model cardDeepSeek / Hugging Face · July 2026 · docs
  42. DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsDeepSeek / arXiv · December 2025 · paper
  43. DeepSeek raises some V4 prices by more than 10xInfoWorld · August 2026 · news
  44. Llama 4 model cardMeta · April 2025 · docs
  45. Llama 3.3 model cardMeta · 6 December 2024 · docs
  46. Llama 3.2 model cardMeta · 25 September 2024 · docs
  47. Qwen3 Technical ReportAlibaba Qwen Team / arXiv · May 2025 · paper
  48. Qwen3.8-27B model cardAlibaba / Hugging Face · August 2026 · docs
  49. Qwen3-4B-Instruct-2507 model cardAlibaba / Hugging Face · July 2025 · docs
  50. Alibaba Model Studio model pricingAlibaba Cloud · September 2026 · pricing
  51. phi-4 model cardMicrosoft / Hugging Face · December 2024 · docs
  52. Phi-4-mini-instruct model cardMicrosoft / Hugging Face · February 2025 · docs
  53. SmolLM3-3B model cardHugging Face · July 2025 · docs
  54. Mistral AI API pricingMistral AI · September 2026 · pricing
  55. Mistral models overviewMistral AI · September 2026 · docs
  56. Ministral-3-14B-Instruct-2512 model cardMistral AI / Hugging Face · 2 December 2025 · docs
  57. Mistral Small 4 model pageOpenRouter · 16 March 2026 · docs
  58. What is Amazon Nova? (teacher/student distillation matrix)Amazon Web Services · 2026 · docs
  59. What’s new in Amazon Nova 2Amazon Web Services · December 2025 · docs
  60. The Amazon Nova Family of Models: Technical Report and Model CardAmazon Science · 17 March 2025 · paper
  61. Amazon Bedrock Model DistillationAmazon Web Services · May 2025 · docs
  62. Amazon Bedrock pricingAmazon Web Services · September 2026 · pricing
  63. Nova Micro API pricingPricePerToken · 2026 · pricing
  64. Nova Lite API pricingPricePerToken · 2026 · pricing
  65. Nova 2 Lite API pricingPricePerToken · 2026 · pricing
  66. Amazon Nova Pro: AWS Bedrock model guide, specs and pricing (2026)UC Strategies · 2026 · news
  67. Together AI pricingTogether AI · September 2026 · pricing
  68. Groq supported models and pricingGroq · September 2026 · pricing
  69. xAI models and pricingxAI · August 2026 · pricing
  70. GPT-5.6 Luna (low) vs Claude 4.5 Haiku (reasoning)Artificial Analysis · September 2026 · news
  71. Gemini 3.5 Flash-Lite vs GPT-5.6 Luna (high)Artificial Analysis · September 2026 · news
  72. Claude Sonnet 5 vs Claude Opus 5Artificial Analysis · September 2026 · news
  73. DeepSeek V4 Flash vs DeepSeek V4 ProArtificial Analysis · September 2026 · news
  74. gpt-oss-120B vs Llama 4 MaverickArtificial Analysis · September 2026 · news
  75. Comparison of AI models across intelligence, performance and priceArtificial Analysis · September 2026 · news
  76. Enterprise AI costs hit 2026 low driven by price wars and Chinese open-source modelsSouth China Morning Post · August 2026 · news
  77. OpenAI discounts GPT-5.6 Luna and TerraAxios · 30 July 2026 · news
  78. The Llama 4 herd: the beginning of a new era of natively multimodal AI innovationMeta AI · 5 April 2025 · blog