{
  "perspective": "customer",
  "title": "The buyer’s view of AI distillation: which small model to actually deploy",
  "updated": "2026-09-04",
  "summary": "For an enterprise buyer in September 2026, the distillation question is no longer \"is the small model good enough\" but \"which small model, and what does the licence let me do with it\". Within a single generation, distilled and small-sibling tiers retain 88–98% of their larger sibling's benchmark score for 3–25% of the price: GPT-5.6 Luna holds about 94% of Sol's GPQA Diamond at 5% of the input rate, DeepSeek-V4-Flash scores 98% of V4-Pro's at a third of the price, and Gemini 2.5 Flash — which Google explicitly documents as a k-sparse logit distillation of 2.5 Pro — holds 96% at 24%; across a large parameter gap, retention falls to 47–76%, as DeepSeek's own R1 students and Google's Gemma E-series show. The gap that remains is not knowledge but agentic reliability: on SWE-bench Verified, Claude Sonnet 5 retains 89% of Opus 5 and Haiku 4.5 only 76%, so long-horizon coding and tool-use workloads are the last place a frontier tier still pays for itself. Open weights have become the buyer's real leverage — DeepSeek V4 (MIT), Gemma 4 (Apache 2.0), Qwen3.8-27B (Apache 2.0) and Mistral Small 4 (Apache 2.0) all clear 71.2–90.1 GPQA Diamond with no per-token fee and no anti-distillation clause — which is why the Silicon Data enterprise inference index fell from $2.04/MTok on 31 May 2026 to $1.16–1.18 in early August. The three things that should drive the decision, in order, are: whether your traffic needs frontier-grade agentic reliability on more than 15% of calls; whether the vendor's terms let you distil its outputs into your own model (OpenAI and Anthropic say no, DeepSeek's MIT licence says yes); and whether you can absorb the migration cost when a tier is repriced or retired — which, on 2026 evidence, happens roughly every quarter.",
  "stats": [
    {
      "label": "Enterprise inference price index",
      "value": 1.17,
      "unit": "USD / M tokens",
      "delta": "-43% since 31 May 2026",
      "note": "Silicon Data index cited by Jefferies; hit a 2026 low of $1.16–$1.18 on 6–8 Aug 2026, down from $2.04 on 31 May and $1.45 in late July.",
      "source": "https://www.scmp.com/tech/tech-trends/article/3363549/enterprise-ai-costs-hit-2026-low-driven-price-wars-chinese-open-source-models-research"
    },
    {
      "label": "Cheapest hosted model per 1M tokens (in)",
      "value": 0.035,
      "unit": "USD",
      "delta": "Amazon Nova Micro",
      "note": "Amazon's cheapest Nova tier. AWS supports it as a distillation student in Bedrock Model Distillation, but does not state that the shipped weights were distilled from Premier or Pro. Text-only, 128K context.",
      "source": "https://pricepertoken.com/pricing-page/model/amazon-nova-micro-v1"
    },
    {
      "label": "GPQA retained by GPT-5.6 Luna vs Sol",
      "value": 94.2,
      "unit": "%",
      "delta": "at 5% of the input price (94.2–97.3% across providers)",
      "note": "87.0 vs 92.4 GPQA Diamond; $0.20 vs $4.00 per M input tokens. Both figures are third-party (OpenRouter auto-routing); provider-specific Luna scores on the same page run to 89.9, and DataLearner reports Sol at 93.5 at max thinking, so retention is best read as a 94.2–97.3% range pending the GPT-5.6 system card.",
      "source": "https://openrouter.ai/openai/gpt-5.6-luna"
    },
    {
      "label": "GPQA of DeepSeek-V4-Flash vs V4-Pro",
      "value": 97.8,
      "unit": "%",
      "delta": "at 33% of the price",
      "note": "88.1 vs 90.1 GPQA Diamond, both MIT-licensed open weights. Same-family comparison: DeepSeek documents V4-Flash as a consolidation of V4 domain experts via on-policy distillation, not as a distillation of V4-Pro.",
      "source": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash"
    },
    {
      "label": "SWE-bench Verified retained by Claude Haiku 4.5 vs Opus 5",
      "value": 76.4,
      "unit": "%",
      "delta": "at 20% of the price",
      "note": "73.3 vs 96.0. Agentic coding is where the distilled tiers still lose most.",
      "source": "https://datanorth.ai/news/claude-opus-5-by-anthropic"
    },
    {
      "label": "Cost per 1M requests, GPT-5.6 Sol vs Luna",
      "value": 18,
      "unit": "x cheaper",
      "delta": "$3,600 → $200",
      "note": "At 400 input + 100 output tokens per request, list price, no caching.",
      "source": "https://developers.openai.com/api/docs/pricing"
    },
    {
      "label": "Best open-weight GPQA Diamond in this dataset",
      "value": 90.1,
      "unit": "%",
      "delta": "DeepSeek-V4-Pro, MIT licence (open-weight range 71.2–90.1)",
      "note": "1.6T-parameter MoE with 49B active; within 4.2 points of Gemini 3.1 Pro.",
      "source": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro"
    },
    {
      "label": "Bedrock Model Distillation claim",
      "value": 75,
      "unit": "% cheaper",
      "delta": "and up to 500% faster",
      "note": "AWS states distilled models are up to 500% faster and up to 75% less expensive than the original, with under 2% accuracy loss on RAG-style use cases.",
      "source": "https://aws.amazon.com/bedrock/model-distillation/"
    }
  ],
  "keyFindings": [
    {
      "title": "The cheap tier is now good enough for roughly 85% of enterprise traffic",
      "detail": "Within a single current generation the small sibling retains 88–98% of its teacher's knowledge benchmark. GPT-5.6 Luna scores 87.0 GPQA Diamond against Sol's 92.4; GPT-5.4 mini scores 88.0 against 93.0; Gemini 2.5 Flash scores 82.8 against 2.5 Pro's 86.4; DeepSeek-V4-Flash scores 88.1 against V4-Pro's 90.1. Across a large parameter gap retention falls to 47–76% — DeepSeek's own R1 students range from 91.2% at 70B down to 47.3% at 1.5B, and Gemma 4 31B to E4B retains 69.5%. The practical rule that falls out of the numbers is to route the hardest 5–15% of traffic to a frontier tier and everything else down to a same-generation small tier, not to the smallest model available.",
      "audience": [
        "customer",
        "developer"
      ],
      "sources": [
        "https://openrouter.ai/openai/gpt-5.6-luna",
        "https://openrouter.ai/openai/gpt-5.6-sol",
        "https://arxiv.org/html/2507.06261v1/",
        "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash"
      ]
    },
    {
      "title": "Agentic coding is the one place the frontier tier still earns its price",
      "detail": "Knowledge benchmarks compress; long-horizon agent benchmarks do not. Claude Opus 5 scores 96.0 on SWE-bench Verified, Sonnet 5 85.2 (89% retention) and Haiku 4.5 73.3 (76% retention). Llama 4 Scout retains 92% of Maverick’s MMLU-Pro but only 76% of its LiveCodeBench. If your workload is a coding agent that must finish a multi-step task unsupervised, the retention curve is much steeper than the GPQA curve suggests and a cheap tier will show up as retries, not as wrong answers.",
      "audience": [
        "customer",
        "developer"
      ],
      "sources": [
        "https://datanorth.ai/news/claude-opus-5-by-anthropic",
        "https://www.morphllm.com/claude-benchmarks",
        "https://www.anthropic.com/news/claude-haiku-4-5",
        "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md"
      ]
    },
    {
      "title": "Open weights are the buyer’s only real negotiating position",
      "detail": "Four open-weight families now clear 71.2–90.1 GPQA Diamond with no per-token fee and no anti-distillation clause: DeepSeek V4-Pro/Flash (MIT, 90.1/88.1), Qwen3.8-27B (Apache 2.0, 89.2), Gemma 4 31B (Apache 2.0, 84.3 GPQA / 85.2 MMLU-Pro) and Mistral Small 4 (Apache 2.0, 71.2 GPQA / 78.0 MMLU-Pro). Having a credible self-host fallback is what makes an API price cut stick — and 2026 has been a year of price cuts.",
      "audience": [
        "customer"
      ],
      "sources": [
        "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro",
        "https://huggingface.co/Qwen/Qwen3.8-27B",
        "https://ai.google.dev/gemma/docs/core/model_card_4",
        "https://openrouter.ai/mistralai/mistral-small-2603"
      ]
    },
    {
      "title": "You may not distil the models you buy — but you may distil the ones you download",
      "detail": "Anthropic’s Commercial Terms section D.4 bars customers from accessing the Services \"to build a competing product or service, including to train competing AI models\". OpenAI’s Services Agreement carries an equivalent restriction. Both apply to enterprise accounts. By contrast DeepSeek R1 and V4 ship under MIT, and the R1 release explicitly encourages distillation; Qwen3.8, Gemma 4, Mistral Small 4 and Ministral 3 are Apache 2.0. If your roadmap includes training a small in-house model on model outputs, the licence question decides your teacher before any benchmark does.",
      "audience": [
        "customer",
        "political"
      ],
      "sources": [
        "https://www.anthropic.com/legal/commercial-terms",
        "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "https://ai.google.dev/gemma/docs/core/model_card_4"
      ]
    },
    {
      "title": "Distillation is now a purchasable service, not just a research technique",
      "detail": "Amazon Bedrock Model Distillation takes only your prompts: it generates synthetic teacher responses and fine-tunes the student. Nova Premier, Claude 3.5 Sonnet v2 and Llama 3.3 70B are supported teachers; Amazon Nova Pro and Llama 3.2 1B/3B are supported students. AWS claims up to 500% faster and up to 75% less expensive inference than the original models, with less than 2% accuracy loss for use cases like RAG. For most buyers this is a cheaper path to a task-specific small model than running a distillation pipeline in-house.",
      "audience": [
        "customer",
        "developer"
      ],
      "sources": [
        "https://aws.amazon.com/bedrock/model-distillation/",
        "https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html"
      ]
    },
    {
      "title": "Distillation transfers narrow skills far better than broad knowledge",
      "detail": "The DeepSeek R1 student series is the cleanest natural experiment available. R1-Distill-Qwen-1.5B keeps 86% of R1’s MATH-500 (83.9 vs 97.3-class teacher performance) but only 47% of its GPQA Diamond (33.8 vs 71.5). At 7B the split is still stark: MATH-500 92.8 but GPQA 49.1. Buyers evaluating a tiny distilled model on a maths or format-following eval will systematically overestimate how it behaves on open-domain knowledge work.",
      "audience": [
        "customer",
        "academic"
      ],
      "sources": [
        "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "https://huggingface.co/deepseek-ai/DeepSeek-R1-0528"
      ]
    },
    {
      "title": "Advertised latency and measured latency have diverged because of adaptive thinking",
      "detail": "Vendor tables still say \"fastest\", but the number a buyer feels now depends on the reasoning effort setting. Artificial Analysis measures GPT-5.6 Luna at 1.70s time-to-first-token at low effort and 19.87s at high; Claude Haiku 4.5 with reasoning on measures 19.92s; Claude Sonnet 5 at max effort measures 177.77s. Gemini 2.5 Flash-Lite in non-reasoning mode measures 0.30s. Any latency SLA written against a reasoning model has to pin the effort level or it is not a specification.",
      "audience": [
        "customer",
        "developer"
      ],
      "sources": [
        "https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning",
        "https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gpt-5-6-luna-high",
        "https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5",
        "https://artificialanalysis.ai/models"
      ]
    },
    {
      "title": "Legacy tiers are the biggest silent line item in most 2026 AI bills",
      "detail": "GPT-4o still lists at $2.50/$10 — the same rate as in 2024 — while GPT-5.6 Luna lists at $0.20/$1.20 with a far larger context window. Claude Sonnet 4.6 costs 50% more than the newer, better Sonnet 5. Gemini 3.5 Flash at $1.50/$9.00 costs twice Gemini 3.8 Flash's $0.75/$3.75 — but that Flash rate is promotional through 2026-12-31 and lists at $1.50/$7.50 from 2027-01-01, i.e. the same input rate as 3.5 Flash. Nothing forces a migration, so pinned model IDs from 2024–25 quietly bill at up to 12x the current market rate for the same job.",
      "audience": [
        "customer"
      ],
      "sources": [
        "https://developers.openai.com/api/docs/pricing",
        "https://platform.claude.com/docs/en/about-claude/pricing",
        "https://ai.google.dev/gemini-api/docs/pricing"
      ]
    },
    {
      "title": "Price stability is not something you can assume any more",
      "detail": "In 2026 alone: OpenAI cut GPT-5.6 Luna 80% and Terra 20% on 30 July, then cut Sol over 20% on 21 August for a three-month window; Anthropic cancelled a scheduled Sonnet 5 increase from $2/$10 to $3/$15 and made the lower price permanent; DeepSeek raised V4 standard rates roughly 3–4.7x on 16 August — and cache-hit input rates by up to 11x — as demand strained capacity. Contract for the workload, not for the price — and keep a second vendor wired up.",
      "audience": [
        "customer",
        "financial"
      ],
      "sources": [
        "https://www.axios.com/2026/07/30/openai-cuts-prices-gpt-terra-luna5",
        "https://platform.claude.com/docs/en/about-claude/pricing",
        "https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html"
      ]
    },
    {
      "title": "Context window is no longer a reason to pay frontier prices",
      "detail": "Every GPT-5.6 tier including the $0.20 Luna carries a 1.05M-token window. Claude Sonnet 5 and Opus 5 both carry 1M. Gemini 3.5 Flash-Lite carries 1M at $0.30/MTok. Llama 4 Scout carries 10M with open weights. Two caveats matter for whole-repository or whole-contract passes: Claude Haiku 4.5 is still capped at 200K, and long prompts are repriced above 200K input tokens — Gemini 3.1 Pro and 2.5 Pro double the input rate above that threshold (and GPT-5.4-class models apply 2x input / 1.5x output), so the cheap headline rate is not the rate you pay on a million-token prompt.",
      "audience": [
        "customer"
      ],
      "sources": [
        "https://developers.openai.com/api/docs/models",
        "https://platform.claude.com/docs/en/about-claude/models/overview",
        "https://ai.google.dev/gemini-api/docs/pricing",
        "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md"
      ]
    }
  ],
  "tables": [
    {
      "id": "model-master",
      "title": "Master comparison: 74 models a buyer could shortlist in September 2026",
      "description": "Price, quality, context and licence for frontier teachers, vendor small siblings and openly-licensed distilled students. \"Quality (GPQA-D)\" is GPQA Diamond; MMLU-Pro is shown where the vendor publishes it. Blank cells mean the figure is not published — nothing here is estimated.",
      "columns": [
        {
          "key": "model",
          "label": "Model",
          "type": "text"
        },
        {
          "key": "vendor",
          "label": "Vendor",
          "type": "text"
        },
        {
          "key": "kind",
          "label": "Role",
          "type": "text"
        },
        {
          "key": "distilled",
          "label": "Distilled?",
          "type": "text"
        },
        {
          "key": "teacher",
          "label": "Teacher",
          "type": "text"
        },
        {
          "key": "in",
          "label": "Input",
          "type": "number",
          "unit": "USD/MTok"
        },
        {
          "key": "out",
          "label": "Output",
          "type": "number",
          "unit": "USD/MTok"
        },
        {
          "key": "gpqa",
          "label": "GPQA-D",
          "type": "number",
          "unit": "%"
        },
        {
          "key": "mmlupro",
          "label": "MMLU-Pro",
          "type": "number",
          "unit": "%"
        },
        {
          "key": "code",
          "label": "Coding",
          "type": "text"
        },
        {
          "key": "ctx",
          "label": "Context",
          "type": "number",
          "unit": "K tokens"
        },
        {
          "key": "license",
          "label": "Licence",
          "type": "text"
        }
      ],
      "rows": [
        {
          "model": "GPT-6 Astra",
          "vendor": "OpenAI",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 10,
          "out": 50,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 1050,
          "license": "Proprietary API",
          "_source": "https://developers.openai.com/api/docs/pricing"
        },
        {
          "model": "GPT-5.6 Sol",
          "vendor": "OpenAI",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 4,
          "out": 20,
          "gpqa": 92.4,
          "mmlupro": null,
          "code": null,
          "ctx": 1050,
          "license": "Proprietary API",
          "_source": "https://openrouter.ai/openai/gpt-5.6-sol"
        },
        {
          "model": "GPT-5.6 Terra",
          "vendor": "OpenAI",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 2,
          "out": 12,
          "gpqa": 88.4,
          "mmlupro": null,
          "code": null,
          "ctx": 1050,
          "license": "Proprietary API",
          "_source": "https://openrouter.ai/openai/gpt-5.6-terra"
        },
        {
          "model": "GPT-5.6 Luna",
          "vendor": "OpenAI",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.2,
          "out": 1.2,
          "gpqa": 87,
          "mmlupro": null,
          "code": null,
          "ctx": 1050,
          "license": "Proprietary API",
          "_source": "https://openrouter.ai/openai/gpt-5.6-luna"
        },
        {
          "model": "GPT-5.5",
          "vendor": "OpenAI",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 5,
          "out": 30,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": null,
          "license": "Proprietary API",
          "_source": "https://developers.openai.com/api/docs/pricing"
        },
        {
          "model": "GPT-5.4",
          "vendor": "OpenAI",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 2.5,
          "out": 15,
          "gpqa": 93,
          "mmlupro": null,
          "code": "57.7 (SWE-bench Pro)",
          "ctx": 1050,
          "license": "Proprietary API",
          "_source": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/"
        },
        {
          "model": "GPT-5.4 mini",
          "vendor": "OpenAI",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.75,
          "out": 4.5,
          "gpqa": 88,
          "mmlupro": null,
          "code": "54.4 (SWE-bench Pro)",
          "ctx": 400,
          "license": "Proprietary API",
          "_source": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/"
        },
        {
          "model": "GPT-5.4 nano",
          "vendor": "OpenAI",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.2,
          "out": 1.25,
          "gpqa": 82.8,
          "mmlupro": null,
          "code": "52.4 (SWE-bench Pro)",
          "ctx": 400,
          "license": "Proprietary API",
          "_source": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/"
        },
        {
          "model": "GPT-5",
          "vendor": "OpenAI",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 1.25,
          "out": 10,
          "gpqa": null,
          "mmlupro": null,
          "code": "74.9 (SWE-bench Verified)",
          "ctx": 400,
          "license": "Proprietary API",
          "_source": "https://arxiv.org/pdf/2601.03267"
        },
        {
          "model": "GPT-5 mini",
          "vendor": "OpenAI",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.25,
          "out": 2,
          "gpqa": 80.3,
          "mmlupro": null,
          "code": "45.7 (SWE-bench Pro)",
          "ctx": 400,
          "license": "Proprietary API",
          "_source": "https://openrouter.ai/openai/gpt-5-mini"
        },
        {
          "model": "GPT-5 nano",
          "vendor": "OpenAI",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.05,
          "out": 0.4,
          "gpqa": 70.9,
          "mmlupro": null,
          "code": null,
          "ctx": 400,
          "license": "Proprietary API",
          "_source": "https://openrouter.ai/openai/gpt-5-nano"
        },
        {
          "model": "GPT-4o",
          "vendor": "OpenAI",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 2.5,
          "out": 10,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 128,
          "license": "Proprietary API",
          "_source": "https://openrouter.ai/openai/gpt-4o"
        },
        {
          "model": "GPT-4o mini",
          "vendor": "OpenAI",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.15,
          "out": 0.6,
          "gpqa": null,
          "mmlupro": null,
          "code": "87.2 (HumanEval)",
          "ctx": 128,
          "license": "Proprietary API",
          "_source": "https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/"
        },
        {
          "model": "o3",
          "vendor": "OpenAI",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 2,
          "out": 8,
          "gpqa": 83.3,
          "mmlupro": null,
          "code": "69.1 (SWE-bench Verified)",
          "ctx": 200,
          "license": "Proprietary API",
          "_source": "https://www.datacamp.com/blog/o4-mini"
        },
        {
          "model": "o4-mini",
          "vendor": "OpenAI",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 1.1,
          "out": 4.4,
          "gpqa": 81.4,
          "mmlupro": null,
          "code": "68.1 (SWE-bench Verified)",
          "ctx": 200,
          "license": "Proprietary API",
          "_source": "https://www.datacamp.com/blog/o4-mini"
        },
        {
          "model": "gpt-oss-120b",
          "vendor": "OpenAI (open weights)",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.15,
          "out": 0.6,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 128,
          "license": "Apache 2.0",
          "_source": "https://huggingface.co/openai/gpt-oss-20b"
        },
        {
          "model": "gpt-oss-20b",
          "vendor": "OpenAI (open weights)",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.075,
          "out": 0.3,
          "gpqa": 58.59,
          "mmlupro": null,
          "code": "53.2 (SWE-bench Verified)",
          "ctx": 128,
          "license": "Apache 2.0",
          "_source": "https://huggingface.co/openai/gpt-oss-20b"
        },
        {
          "model": "Claude Fable 5.1",
          "vendor": "Anthropic",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 10,
          "out": 50,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://platform.claude.com/docs/en/about-claude/models/overview"
        },
        {
          "model": "Claude Opus 5",
          "vendor": "Anthropic",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 5,
          "out": 25,
          "gpqa": null,
          "mmlupro": null,
          "code": "96 (SWE-bench Verified)",
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://datanorth.ai/news/claude-opus-5-by-anthropic"
        },
        {
          "model": "Claude Sonnet 5",
          "vendor": "Anthropic",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 2,
          "out": 10,
          "gpqa": null,
          "mmlupro": null,
          "code": "85.2 (SWE-bench Verified)",
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://www.morphllm.com/claude-benchmarks"
        },
        {
          "model": "Claude Sonnet 4.6",
          "vendor": "Anthropic",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 3,
          "out": 15,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://platform.claude.com/docs/en/about-claude/models/overview"
        },
        {
          "model": "Claude Haiku 4.5",
          "vendor": "Anthropic",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 1,
          "out": 5,
          "gpqa": null,
          "mmlupro": null,
          "code": "73.3 (SWE-bench Verified)",
          "ctx": 200,
          "license": "Proprietary API",
          "_source": "https://www.anthropic.com/news/claude-haiku-4-5"
        },
        {
          "model": "Claude Haiku 3.5",
          "vendor": "Anthropic",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.8,
          "out": 4,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 200,
          "license": "Proprietary API",
          "_source": "https://platform.claude.com/docs/en/about-claude/pricing"
        },
        {
          "model": "Gemini 3.1 Pro (Preview)",
          "vendor": "Google",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 2,
          "out": 12,
          "gpqa": 94.3,
          "mmlupro": 92.6,
          "code": "80.6 (SWE-bench Verified)",
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://deepmind.google/models/gemini/pro/"
        },
        {
          "model": "Gemini 3.8 Flash",
          "vendor": "Google",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.75,
          "out": 3.75,
          "gpqa": null,
          "mmlupro": null,
          "code": "73.7 (DeepSWE v1.1)",
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://deepmind.google/models/gemini/flash/"
        },
        {
          "model": "Gemini 3.5 Flash",
          "vendor": "Google",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 1.5,
          "out": 9,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://ai.google.dev/gemini-api/docs/models"
        },
        {
          "model": "Gemini 3.5 Flash-Lite",
          "vendor": "Google",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.3,
          "out": 2.5,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://ai.google.dev/gemini-api/docs/models"
        },
        {
          "model": "Gemini 3.1 Flash-Lite",
          "vendor": "Google",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.25,
          "out": 1.5,
          "gpqa": 72.2,
          "mmlupro": 83,
          "code": null,
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://layerlens.ai/blog/gemini-3-1-flash-lite-benchmark-results-efficiency-model-comparison"
        },
        {
          "model": "Gemini 2.5 Pro",
          "vendor": "Google",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 1.25,
          "out": 10,
          "gpqa": 86.4,
          "mmlupro": null,
          "code": "67.2 (SWE-bench Verified)",
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://arxiv.org/html/2507.06261v1/"
        },
        {
          "model": "Gemini 2.5 Flash",
          "vendor": "Google",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "Gemini 2.5 Pro (k-sparse logit distillation)",
          "in": 0.3,
          "out": 2.5,
          "gpqa": 82.8,
          "mmlupro": null,
          "code": "60.3 (SWE-bench Verified)",
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://arxiv.org/html/2507.06261v1/"
        },
        {
          "model": "Gemini 2.5 Flash-Lite",
          "vendor": "Google",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "Gemini 2.5 Pro (k-sparse logit distillation)",
          "in": 0.1,
          "out": 0.4,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://arxiv.org/html/2507.06261v1/"
        },
        {
          "model": "Gemma 4 31B",
          "vendor": "Google",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": null,
          "out": null,
          "gpqa": 84.3,
          "mmlupro": 85.2,
          "code": "80 (LiveCodeBench v6)",
          "ctx": 256,
          "license": "Apache 2.0",
          "_source": "https://ai.google.dev/gemma/docs/core/model_card_4"
        },
        {
          "model": "Gemma 4 26B A4B (MoE)",
          "vendor": "Google",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": null,
          "out": null,
          "gpqa": 82.3,
          "mmlupro": 82.6,
          "code": "77.1 (LiveCodeBench v6)",
          "ctx": 256,
          "license": "Apache 2.0",
          "_source": "https://ai.google.dev/gemma/docs/core/model_card_4"
        },
        {
          "model": "Gemma 4 12B Unified",
          "vendor": "Google",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": null,
          "out": null,
          "gpqa": 78.8,
          "mmlupro": 77.2,
          "code": "72 (LiveCodeBench v6)",
          "ctx": 256,
          "license": "Apache 2.0",
          "_source": "https://ai.google.dev/gemma/docs/core/model_card_4"
        },
        {
          "model": "Gemma 4 E4B",
          "vendor": "Google",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": null,
          "out": null,
          "gpqa": 58.6,
          "mmlupro": 69.4,
          "code": "52 (LiveCodeBench v6)",
          "ctx": 128,
          "license": "Apache 2.0",
          "_source": "https://ai.google.dev/gemma/docs/core/model_card_4"
        },
        {
          "model": "Gemma 4 E2B",
          "vendor": "Google",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": null,
          "out": null,
          "gpqa": 43.4,
          "mmlupro": 60,
          "code": "44 (LiveCodeBench v6)",
          "ctx": 128,
          "license": "Apache 2.0",
          "_source": "https://ai.google.dev/gemma/docs/core/model_card_4"
        },
        {
          "model": "Gemma 3 27B IT",
          "vendor": "Google",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": null,
          "out": null,
          "gpqa": 24.3,
          "mmlupro": null,
          "code": "48.8 (HumanEval)",
          "ctx": 128,
          "license": "Gemma Terms of Use (custom)",
          "_source": "https://huggingface.co/google/gemma-3-27b-it"
        },
        {
          "model": "Gemma 3 4B IT",
          "vendor": "Google",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": null,
          "out": null,
          "gpqa": 15,
          "mmlupro": null,
          "code": "36 (HumanEval)",
          "ctx": 128,
          "license": "Gemma Terms of Use (custom)",
          "_source": "https://huggingface.co/google/gemma-3-27b-it"
        },
        {
          "model": "Llama 4 Maverick",
          "vendor": "Meta",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "Llama 4 Behemoth (codistillation)",
          "in": null,
          "out": null,
          "gpqa": 69.8,
          "mmlupro": 80.5,
          "code": "43.4 (LiveCodeBench)",
          "ctx": 1000,
          "license": "Llama 4 Community License (custom commercial)",
          "_source": "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md"
        },
        {
          "model": "Llama 4 Scout",
          "vendor": "Meta",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "Llama 4 Behemoth (codistillation)",
          "in": null,
          "out": null,
          "gpqa": 57.2,
          "mmlupro": 74.3,
          "code": "32.8 (LiveCodeBench)",
          "ctx": 10000,
          "license": "Llama 4 Community License (custom commercial)",
          "_source": "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md"
        },
        {
          "model": "Llama 3.3 70B Instruct",
          "vendor": "Meta",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 1.04,
          "out": 1.04,
          "gpqa": 50.5,
          "mmlupro": 68.9,
          "code": "88.4 (HumanEval)",
          "ctx": 128,
          "license": "Llama 3.3 Community License (custom commercial)",
          "_source": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md"
        },
        {
          "model": "Llama 3.2 3B Instruct",
          "vendor": "Meta",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "Llama 3.1 8B and 70B (token-level logit distillation after pruning)",
          "in": null,
          "out": null,
          "gpqa": 32.8,
          "mmlupro": null,
          "code": null,
          "ctx": 128,
          "license": "Llama 3.2 Community License (custom commercial)",
          "_source": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md"
        },
        {
          "model": "Llama 3.2 1B Instruct",
          "vendor": "Meta",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "Llama 3.1 8B and 70B (token-level logit distillation after pruning)",
          "in": null,
          "out": null,
          "gpqa": 27.2,
          "mmlupro": null,
          "code": null,
          "ctx": 128,
          "license": "Llama 3.2 Community License (custom commercial)",
          "_source": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md"
        },
        {
          "model": "DeepSeek-V4-Pro",
          "vendor": "DeepSeek",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 1.32,
          "out": 3.96,
          "gpqa": 90.1,
          "mmlupro": 87.5,
          "code": "80.6 (SWE-bench Verified)",
          "ctx": 1000,
          "license": "MIT",
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro"
        },
        {
          "model": "DeepSeek-V4-Flash",
          "vendor": "DeepSeek",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "DeepSeek V4 domain experts (on-policy distillation consolidation)",
          "in": 0.44,
          "out": 1.32,
          "gpqa": 88.1,
          "mmlupro": 86.2,
          "code": "79 (SWE-bench Verified)",
          "ctx": 1000,
          "license": "MIT",
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash"
        },
        {
          "model": "DeepSeek-V3.2",
          "vendor": "DeepSeek",
          "kind": "Open weights",
          "distilled": "Yes (documented)",
          "teacher": "DeepSeek specialist models (specialist distillation into the generalist)",
          "in": null,
          "out": null,
          "gpqa": 82.4,
          "mmlupro": 85,
          "code": "73.1 (SWE-bench Verified)",
          "ctx": 128,
          "license": "MIT",
          "_source": "https://arxiv.org/html/2512.02556"
        },
        {
          "model": "DeepSeek-R1 (0528)",
          "vendor": "DeepSeek",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": null,
          "out": null,
          "gpqa": 81,
          "mmlupro": 85,
          "code": "73.3 (LiveCodeBench)",
          "ctx": 128,
          "license": "MIT",
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1-0528"
        },
        {
          "model": "DeepSeek-R1-Distill-Llama-70B",
          "vendor": "DeepSeek / Meta base",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "DeepSeek-R1",
          "in": null,
          "out": null,
          "gpqa": 65.2,
          "mmlupro": null,
          "code": "57.5 (LiveCodeBench)",
          "ctx": 128,
          "license": "MIT (weights) over Llama 3.3 Community License base",
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1"
        },
        {
          "model": "DeepSeek-R1-Distill-Qwen-32B",
          "vendor": "DeepSeek / Qwen base",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "DeepSeek-R1",
          "in": null,
          "out": null,
          "gpqa": 62.1,
          "mmlupro": null,
          "code": "57.2 (LiveCodeBench)",
          "ctx": 128,
          "license": "MIT (weights), Qwen2.5-32B base under Apache 2.0",
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B"
        },
        {
          "model": "DeepSeek-R1-Distill-Qwen-14B",
          "vendor": "DeepSeek / Qwen base",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "DeepSeek-R1",
          "in": null,
          "out": null,
          "gpqa": 59.1,
          "mmlupro": null,
          "code": "53.1 (LiveCodeBench)",
          "ctx": 128,
          "license": "MIT (weights), Qwen2.5-14B base under Apache 2.0",
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1"
        },
        {
          "model": "DeepSeek-R1-Distill-Llama-8B",
          "vendor": "DeepSeek / Meta base",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "DeepSeek-R1",
          "in": null,
          "out": null,
          "gpqa": 49,
          "mmlupro": null,
          "code": "39.6 (LiveCodeBench)",
          "ctx": 128,
          "license": "MIT (weights) over Llama 3.1 Community License base",
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1"
        },
        {
          "model": "DeepSeek-R1-Distill-Qwen-7B",
          "vendor": "DeepSeek / Qwen base",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "DeepSeek-R1",
          "in": null,
          "out": null,
          "gpqa": 49.1,
          "mmlupro": null,
          "code": "37.6 (LiveCodeBench)",
          "ctx": 128,
          "license": "MIT (weights), Qwen2.5-Math-7B base under Apache 2.0",
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1"
        },
        {
          "model": "DeepSeek-R1-Distill-Qwen-1.5B",
          "vendor": "DeepSeek / Qwen base",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "DeepSeek-R1",
          "in": null,
          "out": null,
          "gpqa": 33.8,
          "mmlupro": null,
          "code": "16.9 (LiveCodeBench)",
          "ctx": 128,
          "license": "MIT (weights), Qwen2.5-Math-1.5B base under Apache 2.0",
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1"
        },
        {
          "model": "Qwen3.8-27B",
          "vendor": "Alibaba",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.5,
          "out": 3,
          "gpqa": 89.2,
          "mmlupro": null,
          "code": "61.7 (SWE-bench Pro)",
          "ctx": 262,
          "license": "Apache 2.0",
          "_source": "https://huggingface.co/Qwen/Qwen3.8-27B"
        },
        {
          "model": "qwen3.8-max",
          "vendor": "Alibaba",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 2,
          "out": 6,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": null,
          "license": "Proprietary API",
          "_source": "https://www.alibabacloud.com/help/en/model-studio/model-pricing"
        },
        {
          "model": "qwen3.8-flash",
          "vendor": "Alibaba",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.15,
          "out": 0.47,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": null,
          "license": "Proprietary API",
          "_source": "https://www.alibabacloud.com/help/en/model-studio/model-pricing"
        },
        {
          "model": "qwen-turbo",
          "vendor": "Alibaba",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.05,
          "out": 0.2,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": null,
          "license": "Proprietary API",
          "_source": "https://www.alibabacloud.com/help/en/model-studio/model-pricing"
        },
        {
          "model": "Qwen3-4B-Instruct-2507",
          "vendor": "Alibaba",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "Qwen3-32B / Qwen3-235B-A22B (off-policy + on-policy strong-to-weak distillation)",
          "in": null,
          "out": null,
          "gpqa": 62,
          "mmlupro": 69.6,
          "code": "35.1 (LiveCodeBench v6)",
          "ctx": 262,
          "license": "Apache 2.0",
          "_source": "https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507"
        },
        {
          "model": "qwen3-8b (hosted)",
          "vendor": "Alibaba",
          "kind": "Documented distillation",
          "distilled": "Yes (documented)",
          "teacher": "Qwen3-32B / Qwen3-235B-A22B (strong-to-weak distillation)",
          "in": 0.18,
          "out": 0.7,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 128,
          "license": "Apache 2.0",
          "_source": "https://arxiv.org/pdf/2505.09388"
        },
        {
          "model": "Phi-4 (14B)",
          "vendor": "Microsoft",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": null,
          "out": null,
          "gpqa": 56.1,
          "mmlupro": 70.4,
          "code": "82.6 (HumanEval)",
          "ctx": 16,
          "license": "MIT",
          "_source": "https://huggingface.co/microsoft/phi-4"
        },
        {
          "model": "Phi-4-mini-instruct (3.8B)",
          "vendor": "Microsoft",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": null,
          "out": null,
          "gpqa": 25.2,
          "mmlupro": 52.8,
          "code": null,
          "ctx": 128,
          "license": "MIT",
          "_source": "https://huggingface.co/microsoft/Phi-4-mini-instruct"
        },
        {
          "model": "Mistral Medium 3.5",
          "vendor": "Mistral AI",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 1.5,
          "out": 7.5,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": null,
          "license": "Modified MIT",
          "_source": "https://docs.mistral.ai/getting-started/models/models_overview/"
        },
        {
          "model": "Mistral Large 3",
          "vendor": "Mistral AI",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.5,
          "out": 1.5,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": null,
          "license": "Apache 2.0",
          "_source": "https://docs.mistral.ai/getting-started/models/models_overview/"
        },
        {
          "model": "Mistral Small 4",
          "vendor": "Mistral AI",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.15,
          "out": 0.6,
          "gpqa": 71.2,
          "mmlupro": 78,
          "code": null,
          "ctx": null,
          "license": "Apache 2.0",
          "_source": "https://openrouter.ai/mistralai/mistral-small-2603"
        },
        {
          "model": "Ministral 3 14B Instruct",
          "vendor": "Mistral AI",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.2,
          "out": 0.2,
          "gpqa": 71.2,
          "mmlupro": null,
          "code": "64.6 (LiveCodeBench)",
          "ctx": 256,
          "license": "Apache 2.0",
          "_source": "https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512"
        },
        {
          "model": "Ministral 3 8B",
          "vendor": "Mistral AI",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.15,
          "out": 0.15,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 256,
          "license": "Apache 2.0",
          "_source": "https://docs.mistral.ai/getting-started/models/models_overview/"
        },
        {
          "model": "Ministral 3 3B",
          "vendor": "Mistral AI",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.1,
          "out": 0.1,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 256,
          "license": "Apache 2.0",
          "_source": "https://docs.mistral.ai/getting-started/models/models_overview/"
        },
        {
          "model": "Amazon Nova Premier",
          "vendor": "Amazon",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": null,
          "out": null,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html"
        },
        {
          "model": "Amazon Nova Pro",
          "vendor": "Amazon",
          "kind": "Frontier teacher",
          "distilled": "Undisclosed (Bedrock distillation student)",
          "teacher": "undisclosed",
          "in": 0.8,
          "out": 3.2,
          "gpqa": 46.9,
          "mmlupro": null,
          "code": null,
          "ctx": 300,
          "license": "Proprietary API",
          "_source": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf"
        },
        {
          "model": "Amazon Nova Lite",
          "vendor": "Amazon",
          "kind": "Small tier",
          "distilled": "Undisclosed (Bedrock distillation student)",
          "teacher": "undisclosed",
          "in": 0.06,
          "out": 0.24,
          "gpqa": 42,
          "mmlupro": null,
          "code": null,
          "ctx": 300,
          "license": "Proprietary API",
          "_source": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf"
        },
        {
          "model": "Amazon Nova Micro",
          "vendor": "Amazon",
          "kind": "Small tier",
          "distilled": "Undisclosed (Bedrock distillation student)",
          "teacher": "undisclosed",
          "in": 0.035,
          "out": 0.14,
          "gpqa": 40,
          "mmlupro": null,
          "code": null,
          "ctx": 128,
          "license": "Proprietary API",
          "_source": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf"
        },
        {
          "model": "Amazon Nova 2 Lite",
          "vendor": "Amazon",
          "kind": "Small sibling",
          "distilled": "Undisclosed",
          "teacher": "undisclosed",
          "in": 0.3,
          "out": 2.5,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 1000,
          "license": "Proprietary API",
          "_source": "https://docs.aws.amazon.com/nova/latest/nova2-userguide/whats-new.html"
        },
        {
          "model": "SmolLM3-3B",
          "vendor": "Hugging Face",
          "kind": "Open weights",
          "distilled": "Undisclosed",
          "teacher": "n/a",
          "in": null,
          "out": null,
          "gpqa": 41.7,
          "mmlupro": null,
          "code": "30.48 (HumanEval+)",
          "ctx": 128,
          "license": "Apache 2.0",
          "_source": "https://huggingface.co/HuggingFaceTB/SmolLM3-3B"
        },
        {
          "model": "grok-4.6",
          "vendor": "xAI",
          "kind": "Frontier teacher",
          "distilled": "No",
          "teacher": "n/a",
          "in": 2,
          "out": 6,
          "gpqa": null,
          "mmlupro": null,
          "code": null,
          "ctx": 200,
          "license": "Proprietary API",
          "_source": "https://docs.x.ai/docs/models"
        }
      ],
      "notes": "Prices are vendor list rates in USD per million tokens, before batch (typically -50%) or cache discounts. Open-weight rows with no price are self-host only in this dataset. DeepSeek prices are peak-hour; off-peak is half. Gemini 3.1 Pro and 2.5 Pro input prices double above 200K input tokens. Gemini 3.8 Flash's $0.75/$3.75 is promotional through 2026-12-31; it lists at $1.50/$7.50 from 2027-01-01. Qwen3.8-27B is shown at Alibaba Model Studio's list rate; Groq hosts the same weights at $0.80/$4.00. Gemini 3.7 Flash shipped in August 2026 but is not included in this shortlist.",
      "sources": [
        "https://developers.openai.com/api/docs/pricing",
        "https://platform.claude.com/docs/en/about-claude/pricing",
        "https://ai.google.dev/gemini-api/docs/pricing",
        "https://api-docs.deepseek.com/quick_start/pricing/",
        "https://mistral.ai/pricing/api",
        "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "https://www.together.ai/pricing",
        "https://console.groq.com/docs/models",
        "https://ai.google.dev/gemma/docs/core/model_card_4",
        "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md",
        "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf"
      ]
    },
    {
      "id": "price-vs-quality",
      "title": "Price vs quality: GPQA Diamond points per dollar of blended token price",
      "description": "Blended price uses the 3:1 input:output weighting (0.75 x input + 0.25 x output). \"GPQA per $\" is the crude but decisive ranking metric for knowledge-heavy workloads. Cost per 1M requests assumes 400 input + 100 output tokens per request at list price with no caching.",
      "columns": [
        {
          "key": "model",
          "label": "Model",
          "type": "text"
        },
        {
          "key": "vendor",
          "label": "Vendor",
          "type": "text"
        },
        {
          "key": "tier",
          "label": "Tier",
          "type": "text"
        },
        {
          "key": "gpqa",
          "label": "GPQA Diamond",
          "type": "number",
          "unit": "%"
        },
        {
          "key": "blended",
          "label": "Blended price",
          "type": "number",
          "unit": "USD/MTok"
        },
        {
          "key": "ppp",
          "label": "GPQA points per $",
          "type": "number",
          "unit": "pts/USD"
        },
        {
          "key": "per1m",
          "label": "Cost / 1M requests",
          "type": "number",
          "unit": "USD"
        }
      ],
      "rows": [
        {
          "model": "Amazon Nova Micro",
          "vendor": "Amazon",
          "tier": "Small tier",
          "gpqa": 40,
          "blended": 0.0613,
          "ppp": 652.5,
          "per1m": 28,
          "_source": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf"
        },
        {
          "model": "GPT-5 nano",
          "vendor": "OpenAI",
          "tier": "Small sibling",
          "gpqa": 70.9,
          "blended": 0.1375,
          "ppp": 515.6,
          "per1m": 60,
          "_source": "https://openrouter.ai/openai/gpt-5-nano"
        },
        {
          "model": "gpt-oss-20b",
          "vendor": "OpenAI (open weights)",
          "tier": "Open weights",
          "gpqa": 58.59,
          "blended": 0.1312,
          "ppp": 446.6,
          "per1m": 60,
          "_source": "https://huggingface.co/openai/gpt-oss-20b"
        },
        {
          "model": "Amazon Nova Lite",
          "vendor": "Amazon",
          "tier": "Small tier",
          "gpqa": 42,
          "blended": 0.105,
          "ppp": 400,
          "per1m": 48,
          "_source": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf"
        },
        {
          "model": "Ministral 3 14B Instruct",
          "vendor": "Mistral AI",
          "tier": "Open weights",
          "gpqa": 71.2,
          "blended": 0.2,
          "ppp": 356,
          "per1m": 100,
          "_source": "https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512"
        },
        {
          "model": "Mistral Small 4",
          "vendor": "Mistral AI",
          "tier": "Small sibling",
          "gpqa": 71.2,
          "blended": 0.2625,
          "ppp": 271.2,
          "per1m": 120,
          "_source": "https://openrouter.ai/mistralai/mistral-small-2603"
        },
        {
          "model": "GPT-5.6 Luna",
          "vendor": "OpenAI",
          "tier": "Small sibling",
          "gpqa": 87,
          "blended": 0.45,
          "ppp": 193.3,
          "per1m": 200,
          "_source": "https://openrouter.ai/openai/gpt-5.6-luna"
        },
        {
          "model": "GPT-5.4 nano",
          "vendor": "OpenAI",
          "tier": "Small sibling",
          "gpqa": 82.8,
          "blended": 0.4625,
          "ppp": 179,
          "per1m": 205,
          "_source": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/"
        },
        {
          "model": "Gemini 3.1 Flash-Lite",
          "vendor": "Google",
          "tier": "Small sibling",
          "gpqa": 72.2,
          "blended": 0.5625,
          "ppp": 128.4,
          "per1m": 250,
          "_source": "https://layerlens.ai/blog/gemini-3-1-flash-lite-benchmark-results-efficiency-model-comparison"
        },
        {
          "model": "DeepSeek-V4-Flash",
          "vendor": "DeepSeek",
          "tier": "Distilled",
          "gpqa": 88.1,
          "blended": 0.66,
          "ppp": 133.5,
          "per1m": 308,
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash"
        },
        {
          "model": "GPT-5 mini",
          "vendor": "OpenAI",
          "tier": "Small sibling",
          "gpqa": 80.3,
          "blended": 0.6875,
          "ppp": 116.8,
          "per1m": 300,
          "_source": "https://openrouter.ai/openai/gpt-5-mini"
        },
        {
          "model": "Gemini 2.5 Flash",
          "vendor": "Google",
          "tier": "Distilled",
          "gpqa": 82.8,
          "blended": 0.85,
          "ppp": 97.4,
          "per1m": 370,
          "_source": "https://arxiv.org/html/2507.06261v1/"
        },
        {
          "model": "Qwen3.8-27B",
          "vendor": "Alibaba",
          "tier": "Open weights",
          "gpqa": 89.2,
          "blended": 1.125,
          "ppp": 79.3,
          "per1m": 500,
          "_source": "https://huggingface.co/Qwen/Qwen3.8-27B"
        },
        {
          "model": "GPT-5.4 mini",
          "vendor": "OpenAI",
          "tier": "Small sibling",
          "gpqa": 88,
          "blended": 1.6875,
          "ppp": 52.1,
          "per1m": 750,
          "_source": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/"
        },
        {
          "model": "Llama 3.3 70B Instruct",
          "vendor": "Meta",
          "tier": "Open weights",
          "gpqa": 50.5,
          "blended": 1.04,
          "ppp": 48.6,
          "per1m": 520,
          "_source": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md"
        },
        {
          "model": "DeepSeek-V4-Pro",
          "vendor": "DeepSeek",
          "tier": "Teacher",
          "gpqa": 90.1,
          "blended": 1.98,
          "ppp": 45.5,
          "per1m": 924,
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro"
        },
        {
          "model": "o4-mini",
          "vendor": "OpenAI",
          "tier": "Small sibling",
          "gpqa": 81.4,
          "blended": 1.925,
          "ppp": 42.3,
          "per1m": 880,
          "_source": "https://openrouter.ai/openai/o4-mini"
        },
        {
          "model": "Amazon Nova Pro",
          "vendor": "Amazon",
          "tier": "Teacher",
          "gpqa": 46.9,
          "blended": 1.4,
          "ppp": 33.5,
          "per1m": 640,
          "_source": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf"
        },
        {
          "model": "o3",
          "vendor": "OpenAI",
          "tier": "Teacher",
          "gpqa": 83.3,
          "blended": 3.5,
          "ppp": 23.8,
          "per1m": 1600,
          "_source": "https://www.datacamp.com/blog/o4-mini"
        },
        {
          "model": "Gemini 2.5 Pro",
          "vendor": "Google",
          "tier": "Teacher",
          "gpqa": 86.4,
          "blended": 3.4375,
          "ppp": 25.1,
          "per1m": 1500,
          "_source": "https://arxiv.org/html/2507.06261v1/"
        },
        {
          "model": "Gemini 3.1 Pro (Preview)",
          "vendor": "Google",
          "tier": "Teacher",
          "gpqa": 94.3,
          "blended": 4.5,
          "ppp": 21,
          "per1m": 2000,
          "_source": "https://deepmind.google/models/gemini/pro/"
        },
        {
          "model": "GPT-5.6 Terra",
          "vendor": "OpenAI",
          "tier": "Small sibling",
          "gpqa": 88.4,
          "blended": 4.5,
          "ppp": 19.6,
          "per1m": 2000,
          "_source": "https://openrouter.ai/openai/gpt-5.6-terra"
        },
        {
          "model": "GPT-5.4",
          "vendor": "OpenAI",
          "tier": "Teacher",
          "gpqa": 93,
          "blended": 5.625,
          "ppp": 16.5,
          "per1m": 2500,
          "_source": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/"
        },
        {
          "model": "GPT-5.6 Sol",
          "vendor": "OpenAI",
          "tier": "Teacher",
          "gpqa": 92.4,
          "blended": 8,
          "ppp": 11.6,
          "per1m": 3600,
          "_source": "https://openrouter.ai/openai/gpt-5.6-sol"
        }
      ],
      "notes": "GPT-5 nano tops this ranking on raw efficiency but scores only 70.9 GPQA; GPT-5.6 Luna is the highest-scoring model in the top three, which is why it dominates most 2026 routing configurations. Amazon Nova figures use MMLU-era GPQA methodology from the Nova technical report and are not directly comparable with the 2026 reasoning-model scores.",
      "sources": [
        "https://developers.openai.com/api/docs/pricing",
        "https://openrouter.ai/openai/gpt-5.6-luna",
        "https://api-docs.deepseek.com/quick_start/pricing/",
        "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
        "https://ai.google.dev/gemini-api/docs/pricing",
        "https://arxiv.org/html/2507.06261v1/",
        "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf",
        "https://console.groq.com/docs/models"
      ]
    },
    {
      "id": "latency",
      "title": "Measured latency and throughput (Artificial Analysis, September 2026)",
      "description": "Independently measured time-to-first-token and output speed. On adaptive-thinking models TTFT includes reasoning time, so the effort setting is stated for every row — without it these numbers are not comparable.",
      "columns": [
        {
          "key": "model",
          "label": "Model (setting)",
          "type": "text"
        },
        {
          "key": "ttft",
          "label": "Time to first token",
          "type": "number",
          "unit": "s"
        },
        {
          "key": "tps",
          "label": "Output speed",
          "type": "number",
          "unit": "tokens/s"
        },
        {
          "key": "aaii",
          "label": "AA Intelligence Index",
          "type": "number"
        },
        {
          "key": "blended",
          "label": "Blended price",
          "type": "number",
          "unit": "USD/MTok"
        }
      ],
      "rows": [
        {
          "model": "Gemini 2.5 Flash-Lite (non-reasoning)",
          "ttft": 0.3,
          "tps": null,
          "aaii": null,
          "blended": 0.18,
          "_source": "https://artificialanalysis.ai/models"
        },
        {
          "model": "gpt-oss-120b (high)",
          "ttft": 0.85,
          "tps": 151,
          "aaii": 24,
          "blended": 0.2,
          "_source": "https://artificialanalysis.ai/models/comparisons/gpt-oss-120b-vs-llama-4-maverick"
        },
        {
          "model": "Llama 4 Maverick",
          "ttft": 0.92,
          "tps": 82,
          "aaii": 14,
          "blended": 0.31,
          "_source": "https://artificialanalysis.ai/models/comparisons/gpt-oss-120b-vs-llama-4-maverick"
        },
        {
          "model": "DeepSeek V4-Flash 0731 (reasoning, max)",
          "ttft": 1.19,
          "tps": 140,
          "aaii": 52,
          "blended": 0.23,
          "_source": "https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-deepseek-v4-pro"
        },
        {
          "model": "GPT-5.6 Luna (low)",
          "ttft": 1.7,
          "tps": 109,
          "aaii": 34,
          "blended": 0.17,
          "_source": "https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning"
        },
        {
          "model": "DeepSeek V4-Pro 0813 (reasoning, max)",
          "ttft": 1.9,
          "tps": 60.2,
          "aaii": 53,
          "blended": 0.69,
          "_source": "https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-deepseek-v4-pro"
        },
        {
          "model": "Gemini 3.5 Flash-Lite",
          "ttft": 6.48,
          "tps": 391,
          "aaii": 37,
          "blended": 0.33,
          "_source": "https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gpt-5-6-luna-high"
        },
        {
          "model": "GPT-5.6 Luna (high)",
          "ttft": 19.87,
          "tps": 123,
          "aaii": 47,
          "blended": 0.17,
          "_source": "https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gpt-5-6-luna-high"
        },
        {
          "model": "Claude Haiku 4.5 (reasoning)",
          "ttft": 19.92,
          "tps": 90,
          "aaii": 30,
          "blended": 0.77,
          "_source": "https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning"
        },
        {
          "model": "Claude Opus 5 (adaptive, max effort)",
          "ttft": 77.25,
          "tps": 57,
          "aaii": 63,
          "blended": 3.85,
          "_source": "https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5"
        },
        {
          "model": "Claude Sonnet 5 (adaptive, max effort)",
          "ttft": 177.77,
          "tps": 78,
          "aaii": 55,
          "blended": 1.54,
          "_source": "https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5"
        },
        {
          "model": "Amazon Nova 2 Lite (non-reasoning)",
          "ttft": null,
          "tps": 149,
          "aaii": 12,
          "blended": 0.85,
          "_source": "https://pricepertoken.com/pricing-page/model/amazon-nova-2-lite-v1"
        }
      ],
      "notes": "Gemini 3.5 Flash-Lite is the throughput leader at 391 tokens/s. GPT-5.6 Luna is the only model here that spans both ends of the latency range purely through its effort parameter (1.70s to 19.87s), which makes it unusually easy to run interactive and batch traffic on one model ID. Claude Sonnet 5’s 177.77s figure is max-effort adaptive thinking, not a typical production setting.",
      "sources": [
        "https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning",
        "https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gpt-5-6-luna-high",
        "https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5",
        "https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-deepseek-v4-pro",
        "https://artificialanalysis.ai/models/comparisons/gpt-oss-120b-vs-llama-4-maverick",
        "https://artificialanalysis.ai/models"
      ]
    },
    {
      "id": "license-deployment",
      "title": "Licence and deployment: what you are allowed to do with each family",
      "description": "The question that decides procurement before any benchmark does. \"Distil outputs?\" means: may you legally train your own model on this model’s outputs.",
      "columns": [
        {
          "key": "family",
          "label": "Family",
          "type": "text"
        },
        {
          "key": "license",
          "label": "Licence",
          "type": "text"
        },
        {
          "key": "weights",
          "label": "Weights",
          "type": "text"
        },
        {
          "key": "selfhost",
          "label": "Self-host / air-gap",
          "type": "text"
        },
        {
          "key": "distil",
          "label": "Distil outputs?",
          "type": "text"
        },
        {
          "key": "residency",
          "label": "Data residency options",
          "type": "text"
        },
        {
          "key": "notes",
          "label": "Buyer note",
          "type": "text"
        }
      ],
      "rows": [
        {
          "family": "OpenAI GPT-5.x / GPT-6 / o-series",
          "license": "Proprietary API",
          "weights": "Closed",
          "selfhost": "No",
          "distil": "No — Services Agreement bars using Output to develop competing models",
          "residency": "Azure/Foundry regions",
          "notes": "Deepest price ladder in the market and 1.05M context on every 5.6 tier.",
          "_source": "https://openai.com/policies/services-agreement/"
        },
        {
          "family": "OpenAI gpt-oss 20b / 120b",
          "license": "Apache 2.0",
          "weights": "Open",
          "selfhost": "Yes (120b on one 80GB GPU; 20b in 16GB)",
          "distil": "Yes",
          "residency": "Anywhere",
          "notes": "The escape hatch inside the OpenAI ecosystem. Served by Groq at ~500–1000 tok/s.",
          "_source": "https://huggingface.co/openai/gpt-oss-20b"
        },
        {
          "family": "Anthropic Claude (Fable / Opus / Sonnet / Haiku)",
          "license": "Proprietary API",
          "weights": "Closed",
          "selfhost": "No",
          "distil": "No — Commercial Terms D.4 bars building a competing product or training competing AI models",
          "residency": "inference_geo:\"us\" at a 1.1x multiplier; Bedrock/Vertex regional endpoints at +10%",
          "notes": "Best published SWE-bench Verified (Opus 5, 96.0). 1M context on Fable/Opus/Sonnet, 200K on Haiku 4.5.",
          "_source": "https://www.anthropic.com/legal/commercial-terms"
        },
        {
          "family": "Google Gemini 2.5 / 3.x",
          "license": "Proprietary API",
          "weights": "Closed",
          "selfhost": "No",
          "distil": "No — Gemini API Additional Terms restrict competitive model development",
          "residency": "Vertex AI regions",
          "notes": "Only closed vendor that has publicly documented distilling its own small tier (Gemini 2.5 report).",
          "_source": "https://arxiv.org/html/2507.06261v1/"
        },
        {
          "family": "Google Gemma 4",
          "license": "Apache 2.0",
          "weights": "Open",
          "selfhost": "Yes (2.3B–31B)",
          "distil": "Yes",
          "residency": "Anywhere",
          "notes": "MMLU-Pro 85.2 at 31B under Apache 2.0 — the strongest permissively-licensed model you can put on one node.",
          "_source": "https://ai.google.dev/gemma/docs/core/model_card_4"
        },
        {
          "family": "Google Gemma 3",
          "license": "Gemma Terms of Use (custom, use restrictions apply)",
          "weights": "Open",
          "selfhost": "Yes",
          "distil": "Yes, subject to the Gemma prohibited-use policy",
          "residency": "Anywhere",
          "notes": "Superseded by Gemma 4 on both quality and licence terms.",
          "_source": "https://huggingface.co/google/gemma-3-27b-it"
        },
        {
          "family": "Meta Llama 3.x / 4",
          "license": "Llama Community License (custom commercial; 700M MAU clause)",
          "weights": "Open",
          "selfhost": "Yes",
          "distil": "Yes, with Llama attribution and naming obligations on derivatives",
          "residency": "Anywhere",
          "notes": "Llama 3.2 1B/3B are explicitly documented distillations; Meta's Llama 4 launch post describes Scout/Maverick as codistilled from the unreleased Behemoth (the model card itself does not mention it).",
          "_source": "https://ai.meta.com/blog/llama-4-multimodal-intelligence/"
        },
        {
          "family": "DeepSeek R1 / V3.2 / V4",
          "license": "MIT",
          "weights": "Open",
          "selfhost": "Yes (V4-Pro is a 1.6T MoE — non-trivial)",
          "distil": "Yes — the R1 release shipped six distilled students itself",
          "residency": "Anywhere; first-party API is PRC-hosted",
          "notes": "Highest open-weight GPQA Diamond here (90.1). Many enterprises self-host rather than use the PRC-hosted API.",
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1"
        },
        {
          "family": "Alibaba Qwen3 / Qwen3.8 (open weights)",
          "license": "Apache 2.0",
          "weights": "Open",
          "selfhost": "Yes",
          "distil": "Yes",
          "residency": "Anywhere; Model Studio API is PRC/Singapore",
          "notes": "Qwen3 technical report documents strong-to-weak distillation for the 0.6B–14B dense sizes and 30B-A3B.",
          "_source": "https://huggingface.co/Qwen/Qwen3.8-27B"
        },
        {
          "family": "Microsoft Phi-4 / Phi-4-mini",
          "license": "MIT",
          "weights": "Open",
          "selfhost": "Yes",
          "distil": "Yes",
          "residency": "Anywhere",
          "notes": "Least legally encumbered small models in this dataset. Phi-4’s 16K context is the catch.",
          "_source": "https://huggingface.co/microsoft/phi-4"
        },
        {
          "family": "Mistral Small 4 / Ministral 3 / Large 3",
          "license": "Apache 2.0",
          "weights": "Open",
          "selfhost": "Yes",
          "distil": "Yes",
          "residency": "EU-headquartered vendor; EU hosting available",
          "notes": "The default answer when the requirement is EU sovereignty plus a permissive licence.",
          "_source": "https://docs.mistral.ai/getting-started/models/models_overview/"
        },
        {
          "family": "Mistral Medium 3.5",
          "license": "Modified MIT",
          "weights": "Open",
          "selfhost": "Yes",
          "distil": "Yes, subject to the modified terms",
          "residency": "EU hosting available",
          "notes": "Frontier-class tier of the Mistral line; read the modification before assuming MIT.",
          "_source": "https://docs.mistral.ai/getting-started/models/models_overview/"
        },
        {
          "family": "Amazon Nova / Nova 2",
          "license": "Proprietary API",
          "weights": "Closed",
          "selfhost": "No",
          "distil": "Only through Bedrock Model Distillation, into another Amazon-supported student",
          "residency": "AWS regions incl. GovCloud (US-West)",
          "notes": "The only vendor that documents its teacher/student graph in product docs: Premier → Pro/Lite/Micro, Pro → Lite/Micro.",
          "_source": "https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html"
        },
        {
          "family": "Hugging Face SmolLM3",
          "license": "Apache 2.0",
          "weights": "Open (plus data mixture and training configs)",
          "selfhost": "Yes",
          "distil": "Yes",
          "residency": "Anywhere",
          "notes": "Fully reproducible supply chain — relevant where an auditor asks what the model was trained on.",
          "_source": "https://huggingface.co/HuggingFaceTB/SmolLM3-3B"
        }
      ],
      "notes": "Anti-distillation clauses bind the enterprise account holder, not just individual developers, and survive termination in most of these agreements. If your roadmap includes training an in-house model on teacher outputs, pick an MIT or Apache 2.0 teacher up front rather than seeking a waiver later.",
      "sources": [
        "https://www.anthropic.com/legal/commercial-terms",
        "https://openai.com/policies/services-agreement/",
        "https://huggingface.co/openai/gpt-oss-20b",
        "https://ai.google.dev/gemma/docs/core/model_card_4",
        "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md",
        "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "https://huggingface.co/Qwen/Qwen3.8-27B",
        "https://huggingface.co/microsoft/phi-4",
        "https://docs.mistral.ai/getting-started/models/models_overview/",
        "https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html",
        "https://huggingface.co/HuggingFaceTB/SmolLM3-3B"
      ]
    },
    {
      "id": "retention",
      "title": "Quality retention by pair (documented distillations and same-family size comparisons)",
      "description": "What fraction of the teacher’s score the cheaper model keeps, and what fraction of the teacher’s input price it costs. Rows where the student is documented as a distillation are the ones with a named teacher in the master table; the rest are same-family price ladders.",
      "columns": [
        {
          "key": "pair",
          "label": "Teacher → student",
          "type": "text"
        },
        {
          "key": "relationship",
          "label": "Relationship",
          "type": "text"
        },
        {
          "key": "bench",
          "label": "Benchmark",
          "type": "text"
        },
        {
          "key": "teacher_score",
          "label": "Teacher",
          "type": "number",
          "unit": "%"
        },
        {
          "key": "student_score",
          "label": "Student",
          "type": "number",
          "unit": "%"
        },
        {
          "key": "retention_pct",
          "label": "Retention",
          "type": "number",
          "unit": "%"
        },
        {
          "key": "price_ratio",
          "label": "Student price / teacher price (input)",
          "type": "number",
          "unit": "x"
        }
      ],
      "rows": [
        {
          "pair": "GPT-5.6 Sol → Terra",
          "bench": "GPQA Diamond",
          "teacher_score": 92.4,
          "student_score": 88.4,
          "retention_pct": 95.7,
          "price_ratio": 0.5,
          "_source": "https://openrouter.ai/openai/gpt-5.6-terra",
          "relationship": "Same-generation tier comparison (method undisclosed)"
        },
        {
          "pair": "GPT-5.6 Sol → Luna",
          "bench": "GPQA Diamond",
          "teacher_score": 92.4,
          "student_score": 87,
          "retention_pct": 94.2,
          "price_ratio": 0.05,
          "_source": "https://openrouter.ai/openai/gpt-5.6-luna",
          "relationship": "Same-generation tier comparison (method undisclosed); third-party benchmark, 94.2–97.3% across providers"
        },
        {
          "pair": "GPT-5.4 → GPT-5.4 mini",
          "bench": "GPQA Diamond",
          "teacher_score": 93,
          "student_score": 88,
          "retention_pct": 94.6,
          "price_ratio": 0.3,
          "_source": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/",
          "relationship": "Same-generation tier comparison (method undisclosed)"
        },
        {
          "pair": "GPT-5.4 → GPT-5.4 nano",
          "bench": "GPQA Diamond",
          "teacher_score": 93,
          "student_score": 82.8,
          "retention_pct": 89,
          "price_ratio": 0.08,
          "_source": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/",
          "relationship": "Same-generation tier comparison (method undisclosed)"
        },
        {
          "pair": "o3 → o4-mini",
          "bench": "GPQA Diamond",
          "teacher_score": 83.3,
          "student_score": 81.4,
          "retention_pct": 97.7,
          "price_ratio": 0.55,
          "_source": "https://www.datacamp.com/blog/o4-mini",
          "relationship": "Same-generation tier comparison (method undisclosed)"
        },
        {
          "pair": "Gemini 2.5 Pro → 2.5 Flash",
          "bench": "GPQA Diamond",
          "teacher_score": 86.4,
          "student_score": 82.8,
          "retention_pct": 95.8,
          "price_ratio": 0.24,
          "_source": "https://arxiv.org/html/2507.06261v1/",
          "relationship": "Documented distillation"
        },
        {
          "pair": "DeepSeek V4-Pro vs V4-Flash",
          "bench": "GPQA Diamond",
          "teacher_score": 90.1,
          "student_score": 88.1,
          "retention_pct": 97.8,
          "price_ratio": 0.333,
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
          "relationship": "Same-family price/quality comparison — V4-Flash is a consolidation of V4 domain experts via on-policy distillation, not a compression of V4-Pro"
        },
        {
          "pair": "DeepSeek-R1 → R1-Distill-Llama-70B",
          "bench": "GPQA Diamond",
          "teacher_score": 71.5,
          "student_score": 65.2,
          "retention_pct": 91.2,
          "price_ratio": null,
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
          "relationship": "Documented distillation"
        },
        {
          "pair": "DeepSeek-R1 → R1-Distill-Qwen-32B",
          "bench": "GPQA Diamond",
          "teacher_score": 71.5,
          "student_score": 62.1,
          "retention_pct": 86.9,
          "price_ratio": null,
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
          "relationship": "Documented distillation"
        },
        {
          "pair": "DeepSeek-R1 → R1-Distill-Qwen-14B",
          "bench": "GPQA Diamond",
          "teacher_score": 71.5,
          "student_score": 59.1,
          "retention_pct": 82.7,
          "price_ratio": null,
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
          "relationship": "Documented distillation"
        },
        {
          "pair": "DeepSeek-R1 → R1-Distill-Llama-8B",
          "bench": "GPQA Diamond",
          "teacher_score": 71.5,
          "student_score": 49,
          "retention_pct": 68.5,
          "price_ratio": null,
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
          "relationship": "Documented distillation"
        },
        {
          "pair": "DeepSeek-R1 → R1-Distill-Qwen-7B",
          "bench": "GPQA Diamond",
          "teacher_score": 71.5,
          "student_score": 49.1,
          "retention_pct": 68.7,
          "price_ratio": null,
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
          "relationship": "Documented distillation"
        },
        {
          "pair": "DeepSeek-R1 → R1-Distill-Qwen-1.5B",
          "bench": "GPQA Diamond",
          "teacher_score": 71.5,
          "student_score": 33.8,
          "retention_pct": 47.3,
          "price_ratio": null,
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
          "relationship": "Documented distillation"
        },
        {
          "pair": "Gemma 4 31B vs Gemma 4 E4B",
          "bench": "GPQA Diamond",
          "teacher_score": 84.3,
          "student_score": 58.6,
          "retention_pct": 69.5,
          "price_ratio": null,
          "_source": "https://ai.google.dev/gemma/docs/core/model_card_4",
          "relationship": "Same-family size comparison — Google does not describe E4B as a distillation of 31B"
        },
        {
          "pair": "Llama 4 Maverick → Scout",
          "bench": "MMLU-Pro",
          "teacher_score": 80.5,
          "student_score": 74.3,
          "retention_pct": 92.3,
          "price_ratio": null,
          "_source": "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md",
          "relationship": "Documented codistillation (both from Behemoth), compared here by size"
        },
        {
          "pair": "Nova Pro vs Nova Lite",
          "bench": "MMLU",
          "teacher_score": 85.9,
          "student_score": 80.5,
          "retention_pct": 93.7,
          "price_ratio": 0.075,
          "_source": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf",
          "relationship": "Same-family size comparison — Bedrock supports both as distillation students; shipped provenance undisclosed"
        },
        {
          "pair": "Nova Pro vs Nova Micro",
          "bench": "MMLU",
          "teacher_score": 85.9,
          "student_score": 77.6,
          "retention_pct": 90.3,
          "price_ratio": 0.044,
          "_source": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf",
          "relationship": "Same-family size comparison — Bedrock supports both as distillation students; shipped provenance undisclosed"
        },
        {
          "pair": "Claude Opus 5 → Sonnet 5",
          "bench": "SWE-bench Verified",
          "teacher_score": 96,
          "student_score": 85.2,
          "retention_pct": 88.8,
          "price_ratio": 0.4,
          "_source": "https://www.morphllm.com/claude-benchmarks",
          "relationship": "Same-generation tier comparison (method undisclosed)"
        },
        {
          "pair": "Claude Opus 5 → Haiku 4.5",
          "bench": "SWE-bench Verified",
          "teacher_score": 96,
          "student_score": 73.3,
          "retention_pct": 76.4,
          "price_ratio": 0.2,
          "_source": "https://datanorth.ai/news/claude-opus-5-by-anthropic",
          "relationship": "Cross-generation tier comparison (method undisclosed)"
        },
        {
          "pair": "Llama 4 Maverick → Scout (coding)",
          "bench": "LiveCodeBench",
          "teacher_score": 43.4,
          "student_score": 32.8,
          "retention_pct": 75.6,
          "price_ratio": null,
          "_source": "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md",
          "relationship": "Documented codistillation (both from Behemoth), compared here by size"
        },
        {
          "pair": "DeepSeek V4-Pro vs V4-Flash (coding)",
          "bench": "SWE-bench Verified",
          "teacher_score": 80.6,
          "student_score": 79,
          "retention_pct": 98,
          "price_ratio": 0.333,
          "_source": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
          "relationship": "Same-family price/quality comparison — see row 6"
        },
        {
          "pair": "Gemini 2.5 Pro → 2.5 Flash (coding)",
          "bench": "SWE-bench Verified",
          "teacher_score": 67.2,
          "student_score": 60.3,
          "retention_pct": 89.7,
          "price_ratio": 0.24,
          "_source": "https://arxiv.org/html/2507.06261v1/",
          "relationship": "Documented distillation"
        }
      ],
      "notes": "Two patterns stand out. First, knowledge retention above 94% is now routine inside a family, and DeepSeek V4-Flash reaches 97.8% for a third of the price. Second, retention falls off a cliff for very small students on broad-knowledge benchmarks — R1-Distill-Qwen-1.5B keeps only 47% of R1’s GPQA — while the same model keeps 86% of R1’s MATH-500. Distillation buys you narrow competence cheaply and broad competence expensively. Only rows marked \"Documented distillation\" carry a vendor statement that the student was trained from the teacher. The remaining rows are same-family or same-generation price/quality comparisons and should not be read as training provenance.",
      "sources": [
        "https://openrouter.ai/openai/gpt-5.6-luna",
        "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/",
        "https://arxiv.org/html/2507.06261v1/",
        "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
        "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "https://ai.google.dev/gemma/docs/core/model_card_4",
        "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md",
        "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf",
        "https://www.morphllm.com/claude-benchmarks",
        "https://datanorth.ai/news/claude-opus-5-by-anthropic",
        "https://www.datacamp.com/blog/o4-mini"
      ]
    }
  ],
  "charts": [
    {
      "id": "price-vs-mmlupro",
      "title": "Blended price vs MMLU-Pro",
      "type": "scatter",
      "xLabel": "Blended price (USD per million tokens, 3:1 input:output)",
      "yLabel": "MMLU-Pro (%)",
      "unit": "%",
      "series": [
        {
          "name": "Frontier teachers",
          "data": [
            {
              "x": 4.5,
              "y": 92.6,
              "label": "Gemini 3.1 Pro (Preview)"
            },
            {
              "x": 1.98,
              "y": 87.5,
              "label": "DeepSeek-V4-Pro"
            }
          ]
        },
        {
          "name": "Small siblings",
          "data": [
            {
              "x": 0.5625,
              "y": 83,
              "label": "Gemini 3.1 Flash-Lite"
            },
            {
              "x": 0.2625,
              "y": 78,
              "label": "Mistral Small 4"
            }
          ]
        },
        {
          "name": "Documented distillations",
          "data": [
            {
              "x": 0.66,
              "y": 86.2,
              "label": "DeepSeek-V4-Flash"
            }
          ]
        },
        {
          "name": "Open weights (hosted price)",
          "data": [
            {
              "x": 1.04,
              "y": 68.9,
              "label": "Llama 3.3 70B Instruct"
            }
          ]
        }
      ],
      "notes": "Only models that publish MMLU-Pro AND have a sourced token price appear here. The frontier is almost flat between $0.26 and $2.00: Mistral Small 4 buys 78.0 MMLU-Pro for $0.26 blended, DeepSeek-V4-Flash 86.2 for $0.66, DeepSeek-V4-Pro 87.5 for $1.98, and Gemini 3.1 Pro 92.6 for $4.50 — so the last 6 MMLU-Pro points cost roughly 7x. Gemma 4 31B (85.2) and Gemma 4 26B A4B (82.6) sit off this chart because they are self-host-only and therefore have no list token price.",
      "sources": [
        "https://ai.google.dev/gemini-api/docs/pricing",
        "https://deepmind.google/models/gemini/pro/",
        "https://layerlens.ai/blog/gemini-3-1-flash-lite-benchmark-results-efficiency-model-comparison",
        "https://api-docs.deepseek.com/quick_start/pricing/",
        "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro",
        "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
        "https://mistral.ai/pricing/api",
        "https://openrouter.ai/mistralai/mistral-small-2603",
        "https://www.together.ai/pricing",
        "https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md"
      ]
    },
    {
      "id": "price-vs-gpqa",
      "title": "Blended price vs GPQA Diamond — the whole shortlist on one plot",
      "type": "scatter",
      "xLabel": "Blended price (USD per million tokens, 3:1 input:output)",
      "yLabel": "GPQA Diamond (%)",
      "unit": "%",
      "series": [
        {
          "name": "Frontier teachers",
          "data": [
            {
              "x": 8,
              "y": 92.4,
              "label": "GPT-5.6 Sol"
            },
            {
              "x": 5.625,
              "y": 93,
              "label": "GPT-5.4"
            },
            {
              "x": 3.5,
              "y": 83.3,
              "label": "o3"
            },
            {
              "x": 4.5,
              "y": 94.3,
              "label": "Gemini 3.1 Pro (Preview)"
            },
            {
              "x": 3.4375,
              "y": 86.4,
              "label": "Gemini 2.5 Pro"
            },
            {
              "x": 1.98,
              "y": 90.1,
              "label": "DeepSeek-V4-Pro"
            },
            {
              "x": 1.4,
              "y": 46.9,
              "label": "Amazon Nova Pro"
            }
          ]
        },
        {
          "name": "Small siblings (method undisclosed)",
          "data": [
            {
              "x": 4.5,
              "y": 88.4,
              "label": "GPT-5.6 Terra"
            },
            {
              "x": 0.45,
              "y": 87,
              "label": "GPT-5.6 Luna"
            },
            {
              "x": 1.6875,
              "y": 88,
              "label": "GPT-5.4 mini"
            },
            {
              "x": 0.4625,
              "y": 82.8,
              "label": "GPT-5.4 nano"
            },
            {
              "x": 0.6875,
              "y": 80.3,
              "label": "GPT-5 mini"
            },
            {
              "x": 0.1375,
              "y": 70.9,
              "label": "GPT-5 nano"
            },
            {
              "x": 1.925,
              "y": 81.4,
              "label": "o4-mini"
            },
            {
              "x": 0.5625,
              "y": 72.2,
              "label": "Gemini 3.1 Flash-Lite"
            },
            {
              "x": 0.2625,
              "y": 71.2,
              "label": "Mistral Small 4"
            },
            {
              "x": 0.105,
              "y": 42,
              "label": "Amazon Nova Lite"
            },
            {
              "x": 0.0613,
              "y": 40,
              "label": "Amazon Nova Micro"
            }
          ]
        },
        {
          "name": "Documented distillations",
          "data": [
            {
              "x": 0.85,
              "y": 82.8,
              "label": "Gemini 2.5 Flash"
            },
            {
              "x": 0.66,
              "y": 88.1,
              "label": "DeepSeek-V4-Flash"
            }
          ]
        },
        {
          "name": "Open weights",
          "data": [
            {
              "x": 0.1312,
              "y": 58.59,
              "label": "gpt-oss-20b"
            },
            {
              "x": 1.04,
              "y": 50.5,
              "label": "Llama 3.3 70B Instruct"
            },
            {
              "x": 1.125,
              "y": 89.2,
              "label": "Qwen3.8-27B"
            },
            {
              "x": 0.2,
              "y": 71.2,
              "label": "Ministral 3 14B Instruct"
            }
          ]
        }
      ],
      "notes": "The upper-left corner is the whole story: GPT-5.6 Luna at $0.45 blended / 87.0 GPQA and DeepSeek-V4-Flash at $0.66 / 88.1 sit close to the frontier for a fraction of the price. Gemini 3.1 Flash-Lite is cheaper still at $0.56 blended but scores 72.2 GPQA Diamond, so it trades roughly 15 points of knowledge benchmark for the saving. Paying 8–18x more than Luna or V4-Flash moves you 5–7 GPQA points. Amazon Nova rows use the 2024 Nova technical report’s GPQA methodology and are not comparable with the reasoning-era scores.",
      "sources": [
        "https://developers.openai.com/api/docs/pricing",
        "https://openrouter.ai/openai/gpt-5.6-sol",
        "https://openrouter.ai/openai/gpt-5.6-terra",
        "https://openrouter.ai/openai/gpt-5.6-luna",
        "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/",
        "https://www.datacamp.com/blog/o4-mini",
        "https://ai.google.dev/gemini-api/docs/pricing",
        "https://arxiv.org/html/2507.06261v1/",
        "https://layerlens.ai/blog/gemini-3-1-flash-lite-benchmark-results-efficiency-model-comparison",
        "https://api-docs.deepseek.com/quick_start/pricing/",
        "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro",
        "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
        "https://mistral.ai/pricing/api",
        "https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512",
        "https://console.groq.com/docs/models",
        "https://huggingface.co/Qwen/Qwen3.8-27B",
        "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf"
      ]
    },
    {
      "id": "teacher-vs-distilled-gpqa",
      "title": "Larger vs smaller tier on GPQA Diamond (and MMLU for Nova)",
      "type": "bar",
      "xLabel": "Pair (documented distillations and same-family comparisons)",
      "yLabel": "Benchmark score (%)",
      "unit": "%",
      "series": [
        {
          "name": "Larger / teacher tier",
          "data": [
            {
              "x": "GPT-5.6 Sol → Luna",
              "y": 92.4
            },
            {
              "x": "GPT-5.4 → GPT-5.4 mini",
              "y": 93
            },
            {
              "x": "o3 → o4-mini",
              "y": 83.3
            },
            {
              "x": "Gemini 2.5 Pro → 2.5 Flash",
              "y": 86.4
            },
            {
              "x": "DeepSeek V4-Pro vs V4-Flash",
              "y": 90.1
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Llama-70B",
              "y": 71.5
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Qwen-32B",
              "y": 71.5
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Qwen-7B",
              "y": 71.5
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Qwen-1.5B",
              "y": 71.5
            },
            {
              "x": "Gemma 4 31B vs Gemma 4 E4B",
              "y": 84.3
            },
            {
              "x": "Nova Pro vs Nova Lite",
              "y": 85.9
            }
          ]
        },
        {
          "name": "Smaller tier (distilled or same-family)",
          "data": [
            {
              "x": "GPT-5.6 Sol → Luna",
              "y": 87
            },
            {
              "x": "GPT-5.4 → GPT-5.4 mini",
              "y": 88
            },
            {
              "x": "o3 → o4-mini",
              "y": 81.4
            },
            {
              "x": "Gemini 2.5 Pro → 2.5 Flash",
              "y": 82.8
            },
            {
              "x": "DeepSeek V4-Pro vs V4-Flash",
              "y": 88.1
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Llama-70B",
              "y": 65.2
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Qwen-32B",
              "y": 62.1
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Qwen-7B",
              "y": 49.1
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Qwen-1.5B",
              "y": 33.8
            },
            {
              "x": "Gemma 4 31B vs Gemma 4 E4B",
              "y": 58.6
            },
            {
              "x": "Nova Pro vs Nova Lite",
              "y": 80.5
            }
          ]
        }
      ],
      "notes": "Inside a single generation the bars are nearly the same height. Across a large parameter gap they are not: R1 → R1-Distill-Qwen-1.5B loses more than half the teacher's GPQA, and Gemma 4 31B vs E4B loses 30%. Only the DeepSeek-R1 and Gemini 2.5 pairs are vendor-documented distillations; DeepSeek V4-Pro vs V4-Flash, Gemma 4 and Nova are same-family comparisons — V4-Flash is documented as a consolidation of V4 domain experts via on-policy distillation, not a compression of V4-Pro. Nova uses MMLU rather than GPQA because that is what the Amazon technical report publishes.",
      "sources": [
        "https://openrouter.ai/openai/gpt-5.6-luna",
        "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/",
        "https://www.datacamp.com/blog/o4-mini",
        "https://arxiv.org/html/2507.06261v1/",
        "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
        "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "https://ai.google.dev/gemma/docs/core/model_card_4",
        "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf"
      ]
    },
    {
      "id": "cost-per-1m-requests",
      "title": "Cost per 1,000,000 requests (400 input + 100 output tokens each)",
      "type": "bar",
      "xLabel": "Model",
      "yLabel": "USD per 1M requests",
      "unit": "USD",
      "series": [
        {
          "name": "List price, no caching or batch discount",
          "data": [
            {
              "x": "Amazon Nova Micro",
              "y": 28
            },
            {
              "x": "qwen-turbo",
              "y": 40
            },
            {
              "x": "Amazon Nova Lite",
              "y": 48
            },
            {
              "x": "GPT-5 nano",
              "y": 60
            },
            {
              "x": "gpt-oss-20b",
              "y": 60
            },
            {
              "x": "Ministral 3 14B Instruct",
              "y": 100
            },
            {
              "x": "qwen3.8-flash",
              "y": 107
            },
            {
              "x": "GPT-4o mini",
              "y": 120
            },
            {
              "x": "Mistral Small 4",
              "y": 120
            },
            {
              "x": "gpt-oss-120b",
              "y": 120
            },
            {
              "x": "GPT-5.6 Luna",
              "y": 200
            },
            {
              "x": "GPT-5 mini",
              "y": 300
            },
            {
              "x": "DeepSeek-V4-Flash",
              "y": 308
            },
            {
              "x": "Gemini 2.5 Flash",
              "y": 370
            },
            {
              "x": "Gemini 3.5 Flash-Lite",
              "y": 370
            },
            {
              "x": "Amazon Nova 2 Lite",
              "y": 370
            },
            {
              "x": "Llama 3.3 70B Instruct",
              "y": 520
            },
            {
              "x": "Gemini 3.8 Flash",
              "y": 675
            },
            {
              "x": "GPT-5.4 mini",
              "y": 750
            },
            {
              "x": "o4-mini",
              "y": 880
            },
            {
              "x": "Claude Haiku 4.5",
              "y": 900
            },
            {
              "x": "DeepSeek-V4-Pro",
              "y": 924
            },
            {
              "x": "Claude Sonnet 5",
              "y": 1800
            },
            {
              "x": "Gemini 3.1 Pro (Preview)",
              "y": 2000
            },
            {
              "x": "GPT-5.6 Terra",
              "y": 2000
            },
            {
              "x": "GPT-5.6 Sol",
              "y": 3600
            },
            {
              "x": "Claude Opus 5",
              "y": 4500
            },
            {
              "x": "Claude Fable 5.1",
              "y": 9000
            },
            {
              "x": "GPT-6 Astra",
              "y": 9000
            }
          ]
        }
      ],
      "notes": "A classification or extraction workload at a million requests a month costs $28 on Amazon Nova Micro, $200 on GPT-5.6 Luna, $900 on Claude Haiku 4.5, $1,800 on Claude Sonnet 5 and $9,000 on the two most expensive tiers, Claude Fable 5.1 and GPT-6 Astra (both $10/$50) — a 320x spread across the same shortlist. Batch APIs cut these by 50% at both Anthropic and OpenAI, and prompt caching cuts the input component by up to 90% (97.5% on Claude Fable 5.1), which is usually a bigger lever than moving one tier down.",
      "sources": [
        "https://developers.openai.com/api/docs/pricing",
        "https://platform.claude.com/docs/en/about-claude/pricing",
        "https://ai.google.dev/gemini-api/docs/pricing",
        "https://api-docs.deepseek.com/quick_start/pricing/",
        "https://mistral.ai/pricing/api",
        "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "https://console.groq.com/docs/models",
        "https://www.together.ai/pricing",
        "https://pricepertoken.com/pricing-page/model/amazon-nova-micro-v1",
        "https://pricepertoken.com/pricing-page/model/amazon-nova-lite-v1",
        "https://pricepertoken.com/pricing-page/model/amazon-nova-2-lite-v1"
      ]
    },
    {
      "id": "quality-retention",
      "title": "Quality retention: smaller tier score as a percentage of the larger tier",
      "type": "bar",
      "xLabel": "Pair (documented distillations and same-family comparisons)",
      "yLabel": "Retention (%)",
      "unit": "%",
      "series": [
        {
          "name": "Knowledge benchmarks (GPQA-D / MMLU / MMLU-Pro)",
          "data": [
            {
              "x": "GPT-5.6 Sol → Terra",
              "y": 95.7
            },
            {
              "x": "GPT-5.6 Sol → Luna",
              "y": 94.2
            },
            {
              "x": "GPT-5.4 → GPT-5.4 mini",
              "y": 94.6
            },
            {
              "x": "GPT-5.4 → GPT-5.4 nano",
              "y": 89
            },
            {
              "x": "o3 → o4-mini",
              "y": 97.7
            },
            {
              "x": "Gemini 2.5 Pro → 2.5 Flash",
              "y": 95.8
            },
            {
              "x": "DeepSeek V4-Pro vs V4-Flash",
              "y": 97.8
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Llama-70B",
              "y": 91.2
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Qwen-32B",
              "y": 86.9
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Qwen-14B",
              "y": 82.7
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Llama-8B",
              "y": 68.5
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Qwen-7B",
              "y": 68.7
            },
            {
              "x": "DeepSeek-R1 → R1-Distill-Qwen-1.5B",
              "y": 47.3
            },
            {
              "x": "Gemma 4 31B vs Gemma 4 E4B",
              "y": 69.5
            },
            {
              "x": "Llama 4 Maverick → Scout",
              "y": 92.3
            },
            {
              "x": "Nova Pro vs Nova Lite",
              "y": 93.7
            },
            {
              "x": "Nova Pro vs Nova Micro",
              "y": 90.3
            }
          ]
        },
        {
          "name": "Coding / agentic benchmarks",
          "data": [
            {
              "x": "Claude Opus 5 → Sonnet 5",
              "y": 88.8
            },
            {
              "x": "Claude Opus 5 → Haiku 4.5",
              "y": 76.4
            },
            {
              "x": "Llama 4 Maverick → Scout (coding)",
              "y": 75.6
            },
            {
              "x": "DeepSeek V4-Pro vs V4-Flash (coding)",
              "y": 98
            },
            {
              "x": "Gemini 2.5 Pro → 2.5 Flash (coding)",
              "y": 89.7
            }
          ]
        }
      ],
      "notes": "Read the two series against each other. Within one generation, knowledge retention clusters at 94–98%. Coding and agentic retention is bimodal: DeepSeek V4-Flash keeps 98% of V4-Pro on SWE-bench Verified, but Claude Haiku 4.5 keeps only 76% of Opus 5 and Llama 4 Scout only 76% of Maverick on LiveCodeBench. Only the DeepSeek-R1, Gemini 2.5 and Llama 4 rows are vendor-documented distillations; DeepSeek V4-Pro vs V4-Flash, Gemma 4 and Nova are same-family size comparisons (V4-Flash is documented as a consolidation of V4 domain experts, not a compression of V4-Pro). That spread, not the price sheet, is what should decide whether a workload can be moved down a tier.",
      "sources": [
        "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
        "https://arxiv.org/html/2507.06261v1/",
        "https://www.morphllm.com/claude-benchmarks",
        "https://datanorth.ai/news/claude-opus-5-by-anthropic",
        "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md",
        "https://ai.google.dev/gemma/docs/core/model_card_4",
        "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf",
        "https://openrouter.ai/openai/gpt-5.6-luna",
        "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/"
      ]
    },
    {
      "id": "inference-price-index",
      "title": "Enterprise inference price index, 2026",
      "type": "line",
      "xLabel": "Date",
      "yLabel": "USD per million tokens",
      "unit": "USD",
      "series": [
        {
          "name": "Silicon Data enterprise inference index (cited by Jefferies)",
          "data": [
            {
              "x": "2026-05-31",
              "y": 2.04
            },
            {
              "x": "2026-07-25",
              "y": 1.45
            },
            {
              "x": "2026-08-07",
              "y": 1.17
            }
          ]
        }
      ],
      "notes": "A 43% fall in ten weeks. Jefferies attributes it to OpenAI cutting GPT-5.6 rates by up to 80%, Anthropic matching prior frontier performance at half the price with Claude Opus 5, and Chinese open-weight models such as DeepSeek V4-Flash-0731 competing at roughly $0.03 per task. The 7 August figure is the midpoint of the reported $1.16–$1.18 low.",
      "sources": [
        "https://www.scmp.com/tech/tech-trends/article/3363549/enterprise-ai-costs-hit-2026-low-driven-price-wars-chinese-open-source-models-research",
        "https://www.axios.com/2026/07/30/openai-cuts-prices-gpt-terra-luna5"
      ]
    }
  ],
  "timeline": [
    {
      "date": "2024-07-18",
      "title": "GPT-4o mini launches at $0.15/$0.60",
      "detail": "MMLU 82.0, HumanEval 87.2, MMMU 59.4 — over 60% cheaper than GPT-3.5 Turbo. Establishes the \"cheap tier\" as a permanent product line every vendor now copies.",
      "category": "product",
      "source": "https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/"
    },
    {
      "date": "2024-09-25",
      "title": "Llama 3.2 1B/3B ship as documented distillations",
      "detail": "Meta states that logits from Llama 3.1 8B and 70B were used as token-level targets in pretraining, with distillation applied after pruning to recover performance. The first mainstream model card to spell out the recipe.",
      "category": "product",
      "source": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md"
    },
    {
      "date": "2024-12-03",
      "title": "Amazon Bedrock Model Distillation announced",
      "detail": "Turns distillation into a managed service: the customer supplies prompts, AWS generates teacher responses and fine-tunes the student. Later reaches GA with claims of up to 500% faster and 75% cheaper inference.",
      "category": "product",
      "source": "https://aws.amazon.com/bedrock/model-distillation/"
    },
    {
      "date": "2024-12-06",
      "title": "Llama 3.3 70B released",
      "detail": "MMLU 86.0 CoT, MMLU-Pro 68.9, HumanEval 88.4, 128K context under the Llama 3.3 Community License. Becomes the default open student base for enterprise distillation projects.",
      "category": "product",
      "source": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md"
    },
    {
      "date": "2025-01-20",
      "title": "DeepSeek ships R1 plus six distilled students under MIT",
      "detail": "R1-Distill-Qwen-32B beats o1-mini on AIME 2024 (72.6 vs 63.6), MATH-500 (94.3 vs 90.0) and GPQA Diamond (62.1 vs 60.0). For buyers this was the first time a free download credibly replaced a paid reasoning tier.",
      "category": "product",
      "source": "https://huggingface.co/deepseek-ai/DeepSeek-R1"
    },
    {
      "date": "2025-04",
      "title": "Llama 4 Scout and Maverick ship as codistilled models",
      "detail": "Meta's launch post says both are codistilled from the ~2T-parameter Llama 4 Behemoth, which was never released; the Llama 4 model card carries the benchmarks but not the provenance claim. Maverick posts MMLU-Pro 80.5 / GPQA-D 69.8; Scout adds a 10M-token context window.",
      "category": "product",
      "source": "https://ai.meta.com/blog/llama-4-multimodal-intelligence/"
    },
    {
      "date": "2025-06",
      "title": "Gemini 2.5 report confirms the Flash line is distilled",
      "detail": "Google states the smaller 2.5 models are distilled and that the teacher’s next-token distribution is approximated with a k-sparse distribution to cut storage cost. Rare public confirmation from a closed-model vendor.",
      "category": "research",
      "source": "https://arxiv.org/html/2507.06261v1/"
    },
    {
      "date": "2025-08-07",
      "title": "GPT-5, GPT-5 mini and GPT-5 nano launch together",
      "detail": "A three-tier release at $1.25/$10, $0.25/$2 and $0.05/$0.40 with 400K context on all three. Tiered families become the default shape of a frontier launch.",
      "category": "product",
      "source": "https://developers.openai.com/api/docs/pricing"
    },
    {
      "date": "2025-10-15",
      "title": "Claude Haiku 4.5 released at $1/$5",
      "detail": "SWE-bench Verified 73.3. Anthropic positions it as matching Claude Sonnet 4 coding performance at a third of the cost and more than twice the speed.",
      "category": "product",
      "source": "https://www.anthropic.com/news/claude-haiku-4-5"
    },
    {
      "date": "2025-12-02",
      "title": "Ministral 3 (3B/8B/14B) ships Apache 2.0 with 256K context",
      "detail": "The 14B posts GPQA Diamond 71.2, AIME25 85.0 and MATH 90.4 — edge-class weights with no licence friction, which matters for EU buyers.",
      "category": "product",
      "source": "https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512"
    },
    {
      "date": "2026-03-16",
      "title": "Mistral Small 4 released under Apache 2.0",
      "detail": "MMLU-Pro 78.0, GPQA Diamond 71.2, priced at $0.15/$0.60 on the Mistral API. Becomes the reference EU-sovereign small model.",
      "category": "product",
      "source": "https://openrouter.ai/mistralai/mistral-small-2603"
    },
    {
      "date": "2026-04-02",
      "title": "Gemma 4 ships Apache 2.0 with frontier-class small models",
      "detail": "The 31B dense posts MMLU-Pro 85.2 and GPQA Diamond 84.3; the 26B MoE reaches 82.3 GPQA with only 3.8B active parameters. Google moves the Gemma line from a custom licence to Apache 2.0.",
      "category": "product",
      "source": "https://deepmind.google/models/gemma/gemma-4/"
    },
    {
      "date": "2026-04-26",
      "title": "DeepSeek-V4-Pro released under MIT",
      "detail": "1.6T total / 49B active parameters, 1M context, MMLU-Pro 87.5, GPQA Diamond 90.1, SWE-bench Verified 80.6. The strongest openly-licensed model a buyer can self-host.",
      "category": "product",
      "source": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro"
    },
    {
      "date": "2026-06-30",
      "title": "Claude Sonnet 5 launches at $2/$10 with 1M context",
      "detail": "SWE-bench Verified 85.2, Terminal-Bench 2.1 80.4 — close to Opus 4.8 at a fraction of the price. Introductory pricing later made permanent.",
      "category": "product",
      "source": "https://www.anthropic.com/news/claude-sonnet-5"
    },
    {
      "date": "2026-07-09",
      "title": "GPT-5.6 Sol, Terra and Luna launch as one price ladder",
      "detail": "All three carry a 1.05M-token context. GPQA Diamond 92.4 / 88.4 / 87.0 across a 20x price spread — the clearest published price-vs-quality ladder in the market.",
      "category": "product",
      "source": "https://openrouter.ai/openai/gpt-5.6-sol"
    },
    {
      "date": "2026-07-24",
      "title": "Claude Opus 5 released at $5/$25",
      "detail": "SWE-bench Verified 96.0, SWE-bench Pro 79.2, OSWorld 2.0 70.6. Anthropic positions it as near-Fable frontier quality at half the price, holding the Opus price flat.",
      "category": "product",
      "source": "https://platform.claude.com/docs/en/models/opus-5/overview"
    },
    {
      "date": "2026-07-30",
      "title": "OpenAI cuts GPT-5.6 Luna by 80% and Terra by 20%",
      "detail": "Luna drops to $0.20/$1.20. This single move reset the floor for closed-model pricing and is a major driver of the mid-2026 fall in the enterprise inference index.",
      "category": "market",
      "source": "https://www.axios.com/2026/07/30/openai-cuts-prices-gpt-terra-luna5"
    },
    {
      "date": "2026-08-11",
      "title": "Anthropic makes Sonnet 5 $2/$10 permanent",
      "detail": "The scheduled 1 September increase to $3/$15 is cancelled. A rare case of a vendor withdrawing an announced price rise under competitive pressure.",
      "category": "market",
      "source": "https://platform.claude.com/docs/en/about-claude/pricing"
    },
    {
      "date": "2026-08-08",
      "title": "Enterprise inference index hits a 2026 low of $1.16–$1.18/MTok",
      "detail": "Silicon Data index cited by Jefferies, down from $2.04 on 31 May and $1.45 in late July, driven by OpenAI’s cuts, Anthropic’s Opus 5 repricing and Chinese open-weight competition.",
      "category": "market",
      "source": "https://www.scmp.com/tech/tech-trends/article/3363549/enterprise-ai-costs-hit-2026-low-driven-price-wars-chinese-open-source-models-research"
    },
    {
      "date": "2026-08-16",
      "title": "DeepSeek raises V4 standard rates ~3–4.7x, and cache-hit input rates by up to 11x",
      "detail": "Standard rates: V4-Flash goes from $0.14/$0.28 to $0.44/$1.32 at peak (about 3.1x input, 4.7x output); V4-Pro from $0.435/$0.87 to $1.32/$3.96 (about 3.0x / 4.6x). InfoWorld's \"more than 10x\" headline refers specifically to cache-hit input tokens, which rose between 52% and 1,100%. Capacity, not competition, sets the floor for the cheapest tiers.",
      "category": "market",
      "source": "https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html"
    },
    {
      "date": "2026-09-02",
      "title": "Gemini 3.8 Flash ships at $0.75/$3.75",
      "detail": "HLE-Verified 54.9 and 1M context per 9to5Google; Terminal-Bench 2.1 89.4 and DeepSWE v1.1 73.7 per Google's Gemini Flash model page (https://deepmind.google/models/gemini/flash/). Google's third Flash update in three months — the previous model, Gemini 3.7 Flash, shipped three weeks earlier — evidence that the cheap tier is now the fastest-moving part of the market.",
      "category": "product",
      "source": "https://9to5google.com/2026/09/02/gemini-3-8-flash-launch/"
    }
  ],
  "glossary": [
    {
      "term": "Knowledge distillation",
      "definition": "Training a small \"student\" model to reproduce the behaviour of a large \"teacher\" — either from its output text (black-box) or from its output probability distribution over tokens (white-box logit distillation)."
    },
    {
      "term": "Teacher / student",
      "definition": "The large source model and the small target model in a distillation. AWS documents Nova Premier as a teacher to Nova Pro, Lite and Micro, and Nova Pro as a teacher to Lite and Micro."
    },
    {
      "term": "Quality retention",
      "definition": "The student’s benchmark score as a percentage of the teacher’s on the same benchmark. Useful for buyers because it is scale-free, but it varies enormously by benchmark: broad knowledge retains worse than narrow maths."
    },
    {
      "term": "Small sibling",
      "definition": "A cheaper tier released alongside a frontier model (mini, nano, Flash, Haiku, Luna) where the vendor has not publicly documented how it was built. Behaves like a distillation commercially whether or not it is one technically."
    },
    {
      "term": "k-sparse distillation",
      "definition": "Storing only the top-k entries of the teacher’s next-token probability distribution instead of the full vocabulary, to make logit distillation affordable at scale. Documented in the Gemini 2.5 technical report."
    },
    {
      "term": "Codistillation",
      "definition": "Training several student models jointly against a shared teacher during the teacher’s own training run. Meta used this for Llama 4 Scout and Maverick against Llama 4 Behemoth."
    },
    {
      "term": "Strong-to-weak distillation",
      "definition": "Alibaba’s two-phase Qwen3 pipeline: off-policy distillation on teacher outputs in both thinking and non-thinking modes, then on-policy distillation aligning student logits with a Qwen3-32B or 235B-A22B teacher by KL divergence."
    },
    {
      "term": "GPQA Diamond",
      "definition": "A 198-question set of graduate-level science problems written to be resistant to web search. The most commonly quoted knowledge benchmark for 2025–26 frontier models."
    },
    {
      "term": "MMLU-Pro",
      "definition": "A harder, ten-choice successor to MMLU with more reasoning-heavy questions. Scores are typically 10–20 points below MMLU for the same model, so the two are not interchangeable in a comparison table."
    },
    {
      "term": "SWE-bench Verified",
      "definition": "A 500-issue human-validated subset of SWE-bench measuring whether a model can resolve a real GitHub issue end to end. The benchmark where distilled tiers lose the most ground."
    },
    {
      "term": "Time to first token (TTFT)",
      "definition": "Latency from request to first streamed token. On reasoning models it now includes thinking time, so a \"fast\" model at high effort can measure slower than a \"slow\" model at low effort."
    },
    {
      "term": "Blended price",
      "definition": "A single price per million tokens computed as 0.75 x input + 0.25 x output, the 3:1 weighting Artificial Analysis uses. Handy for ranking, misleading for workloads with unusual input:output ratios."
    },
    {
      "term": "Prompt caching",
      "definition": "Charging a reduced rate for repeated prefix tokens. Cache reads cost 10% of base input on most Claude models and 2.5% on Fable 5.1; OpenAI caches at roughly 10% of input. Often a larger saving than switching model tiers."
    },
    {
      "term": "Anti-distillation clause",
      "definition": "Contract language barring customers from using a vendor’s outputs to train a competing model. Anthropic’s Commercial Terms D.4 and OpenAI’s Services Agreement both contain one; MIT and Apache 2.0 open-weight models do not."
    },
    {
      "term": "Model tiering / routing",
      "definition": "Sending most traffic to a cheap model and escalating only hard requests to a frontier tier. The dominant 2026 cost-control pattern and the main practical way distillation shows up on a buyer’s invoice."
    },
    {
      "term": "Effective parameters",
      "definition": "For MoE and Matformer-style models, the parameters actually activated per token (e.g. Gemma 4 26B A4B activates 3.8B of 25.2B; DeepSeek-V4-Flash activates 13B of 284B). Determines serving cost far more than total parameter count."
    }
  ],
  "sources": [
    {
      "title": "OpenAI API pricing",
      "url": "https://developers.openai.com/api/docs/pricing",
      "publisher": "OpenAI",
      "date": "2026-09",
      "type": "pricing"
    },
    {
      "title": "OpenAI API models reference",
      "url": "https://developers.openai.com/api/docs/models",
      "publisher": "OpenAI",
      "date": "2026-09",
      "type": "docs"
    },
    {
      "title": "GPT-4o mini: advancing cost-efficient intelligence",
      "url": "https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/",
      "publisher": "OpenAI",
      "date": "2024-07-18",
      "type": "blog"
    },
    {
      "title": "OpenAI GPT-5 System Card",
      "url": "https://arxiv.org/pdf/2601.03267",
      "publisher": "OpenAI / arXiv",
      "date": "2026-01",
      "type": "paper"
    },
    {
      "title": "OpenAI Services Agreement",
      "url": "https://openai.com/policies/services-agreement/",
      "publisher": "OpenAI",
      "date": "2026-01-01",
      "type": "law"
    },
    {
      "title": "GPT-5.6 Sol model page",
      "url": "https://openrouter.ai/openai/gpt-5.6-sol",
      "publisher": "OpenRouter",
      "date": "2026-07-09",
      "type": "docs"
    },
    {
      "title": "GPT-5.6 Terra model page",
      "url": "https://openrouter.ai/openai/gpt-5.6-terra",
      "publisher": "OpenRouter",
      "date": "2026-07-09",
      "type": "docs"
    },
    {
      "title": "GPT-5.6 Luna model page",
      "url": "https://openrouter.ai/openai/gpt-5.6-luna",
      "publisher": "OpenRouter",
      "date": "2026-07-09",
      "type": "docs"
    },
    {
      "title": "GPT-5 mini model page",
      "url": "https://openrouter.ai/openai/gpt-5-mini",
      "publisher": "OpenRouter",
      "date": "2025-08-07",
      "type": "docs"
    },
    {
      "title": "GPT-5 nano model page",
      "url": "https://openrouter.ai/openai/gpt-5-nano",
      "publisher": "OpenRouter",
      "date": "2025-08-07",
      "type": "docs"
    },
    {
      "title": "GPT-4o model page",
      "url": "https://openrouter.ai/openai/gpt-4o",
      "publisher": "OpenRouter",
      "date": "2024-05-13",
      "type": "docs"
    },
    {
      "title": "o4-mini model page",
      "url": "https://openrouter.ai/openai/o4-mini",
      "publisher": "OpenRouter",
      "date": "2025-04-16",
      "type": "docs"
    },
    {
      "title": "OpenAI ships GPT-5.4 mini and nano",
      "url": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/",
      "publisher": "The Decoder",
      "date": "2026-02",
      "type": "news"
    },
    {
      "title": "o4-mini: tests, features, o3 comparison, benchmarks",
      "url": "https://www.datacamp.com/blog/o4-mini",
      "publisher": "DataCamp",
      "date": "2025-04",
      "type": "news"
    },
    {
      "title": "gpt-oss-20b model card",
      "url": "https://huggingface.co/openai/gpt-oss-20b",
      "publisher": "OpenAI / Hugging Face",
      "date": "2025-08",
      "type": "docs"
    },
    {
      "title": "Claude platform pricing",
      "url": "https://platform.claude.com/docs/en/about-claude/pricing",
      "publisher": "Anthropic",
      "date": "2026-09",
      "type": "pricing"
    },
    {
      "title": "Claude models overview",
      "url": "https://platform.claude.com/docs/en/about-claude/models/overview",
      "publisher": "Anthropic",
      "date": "2026-09",
      "type": "docs"
    },
    {
      "title": "Claude Opus 5 model page",
      "url": "https://platform.claude.com/docs/en/models/opus-5/overview",
      "publisher": "Anthropic",
      "date": "2026-07-24",
      "type": "docs"
    },
    {
      "title": "Introducing Claude Sonnet 5",
      "url": "https://www.anthropic.com/news/claude-sonnet-5",
      "publisher": "Anthropic",
      "date": "2026-06-30",
      "type": "blog"
    },
    {
      "title": "Introducing Claude Haiku 4.5",
      "url": "https://www.anthropic.com/news/claude-haiku-4-5",
      "publisher": "Anthropic",
      "date": "2025-10-15",
      "type": "blog"
    },
    {
      "title": "Anthropic Commercial Terms of Service",
      "url": "https://www.anthropic.com/legal/commercial-terms",
      "publisher": "Anthropic",
      "date": "2025-06-17",
      "type": "law"
    },
    {
      "title": "Claude Opus 5 by Anthropic: benchmarks and pricing",
      "url": "https://datanorth.ai/news/claude-opus-5-by-anthropic",
      "publisher": "DataNorth",
      "date": "2026-07",
      "type": "news"
    },
    {
      "title": "Claude benchmarks 2026",
      "url": "https://www.morphllm.com/claude-benchmarks",
      "publisher": "Morph",
      "date": "2026-09",
      "type": "news"
    },
    {
      "title": "Anthropic launches Opus 5",
      "url": "https://techcrunch.com/2026/07/24/anthropic-launches-opus-5/",
      "publisher": "TechCrunch",
      "date": "2026-07-24",
      "type": "news"
    },
    {
      "title": "Gemini API pricing",
      "url": "https://ai.google.dev/gemini-api/docs/pricing",
      "publisher": "Google",
      "date": "2026-09",
      "type": "pricing"
    },
    {
      "title": "Gemini API models",
      "url": "https://ai.google.dev/gemini-api/docs/models",
      "publisher": "Google",
      "date": "2026-09",
      "type": "docs"
    },
    {
      "title": "Gemini 3.1 Pro model page",
      "url": "https://deepmind.google/models/gemini/pro/",
      "publisher": "Google DeepMind",
      "date": "2026-02",
      "type": "docs"
    },
    {
      "title": "Gemini Flash model page",
      "url": "https://deepmind.google/models/gemini/flash/",
      "publisher": "Google DeepMind",
      "date": "2026-09-02",
      "type": "docs"
    },
    {
      "title": "Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context and Next Generation Agentic Capabilities",
      "url": "https://arxiv.org/html/2507.06261v1/",
      "publisher": "Google DeepMind / arXiv",
      "date": "2025-07",
      "type": "paper"
    },
    {
      "title": "Gemma 4 model card",
      "url": "https://ai.google.dev/gemma/docs/core/model_card_4",
      "publisher": "Google",
      "date": "2026-04-02",
      "type": "docs"
    },
    {
      "title": "Gemma 4",
      "url": "https://deepmind.google/models/gemma/gemma-4/",
      "publisher": "Google DeepMind",
      "date": "2026-04-02",
      "type": "blog"
    },
    {
      "title": "gemma-3-27b-it model card",
      "url": "https://huggingface.co/google/gemma-3-27b-it",
      "publisher": "Google / Hugging Face",
      "date": "2025-03",
      "type": "docs"
    },
    {
      "title": "Gemini 3.8 Flash rolling out three weeks after last release",
      "url": "https://9to5google.com/2026/09/02/gemini-3-8-flash-launch/",
      "publisher": "9to5Google",
      "date": "2026-09-02",
      "type": "news"
    },
    {
      "title": "Gemini 3.8 Flash model stats",
      "url": "https://llm-stats.com/models/gemini-3.8-flash",
      "publisher": "LLM-Stats",
      "date": "2026-09-02",
      "type": "news"
    },
    {
      "title": "Gemini 3.1 Flash-Lite benchmark results",
      "url": "https://layerlens.ai/blog/gemini-3-1-flash-lite-benchmark-results-efficiency-model-comparison",
      "publisher": "LayerLens",
      "date": "2026-03",
      "type": "news"
    },
    {
      "title": "DeepSeek API pricing",
      "url": "https://api-docs.deepseek.com/quick_start/pricing/",
      "publisher": "DeepSeek",
      "date": "2026-08",
      "type": "pricing"
    },
    {
      "title": "DeepSeek-R1 model card (with distilled model evaluations)",
      "url": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
      "publisher": "DeepSeek / Hugging Face",
      "date": "2025-01-20",
      "type": "docs"
    },
    {
      "title": "DeepSeek-R1-0528 model card",
      "url": "https://huggingface.co/deepseek-ai/DeepSeek-R1-0528",
      "publisher": "DeepSeek / Hugging Face",
      "date": "2025-05-28",
      "type": "docs"
    },
    {
      "title": "DeepSeek-R1-Distill-Qwen-32B model card",
      "url": "https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B",
      "publisher": "DeepSeek / Hugging Face",
      "date": "2025-01-20",
      "type": "docs"
    },
    {
      "title": "DeepSeek-V4-Pro model card",
      "url": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro",
      "publisher": "DeepSeek / Hugging Face",
      "date": "2026-04-26",
      "type": "docs"
    },
    {
      "title": "DeepSeek-V4-Flash model card",
      "url": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
      "publisher": "DeepSeek / Hugging Face",
      "date": "2026-07",
      "type": "docs"
    },
    {
      "title": "DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models",
      "url": "https://arxiv.org/html/2512.02556",
      "publisher": "DeepSeek / arXiv",
      "date": "2025-12",
      "type": "paper"
    },
    {
      "title": "DeepSeek raises some V4 prices by more than 10x",
      "url": "https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html",
      "publisher": "InfoWorld",
      "date": "2026-08",
      "type": "news"
    },
    {
      "title": "Llama 4 model card",
      "url": "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md",
      "publisher": "Meta",
      "date": "2025-04",
      "type": "docs"
    },
    {
      "title": "Llama 3.3 model card",
      "url": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md",
      "publisher": "Meta",
      "date": "2024-12-06",
      "type": "docs"
    },
    {
      "title": "Llama 3.2 model card",
      "url": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md",
      "publisher": "Meta",
      "date": "2024-09-25",
      "type": "docs"
    },
    {
      "title": "Qwen3 Technical Report",
      "url": "https://arxiv.org/pdf/2505.09388",
      "publisher": "Alibaba Qwen Team / arXiv",
      "date": "2025-05",
      "type": "paper"
    },
    {
      "title": "Qwen3.8-27B model card",
      "url": "https://huggingface.co/Qwen/Qwen3.8-27B",
      "publisher": "Alibaba / Hugging Face",
      "date": "2026-08",
      "type": "docs"
    },
    {
      "title": "Qwen3-4B-Instruct-2507 model card",
      "url": "https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507",
      "publisher": "Alibaba / Hugging Face",
      "date": "2025-07",
      "type": "docs"
    },
    {
      "title": "Alibaba Model Studio model pricing",
      "url": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
      "publisher": "Alibaba Cloud",
      "date": "2026-09",
      "type": "pricing"
    },
    {
      "title": "phi-4 model card",
      "url": "https://huggingface.co/microsoft/phi-4",
      "publisher": "Microsoft / Hugging Face",
      "date": "2024-12",
      "type": "docs"
    },
    {
      "title": "Phi-4-mini-instruct model card",
      "url": "https://huggingface.co/microsoft/Phi-4-mini-instruct",
      "publisher": "Microsoft / Hugging Face",
      "date": "2025-02",
      "type": "docs"
    },
    {
      "title": "SmolLM3-3B model card",
      "url": "https://huggingface.co/HuggingFaceTB/SmolLM3-3B",
      "publisher": "Hugging Face",
      "date": "2025-07",
      "type": "docs"
    },
    {
      "title": "Mistral AI API pricing",
      "url": "https://mistral.ai/pricing/api",
      "publisher": "Mistral AI",
      "date": "2026-09",
      "type": "pricing"
    },
    {
      "title": "Mistral models overview",
      "url": "https://docs.mistral.ai/getting-started/models/models_overview/",
      "publisher": "Mistral AI",
      "date": "2026-09",
      "type": "docs"
    },
    {
      "title": "Ministral-3-14B-Instruct-2512 model card",
      "url": "https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512",
      "publisher": "Mistral AI / Hugging Face",
      "date": "2025-12-02",
      "type": "docs"
    },
    {
      "title": "Mistral Small 4 model page",
      "url": "https://openrouter.ai/mistralai/mistral-small-2603",
      "publisher": "OpenRouter",
      "date": "2026-03-16",
      "type": "docs"
    },
    {
      "title": "What is Amazon Nova? (teacher/student distillation matrix)",
      "url": "https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html",
      "publisher": "Amazon Web Services",
      "date": "2026",
      "type": "docs"
    },
    {
      "title": "What’s new in Amazon Nova 2",
      "url": "https://docs.aws.amazon.com/nova/latest/nova2-userguide/whats-new.html",
      "publisher": "Amazon Web Services",
      "date": "2025-12",
      "type": "docs"
    },
    {
      "title": "The Amazon Nova Family of Models: Technical Report and Model Card",
      "url": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf",
      "publisher": "Amazon Science",
      "date": "2025-03-17",
      "type": "paper"
    },
    {
      "title": "Amazon Bedrock Model Distillation",
      "url": "https://aws.amazon.com/bedrock/model-distillation/",
      "publisher": "Amazon Web Services",
      "date": "2025-05",
      "type": "docs"
    },
    {
      "title": "Amazon Bedrock pricing",
      "url": "https://aws.amazon.com/bedrock/pricing/",
      "publisher": "Amazon Web Services",
      "date": "2026-09",
      "type": "pricing"
    },
    {
      "title": "Nova Micro API pricing",
      "url": "https://pricepertoken.com/pricing-page/model/amazon-nova-micro-v1",
      "publisher": "PricePerToken",
      "date": "2026",
      "type": "pricing"
    },
    {
      "title": "Nova Lite API pricing",
      "url": "https://pricepertoken.com/pricing-page/model/amazon-nova-lite-v1",
      "publisher": "PricePerToken",
      "date": "2026",
      "type": "pricing"
    },
    {
      "title": "Nova 2 Lite API pricing",
      "url": "https://pricepertoken.com/pricing-page/model/amazon-nova-2-lite-v1",
      "publisher": "PricePerToken",
      "date": "2026",
      "type": "pricing"
    },
    {
      "title": "Amazon Nova Pro: AWS Bedrock model guide, specs and pricing (2026)",
      "url": "https://ucstrategies.com/news/amazon-nova-pro-aws-bedrock-model-guide-specs-pricing-2026/",
      "publisher": "UC Strategies",
      "date": "2026",
      "type": "news"
    },
    {
      "title": "Together AI pricing",
      "url": "https://www.together.ai/pricing",
      "publisher": "Together AI",
      "date": "2026-09",
      "type": "pricing"
    },
    {
      "title": "Groq supported models and pricing",
      "url": "https://console.groq.com/docs/models",
      "publisher": "Groq",
      "date": "2026-09",
      "type": "pricing"
    },
    {
      "title": "xAI models and pricing",
      "url": "https://docs.x.ai/docs/models",
      "publisher": "xAI",
      "date": "2026-08",
      "type": "pricing"
    },
    {
      "title": "GPT-5.6 Luna (low) vs Claude 4.5 Haiku (reasoning)",
      "url": "https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning",
      "publisher": "Artificial Analysis",
      "date": "2026-09",
      "type": "news"
    },
    {
      "title": "Gemini 3.5 Flash-Lite vs GPT-5.6 Luna (high)",
      "url": "https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gpt-5-6-luna-high",
      "publisher": "Artificial Analysis",
      "date": "2026-09",
      "type": "news"
    },
    {
      "title": "Claude Sonnet 5 vs Claude Opus 5",
      "url": "https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5",
      "publisher": "Artificial Analysis",
      "date": "2026-09",
      "type": "news"
    },
    {
      "title": "DeepSeek V4 Flash vs DeepSeek V4 Pro",
      "url": "https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-deepseek-v4-pro",
      "publisher": "Artificial Analysis",
      "date": "2026-09",
      "type": "news"
    },
    {
      "title": "gpt-oss-120B vs Llama 4 Maverick",
      "url": "https://artificialanalysis.ai/models/comparisons/gpt-oss-120b-vs-llama-4-maverick",
      "publisher": "Artificial Analysis",
      "date": "2026-09",
      "type": "news"
    },
    {
      "title": "Comparison of AI models across intelligence, performance and price",
      "url": "https://artificialanalysis.ai/models",
      "publisher": "Artificial Analysis",
      "date": "2026-09",
      "type": "news"
    },
    {
      "title": "Enterprise AI costs hit 2026 low driven by price wars and Chinese open-source models",
      "url": "https://www.scmp.com/tech/tech-trends/article/3363549/enterprise-ai-costs-hit-2026-low-driven-price-wars-chinese-open-source-models-research",
      "publisher": "South China Morning Post",
      "date": "2026-08",
      "type": "news"
    },
    {
      "title": "OpenAI discounts GPT-5.6 Luna and Terra",
      "url": "https://www.axios.com/2026/07/30/openai-cuts-prices-gpt-terra-luna5",
      "publisher": "Axios",
      "date": "2026-07-30",
      "type": "news"
    },
    {
      "title": "The Llama 4 herd: the beginning of a new era of natively multimodal AI innovation",
      "url": "https://ai.meta.com/blog/llama-4-multimodal-intelligence/",
      "publisher": "Meta AI",
      "date": "2025-04-05",
      "type": "blog"
    }
  ],
  "extras": {
    "models": [
      {
        "model": "GPT-6 Astra",
        "vendor": "OpenAI",
        "family": "GPT-6",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 10,
        "output_per_mtok_usd": 50,
        "blended_per_mtok_usd": 20,
        "cost_per_1m_requests_usd": 9000,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1050,
        "maxOutputK": 128,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2026",
        "note": "Top of the OpenAI stack; benchmark table not published at the time of writing.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": null,
        "latencySource": null
      },
      {
        "model": "GPT-5.6 Sol",
        "vendor": "OpenAI",
        "family": "GPT-5.6",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 4,
        "output_per_mtok_usd": 20,
        "blended_per_mtok_usd": 8,
        "cost_per_1m_requests_usd": 3600,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 92.4,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1050,
        "maxOutputK": 128,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2026-07-09",
        "note": "Largest GPT-5.6 tier. List price cut ~20% on 2026-08-21 for three months.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://openrouter.ai/openai/gpt-5.6-sol",
        "latencySource": null
      },
      {
        "model": "GPT-5.6 Terra",
        "vendor": "OpenAI",
        "family": "GPT-5.6",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 2,
        "output_per_mtok_usd": 12,
        "blended_per_mtok_usd": 4.5,
        "cost_per_1m_requests_usd": 2000,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 88.4,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1050,
        "maxOutputK": 128,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2026-07-09",
        "note": "Mid tier of the GPT-5.6 family; OpenAI has not documented how it relates to Sol. TAU-Bench 72.0%.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://openrouter.ai/openai/gpt-5.6-terra",
        "latencySource": null
      },
      {
        "model": "GPT-5.6 Luna",
        "vendor": "OpenAI",
        "family": "GPT-5.6",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.2,
        "output_per_mtok_usd": 1.2,
        "blended_per_mtok_usd": 0.45,
        "cost_per_1m_requests_usd": 200,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 87,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": 1700,
        "contextK": 1050,
        "maxOutputK": 128,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2026-07-09",
        "note": "The single best price/quality point in the closed-model market as of Sep 2026: roughly 94% of Sol GPQA at 5% of the input price. Price cut 80% on 2026-07-30. TTFT measured at low reasoning effort. The 87.0 GPQA Diamond figure is OpenRouter's auto-routing measurement; provider-specific scores on the same page range 87.0–89.9 (Azure US 89.9, Azure EU 89.1, Bedrock US 88.6, Azure 88.2), and no OpenAI primary source for it is currently reachable.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://openrouter.ai/openai/gpt-5.6-luna",
        "latencySource": "https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning"
      },
      {
        "model": "GPT-5.5",
        "vendor": "OpenAI",
        "family": "GPT-5.5",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 5,
        "output_per_mtok_usd": 30,
        "blended_per_mtok_usd": 11.25,
        "cost_per_1m_requests_usd": 5000,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": null,
        "maxOutputK": null,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2026-04",
        "note": "Superseded by GPT-5.6 for most buyers; still priced at the old frontier rate. Context window undisclosed here: the 272K figure previously shown is OpenAI's long-context pricing threshold, not a window size, and no primary source in this dataset states GPT-5.5's window.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": null,
        "latencySource": null
      },
      {
        "model": "GPT-5.4",
        "vendor": "OpenAI",
        "family": "GPT-5.4",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 2.5,
        "output_per_mtok_usd": 15,
        "blended_per_mtok_usd": 5.625,
        "cost_per_1m_requests_usd": 2500,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 93,
        "humaneval_or_swe": 57.7,
        "coding_benchmark": "SWE-bench Pro",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1050,
        "maxOutputK": null,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2026-02",
        "note": "Teacher tier for the 5.4 mini/nano pair. 1,050,000-token context window / 128,000 max output tokens per OpenAI's model page.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/",
        "latencySource": null,
        "contextSource": "https://developers.openai.com/api/docs/models/gpt-5.4"
      },
      {
        "model": "GPT-5.4 mini",
        "vendor": "OpenAI",
        "family": "GPT-5.4",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.75,
        "output_per_mtok_usd": 4.5,
        "blended_per_mtok_usd": 1.6875,
        "cost_per_1m_requests_usd": 750,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 88,
        "humaneval_or_swe": 54.4,
        "coding_benchmark": "SWE-bench Pro",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 400,
        "maxOutputK": null,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2026-02",
        "note": "Retains 94.6% of GPT-5.4 GPQA at 30% of the input price. 400K context window.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/",
        "latencySource": null,
        "contextSource": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/"
      },
      {
        "model": "GPT-5.4 nano",
        "vendor": "OpenAI",
        "family": "GPT-5.4",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.2,
        "output_per_mtok_usd": 1.25,
        "blended_per_mtok_usd": 0.4625,
        "cost_per_1m_requests_usd": 205,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 82.8,
        "humaneval_or_swe": 52.4,
        "coding_benchmark": "SWE-bench Pro",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 400,
        "maxOutputK": null,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2026-02",
        "note": "Cheapest OpenAI tier that still clears 80% GPQA. 400K context window.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/",
        "latencySource": null,
        "contextSource": "https://the-decoder.com/openai-ships-gpt-5-4-mini-and-nano-faster-and-more-capable-but-up-to-4x-pricier/"
      },
      {
        "model": "GPT-5",
        "vendor": "OpenAI",
        "family": "GPT-5",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 1.25,
        "output_per_mtok_usd": 10,
        "blended_per_mtok_usd": 3.4375,
        "cost_per_1m_requests_usd": 1500,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": 74.9,
        "coding_benchmark": "SWE-bench Verified",
        "aime": 94.6,
        "latency_ttft_ms": null,
        "contextK": 400,
        "maxOutputK": 128,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2025-08-07",
        "note": "Now a value tier rather than a frontier tier; still the default in a lot of 2025-era enterprise code.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://arxiv.org/pdf/2601.03267",
        "latencySource": null
      },
      {
        "model": "GPT-5 mini",
        "vendor": "OpenAI",
        "family": "GPT-5",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.25,
        "output_per_mtok_usd": 2,
        "blended_per_mtok_usd": 0.6875,
        "cost_per_1m_requests_usd": 300,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 80.3,
        "humaneval_or_swe": 45.7,
        "coding_benchmark": "SWE-bench Pro",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 400,
        "maxOutputK": 128,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2025-08-07",
        "note": "TAU-Bench 75.3%. Widely used as the default cheap tier in 2025-26 agent stacks.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://openrouter.ai/openai/gpt-5-mini",
        "latencySource": null
      },
      {
        "model": "GPT-5 nano",
        "vendor": "OpenAI",
        "family": "GPT-5",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.05,
        "output_per_mtok_usd": 0.4,
        "blended_per_mtok_usd": 0.1375,
        "cost_per_1m_requests_usd": 60,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 70.9,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 400,
        "maxOutputK": 128,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2025-08-07",
        "note": "Cheapest first-party OpenAI text model. TAU-Bench 52.0%.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://openrouter.ai/openai/gpt-5-nano",
        "latencySource": null
      },
      {
        "model": "GPT-4o",
        "vendor": "OpenAI",
        "family": "GPT-4o",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 2.5,
        "output_per_mtok_usd": 10,
        "blended_per_mtok_usd": 4.375,
        "cost_per_1m_requests_usd": 2000,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": 16,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2024-05-13",
        "note": "Legacy. Same list price as it had in 2024 while newer tiers undercut it 10x — a common source of silent overspend.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://openrouter.ai/openai/gpt-4o",
        "latencySource": null
      },
      {
        "model": "GPT-4o mini",
        "vendor": "OpenAI",
        "family": "GPT-4o",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.15,
        "output_per_mtok_usd": 0.6,
        "blended_per_mtok_usd": 0.2625,
        "cost_per_1m_requests_usd": 120,
        "mmlu": 82,
        "mmlu_variant": "MMLU",
        "gpqa": null,
        "humaneval_or_swe": 87.2,
        "coding_benchmark": "HumanEval",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": 16,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2024-07-18",
        "note": "The model that made \"cheap tier\" a standard product line. MMMU 59.4%. OpenAI has never said how it was built.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/",
        "latencySource": null
      },
      {
        "model": "o3",
        "vendor": "OpenAI",
        "family": "o-series",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 2,
        "output_per_mtok_usd": 8,
        "blended_per_mtok_usd": 3.5,
        "cost_per_1m_requests_usd": 1600,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 83.3,
        "humaneval_or_swe": 69.1,
        "coding_benchmark": "SWE-bench Verified",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 200,
        "maxOutputK": 100,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2025-04-16",
        "note": "Reasoning teacher of the 2025 generation.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://www.datacamp.com/blog/o4-mini",
        "latencySource": null
      },
      {
        "model": "o4-mini",
        "vendor": "OpenAI",
        "family": "o-series",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 1.1,
        "output_per_mtok_usd": 4.4,
        "blended_per_mtok_usd": 1.925,
        "cost_per_1m_requests_usd": 880,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 81.4,
        "humaneval_or_swe": 68.1,
        "coding_benchmark": "SWE-bench Verified",
        "aime": 88.9,
        "latency_ttft_ms": null,
        "contextK": 200,
        "maxOutputK": 100,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "OpenAI API",
        "releaseDate": "2025-04-16",
        "note": "97.7% of o3 GPQA Diamond (81.4 vs 83.3) and 98.6% of its SWE-bench Verified (68.1 vs 69.1) at just over half the price — but now beaten on both axes by GPT-5.6 Luna.",
        "source": "https://developers.openai.com/api/docs/pricing",
        "benchmarkSource": "https://www.datacamp.com/blog/o4-mini",
        "latencySource": null
      },
      {
        "model": "gpt-oss-120b",
        "vendor": "OpenAI (open weights)",
        "family": "gpt-oss",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 116.8,
        "active_params_b": 5.1,
        "input_per_mtok_usd": 0.15,
        "output_per_mtok_usd": 0.6,
        "blended_per_mtok_usd": 0.2625,
        "cost_per_1m_requests_usd": 120,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": 850,
        "contextK": 128,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Groq / Together / self-host",
        "releaseDate": "2025-08",
        "note": "MoE, MXFP4 quantised — runs on one 80GB GPU. The default \"we must own the weights\" option for US buyers.",
        "source": "https://console.groq.com/docs/models",
        "benchmarkSource": "https://huggingface.co/openai/gpt-oss-20b",
        "latencySource": "https://artificialanalysis.ai/models/comparisons/gpt-oss-120b-vs-llama-4-maverick"
      },
      {
        "model": "gpt-oss-20b",
        "vendor": "OpenAI (open weights)",
        "family": "gpt-oss",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 20.9,
        "active_params_b": 3.6,
        "input_per_mtok_usd": 0.075,
        "output_per_mtok_usd": 0.3,
        "blended_per_mtok_usd": 0.1312,
        "cost_per_1m_requests_usd": 60,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 58.59,
        "humaneval_or_swe": 53.2,
        "coding_benchmark": "SWE-bench Verified",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Groq / self-host (16GB)",
        "releaseDate": "2025-08",
        "note": "Fits in 16GB. Groq serves it at ~1000 tok/s.",
        "source": "https://console.groq.com/docs/models",
        "benchmarkSource": "https://huggingface.co/openai/gpt-oss-20b",
        "latencySource": null
      },
      {
        "model": "Claude Fable 5.1",
        "vendor": "Anthropic",
        "family": "Claude 5",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 10,
        "output_per_mtok_usd": 50,
        "blended_per_mtok_usd": 20,
        "cost_per_1m_requests_usd": 9000,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1000,
        "maxOutputK": 128,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Claude API / Bedrock / Vertex / Foundry",
        "releaseDate": "2026",
        "note": "Cache hits cost 2.5% of base input (vs 10% elsewhere) — the cheapest long-context re-read in the market.",
        "source": "https://platform.claude.com/docs/en/about-claude/pricing",
        "benchmarkSource": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "latencySource": null
      },
      {
        "model": "Claude Opus 5",
        "vendor": "Anthropic",
        "family": "Claude 5",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 5,
        "output_per_mtok_usd": 25,
        "blended_per_mtok_usd": 10,
        "cost_per_1m_requests_usd": 4500,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": 96,
        "coding_benchmark": "SWE-bench Verified",
        "aime": null,
        "latency_ttft_ms": 77250,
        "contextK": 1000,
        "maxOutputK": 128,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Claude API / Bedrock / Vertex / Foundry",
        "releaseDate": "2026-07-24",
        "note": "SWE-bench Pro 79.2%, OSWorld 2.0 70.6%, Frontier-Bench 43.3%. Latency shown is Artificial Analysis TTFT at max effort (adaptive thinking inflates it).",
        "source": "https://platform.claude.com/docs/en/about-claude/pricing",
        "benchmarkSource": "https://datanorth.ai/news/claude-opus-5-by-anthropic",
        "latencySource": "https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5"
      },
      {
        "model": "Claude Sonnet 5",
        "vendor": "Anthropic",
        "family": "Claude 5",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 2,
        "output_per_mtok_usd": 10,
        "blended_per_mtok_usd": 4,
        "cost_per_1m_requests_usd": 1800,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": 85.2,
        "coding_benchmark": "SWE-bench Verified",
        "aime": null,
        "latency_ttft_ms": 177770,
        "contextK": 1000,
        "maxOutputK": 128,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Claude API / Bedrock / Vertex / Foundry",
        "releaseDate": "2026-06-30",
        "note": "Terminal-Bench 2.1 80.4%, SWE-bench Pro 63.2%. Introductory $2/$10 made permanent on 2026-08-11; the planned rise to $3/$15 was cancelled.",
        "source": "https://platform.claude.com/docs/en/about-claude/pricing",
        "benchmarkSource": "https://www.morphllm.com/claude-benchmarks",
        "latencySource": "https://artificialanalysis.ai/models/comparisons/claude-sonnet-5-vs-claude-opus-5"
      },
      {
        "model": "Claude Sonnet 4.6",
        "vendor": "Anthropic",
        "family": "Claude 4.6",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 3,
        "output_per_mtok_usd": 15,
        "blended_per_mtok_usd": 6,
        "cost_per_1m_requests_usd": 2700,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1000,
        "maxOutputK": 128,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Claude API / Bedrock / Vertex / Foundry",
        "releaseDate": "2026-02",
        "note": "Costs 50% more than Sonnet 5 and scores lower — migrate.",
        "source": "https://platform.claude.com/docs/en/about-claude/pricing",
        "benchmarkSource": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "latencySource": null
      },
      {
        "model": "Claude Haiku 4.5",
        "vendor": "Anthropic",
        "family": "Claude 4.5",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 1,
        "output_per_mtok_usd": 5,
        "blended_per_mtok_usd": 2,
        "cost_per_1m_requests_usd": 900,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": 73.3,
        "coding_benchmark": "SWE-bench Verified",
        "aime": null,
        "latency_ttft_ms": 19920,
        "contextK": 200,
        "maxOutputK": 64,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Claude API / Bedrock / Vertex / Foundry",
        "releaseDate": "2025-10-15",
        "note": "Anthropic positions it as \"similar coding performance to Sonnet 4 at one-third the cost and twice the speed\". Only 200K context. Latency is AA TTFT with reasoning on.",
        "source": "https://platform.claude.com/docs/en/about-claude/pricing",
        "benchmarkSource": "https://www.anthropic.com/news/claude-haiku-4-5",
        "latencySource": "https://artificialanalysis.ai/models/comparisons/gpt-5-6-luna-low-vs-claude-4-5-haiku-reasoning"
      },
      {
        "model": "Claude Haiku 3.5",
        "vendor": "Anthropic",
        "family": "Claude 3.5",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.8,
        "output_per_mtok_usd": 4,
        "blended_per_mtok_usd": 1.6,
        "cost_per_1m_requests_usd": 720,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 200,
        "maxOutputK": 8,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Bedrock / Google Cloud only (retired on first-party API)",
        "releaseDate": "2024-11",
        "note": "Retired except on Bedrock and Google Cloud — do not start new builds here.",
        "source": "https://platform.claude.com/docs/en/about-claude/pricing",
        "benchmarkSource": "https://platform.claude.com/docs/en/about-claude/pricing",
        "latencySource": null
      },
      {
        "model": "Gemini 3.1 Pro (Preview)",
        "vendor": "Google",
        "family": "Gemini 3",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 2,
        "output_per_mtok_usd": 12,
        "blended_per_mtok_usd": 4.5,
        "cost_per_1m_requests_usd": 2000,
        "mmlu": 92.6,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 94.3,
        "humaneval_or_swe": 80.6,
        "coding_benchmark": "SWE-bench Verified",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1000,
        "maxOutputK": 64,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Gemini API / Vertex AI",
        "releaseDate": "2026-02",
        "note": "Highest published MMLU-Pro and GPQA Diamond of any model in this table. Terminal-Bench 2.0 68.5%, HLE 44.4%. Price doubles above 200K input tokens.",
        "source": "https://ai.google.dev/gemini-api/docs/pricing",
        "benchmarkSource": "https://deepmind.google/models/gemini/pro/",
        "latencySource": null
      },
      {
        "model": "Gemini 3.8 Flash",
        "vendor": "Google",
        "family": "Gemini 3",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.75,
        "output_per_mtok_usd": 3.75,
        "blended_per_mtok_usd": 1.5,
        "cost_per_1m_requests_usd": 675,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": 73.7,
        "coding_benchmark": "DeepSWE v1.1",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1000,
        "maxOutputK": 66,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Gemini API / Vertex AI",
        "releaseDate": "2026-09-02",
        "note": "Newest model in this table (two days old). Terminal-Bench 2.1 89.4%, HLE-Verified 54.9%. Promo price runs through 2026-12-31.",
        "source": "https://ai.google.dev/gemini-api/docs/pricing",
        "benchmarkSource": "https://deepmind.google/models/gemini/flash/",
        "latencySource": null
      },
      {
        "model": "Gemini 3.5 Flash",
        "vendor": "Google",
        "family": "Gemini 3",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 1.5,
        "output_per_mtok_usd": 9,
        "blended_per_mtok_usd": 3.375,
        "cost_per_1m_requests_usd": 1500,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1000,
        "maxOutputK": 66,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Gemini API / Vertex AI",
        "releaseDate": "2026-06",
        "note": "Costs 2x Gemini 3.8 Flash for older capability — a live migration target.",
        "source": "https://ai.google.dev/gemini-api/docs/pricing",
        "benchmarkSource": "https://ai.google.dev/gemini-api/docs/models",
        "latencySource": null
      },
      {
        "model": "Gemini 3.5 Flash-Lite",
        "vendor": "Google",
        "family": "Gemini 3",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.3,
        "output_per_mtok_usd": 2.5,
        "blended_per_mtok_usd": 0.85,
        "cost_per_1m_requests_usd": 370,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": 6480,
        "contextK": 1000,
        "maxOutputK": 64,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Gemini API / Vertex AI",
        "releaseDate": "2026-06",
        "note": "391 output tok/s in Artificial Analysis testing — the fastest generally-available hosted model in this table.",
        "source": "https://ai.google.dev/gemini-api/docs/pricing",
        "benchmarkSource": "https://ai.google.dev/gemini-api/docs/models",
        "latencySource": "https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gpt-5-6-luna-high"
      },
      {
        "model": "Gemini 3.1 Flash-Lite",
        "vendor": "Google",
        "family": "Gemini 3",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.25,
        "output_per_mtok_usd": 1.5,
        "blended_per_mtok_usd": 0.5625,
        "cost_per_1m_requests_usd": 250,
        "mmlu": 83,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 72.2,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1000,
        "maxOutputK": 64,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Gemini API / Vertex AI",
        "releaseDate": "2026-03-03",
        "note": "72.2% GPQA Diamond (198 items) and 83.0% MMLU-Pro at $0.25/MTok input — the strongest quality-per-dollar in the Google line, though well below the frontier tiers on GPQA.",
        "source": "https://ai.google.dev/gemini-api/docs/pricing",
        "benchmarkSource": "https://layerlens.ai/blog/gemini-3-1-flash-lite-benchmark-results-efficiency-model-comparison",
        "latencySource": null
      },
      {
        "model": "Gemini 2.5 Pro",
        "vendor": "Google",
        "family": "Gemini 2.5",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 1.25,
        "output_per_mtok_usd": 10,
        "blended_per_mtok_usd": 3.4375,
        "cost_per_1m_requests_usd": 1500,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 86.4,
        "humaneval_or_swe": 67.2,
        "coding_benchmark": "SWE-bench Verified",
        "aime": 88,
        "latency_ttft_ms": null,
        "contextK": 1000,
        "maxOutputK": 64,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Gemini API / Vertex AI",
        "releaseDate": "2025-03",
        "note": "Documented teacher for the 2.5 Flash line. LiveCodeBench 74.2%.",
        "source": "https://ai.google.dev/gemini-api/docs/pricing",
        "benchmarkSource": "https://arxiv.org/html/2507.06261v1/",
        "latencySource": null
      },
      {
        "model": "Gemini 2.5 Flash",
        "vendor": "Google",
        "family": "Gemini 2.5",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "Gemini 2.5 Pro (k-sparse logit distillation)",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.3,
        "output_per_mtok_usd": 2.5,
        "blended_per_mtok_usd": 0.85,
        "cost_per_1m_requests_usd": 370,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 82.8,
        "humaneval_or_swe": 60.3,
        "coding_benchmark": "SWE-bench Verified",
        "aime": 72,
        "latency_ttft_ms": null,
        "contextK": 1000,
        "maxOutputK": 64,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Gemini API / Vertex AI",
        "releaseDate": "2025-04",
        "note": "The Gemini 2.5 report states the smaller models are distilled, approximating the teacher next-token distribution with a k-sparse distribution. 95.8% of Pro GPQA at 24% of the input price.",
        "source": "https://ai.google.dev/gemini-api/docs/pricing",
        "benchmarkSource": "https://arxiv.org/html/2507.06261v1/",
        "latencySource": null
      },
      {
        "model": "Gemini 2.5 Flash-Lite",
        "vendor": "Google",
        "family": "Gemini 2.5",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "Gemini 2.5 Pro (k-sparse logit distillation)",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.1,
        "output_per_mtok_usd": 0.4,
        "blended_per_mtok_usd": 0.175,
        "cost_per_1m_requests_usd": 80,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": 300,
        "contextK": 1000,
        "maxOutputK": 64,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Gemini API / Vertex AI",
        "releaseDate": "2025-06",
        "note": "0.30s time-to-first-token in non-reasoning mode — the lowest measured in the Artificial Analysis set. Cheapest 1M-context model from a US hyperscaler.",
        "source": "https://ai.google.dev/gemini-api/docs/pricing",
        "benchmarkSource": "https://arxiv.org/html/2507.06261v1/",
        "latencySource": "https://artificialanalysis.ai/models"
      },
      {
        "model": "Gemma 4 31B",
        "vendor": "Google",
        "family": "Gemma 4",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 30.7,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 85.2,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 84.3,
        "humaneval_or_swe": 80,
        "coding_benchmark": "LiveCodeBench v6",
        "aime": 89.2,
        "latency_ttft_ms": null,
        "contextK": 256,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Self-host / Vertex / LM Studio",
        "releaseDate": "2026-04-02",
        "note": "Apache 2.0, 31B dense, MMLU-Pro 85.2 — beats every hosted mid-tier in this table on paper while being fully self-hostable. Built from Gemini 3 research; Google does not call it a distillation.",
        "source": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "benchmarkSource": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "latencySource": null
      },
      {
        "model": "Gemma 4 26B A4B (MoE)",
        "vendor": "Google",
        "family": "Gemma 4",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 25.2,
        "active_params_b": 3.8,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 82.6,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 82.3,
        "humaneval_or_swe": 77.1,
        "coding_benchmark": "LiveCodeBench v6",
        "aime": 88.3,
        "latency_ttft_ms": null,
        "contextK": 256,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Self-host / Vertex",
        "releaseDate": "2026-04-02",
        "note": "3.8B active parameters at 82.3 GPQA — the best FLOP-per-point ratio in the open set.",
        "source": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "benchmarkSource": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "latencySource": null
      },
      {
        "model": "Gemma 4 12B Unified",
        "vendor": "Google",
        "family": "Gemma 4",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 11.95,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 77.2,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 78.8,
        "humaneval_or_swe": 72,
        "coding_benchmark": "LiveCodeBench v6",
        "aime": 77.5,
        "latency_ttft_ms": null,
        "contextK": 256,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Self-host",
        "releaseDate": "2026-04-02",
        "note": "Single-GPU class with 256K context.",
        "source": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "benchmarkSource": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "latencySource": null
      },
      {
        "model": "Gemma 4 E4B",
        "vendor": "Google",
        "family": "Gemma 4",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 4.5,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 69.4,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 58.6,
        "humaneval_or_swe": 52,
        "coding_benchmark": "LiveCodeBench v6",
        "aime": 42.5,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Self-host / on-device",
        "releaseDate": "2026-04-02",
        "note": "On-device tier: 4.5B effective params, MMLU-Pro 69.4 — above GPT-4o mini class quality offline.",
        "source": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "benchmarkSource": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "latencySource": null
      },
      {
        "model": "Gemma 4 E2B",
        "vendor": "Google",
        "family": "Gemma 4",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 2.3,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 60,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 43.4,
        "humaneval_or_swe": 44,
        "coding_benchmark": "LiveCodeBench v6",
        "aime": 37.5,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Self-host / on-device",
        "releaseDate": "2026-04-02",
        "note": "Smallest Gemma 4; phone-class.",
        "source": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "benchmarkSource": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "latencySource": null
      },
      {
        "model": "Gemma 3 27B IT",
        "vendor": "Google",
        "family": "Gemma 3",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 27,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 78.6,
        "mmlu_variant": "MMLU",
        "gpqa": 24.3,
        "humaneval_or_swe": 48.8,
        "coding_benchmark": "HumanEval",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "Gemma Terms of Use (custom)",
        "weights": "open",
        "hosting": "Self-host",
        "releaseDate": "2025-03",
        "note": "Superseded by Gemma 4 on both quality and licence (Gemma 4 is Apache 2.0). Kept here because a lot of 2025 fine-tunes sit on it.",
        "source": "https://huggingface.co/google/gemma-3-27b-it",
        "benchmarkSource": "https://huggingface.co/google/gemma-3-27b-it",
        "latencySource": null
      },
      {
        "model": "Gemma 3 4B IT",
        "vendor": "Google",
        "family": "Gemma 3",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 4,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 59.6,
        "mmlu_variant": "MMLU",
        "gpqa": 15,
        "humaneval_or_swe": 36,
        "coding_benchmark": "HumanEval",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "Gemma Terms of Use (custom)",
        "weights": "open",
        "hosting": "Self-host / on-device",
        "releaseDate": "2025-03",
        "note": "Superseded by Gemma 4 E4B.",
        "source": "https://huggingface.co/google/gemma-3-27b-it",
        "benchmarkSource": "https://huggingface.co/google/gemma-3-27b-it",
        "latencySource": null
      },
      {
        "model": "Llama 4 Maverick",
        "vendor": "Meta",
        "family": "Llama 4",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "Llama 4 Behemoth (codistillation)",
        "params_b": 400,
        "active_params_b": 17,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 80.5,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 69.8,
        "humaneval_or_swe": 43.4,
        "coding_benchmark": "LiveCodeBench",
        "aime": null,
        "latency_ttft_ms": 920,
        "contextK": 1000,
        "maxOutputK": null,
        "license": "Llama 4 Community License (custom commercial)",
        "weights": "open",
        "hosting": "Self-host / Bedrock / Together / Groq",
        "releaseDate": "2025-04",
        "note": "Meta's Llama 4 launch post says Maverick was codistilled from the ~2T Llama 4 Behemoth teacher, which was never released; the Llama 4 model card documents the benchmarks only and does not mention Behemoth. 128 experts, 17B active. Artificial Analysis blended price $0.31/MTok.",
        "source": "https://ai.meta.com/blog/llama-4-multimodal-intelligence/",
        "benchmarkSource": "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md",
        "latencySource": "https://artificialanalysis.ai/models/comparisons/gpt-oss-120b-vs-llama-4-maverick"
      },
      {
        "model": "Llama 4 Scout",
        "vendor": "Meta",
        "family": "Llama 4",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "Llama 4 Behemoth (codistillation)",
        "params_b": 109,
        "active_params_b": 17,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 74.3,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 57.2,
        "humaneval_or_swe": 32.8,
        "coding_benchmark": "LiveCodeBench",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 10000,
        "maxOutputK": null,
        "license": "Llama 4 Community License (custom commercial)",
        "weights": "open",
        "hosting": "Self-host / Bedrock / Together",
        "releaseDate": "2025-04",
        "note": "10M-token context window — the longest in this table by an order of magnitude. Retains 92% of Maverick MMLU-Pro at a quarter of the total parameters.",
        "source": "https://ai.meta.com/blog/llama-4-multimodal-intelligence/",
        "benchmarkSource": "https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md",
        "latencySource": null
      },
      {
        "model": "Llama 3.3 70B Instruct",
        "vendor": "Meta",
        "family": "Llama 3",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 70,
        "active_params_b": null,
        "input_per_mtok_usd": 1.04,
        "output_per_mtok_usd": 1.04,
        "blended_per_mtok_usd": 1.04,
        "cost_per_1m_requests_usd": 520,
        "mmlu": 68.9,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 50.5,
        "humaneval_or_swe": 88.4,
        "coding_benchmark": "HumanEval",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "Llama 3.3 Community License (custom commercial)",
        "weights": "open",
        "hosting": "Together / Groq / Bedrock / self-host",
        "releaseDate": "2024-12-06",
        "note": "MMLU 86.0 CoT, IFEval 92.1. Still the most common open student base for enterprise distillation projects; also a supported student in Amazon Bedrock Model Distillation.",
        "source": "https://www.together.ai/pricing",
        "benchmarkSource": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md",
        "latencySource": null
      },
      {
        "model": "Llama 3.2 3B Instruct",
        "vendor": "Meta",
        "family": "Llama 3",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "Llama 3.1 8B and 70B (token-level logit distillation after pruning)",
        "params_b": 3,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 63.4,
        "mmlu_variant": "MMLU",
        "gpqa": 32.8,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "Llama 3.2 Community License (custom commercial)",
        "weights": "open",
        "hosting": "Self-host / edge / Bedrock",
        "releaseDate": "2024-09-25",
        "note": "Meta documents this explicitly: logits from Llama 3.1 8B and 70B were used as token-level targets during pretraining, with distillation applied after pruning to recover performance. MATH CoT 48.0.",
        "source": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md",
        "benchmarkSource": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md",
        "latencySource": null
      },
      {
        "model": "Llama 3.2 1B Instruct",
        "vendor": "Meta",
        "family": "Llama 3",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "Llama 3.1 8B and 70B (token-level logit distillation after pruning)",
        "params_b": 1,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 49.3,
        "mmlu_variant": "MMLU",
        "gpqa": 27.2,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "Llama 3.2 Community License (custom commercial)",
        "weights": "open",
        "hosting": "Self-host / on-device / Bedrock",
        "releaseDate": "2024-09-25",
        "note": "The reference \"distilled onto a phone\" model. MATH CoT 30.6, ARC-C 59.4.",
        "source": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md",
        "benchmarkSource": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md",
        "latencySource": null
      },
      {
        "model": "DeepSeek-V4-Pro",
        "vendor": "DeepSeek",
        "family": "DeepSeek V4",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": 1600,
        "active_params_b": 49,
        "input_per_mtok_usd": 1.32,
        "output_per_mtok_usd": 3.96,
        "blended_per_mtok_usd": 1.98,
        "cost_per_1m_requests_usd": 924,
        "mmlu": 87.5,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 90.1,
        "humaneval_or_swe": 80.6,
        "coding_benchmark": "SWE-bench Verified",
        "aime": null,
        "latency_ttft_ms": 1900,
        "contextK": 1000,
        "maxOutputK": 384,
        "license": "MIT",
        "weights": "open",
        "hosting": "DeepSeek API / self-host (1.6T MoE)",
        "releaseDate": "2026-04-26",
        "note": "MIT-licensed 1.6T-parameter MoE that scores within 4 points of Gemini 3.1 Pro on GPQA. Peak price shown; off-peak is half ($0.66/$1.98). LiveCodeBench 93.5, HLE (with tools) 48.2.",
        "source": "https://api-docs.deepseek.com/quick_start/pricing/",
        "benchmarkSource": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro",
        "latencySource": "https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-deepseek-v4-pro"
      },
      {
        "model": "DeepSeek-V4-Flash",
        "vendor": "DeepSeek",
        "family": "DeepSeek V4",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "DeepSeek V4 domain experts (on-policy distillation consolidation)",
        "params_b": 284,
        "active_params_b": 13,
        "input_per_mtok_usd": 0.44,
        "output_per_mtok_usd": 1.32,
        "blended_per_mtok_usd": 0.66,
        "cost_per_1m_requests_usd": 308,
        "mmlu": 86.2,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 88.1,
        "humaneval_or_swe": 79,
        "coding_benchmark": "SWE-bench Verified",
        "aime": null,
        "latency_ttft_ms": 1190,
        "contextK": 1000,
        "maxOutputK": 384,
        "license": "MIT",
        "weights": "open",
        "hosting": "DeepSeek API / self-host",
        "releaseDate": "2026-07",
        "note": "The single best value in this table on GPQA-per-dollar: 97.8% of V4-Pro GPQA at a third of the price and 2.3x the output speed. Off-peak $0.22/$0.66. Standard rates rose about 3.1x on input and 4.7x on output from the pre-August-2026 rate; cache-hit input rates rose by up to 11x.",
        "source": "https://api-docs.deepseek.com/quick_start/pricing/",
        "benchmarkSource": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
        "latencySource": "https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-deepseek-v4-pro"
      },
      {
        "model": "DeepSeek-V3.2",
        "vendor": "DeepSeek",
        "family": "DeepSeek V3",
        "role": "open",
        "isDistilled": true,
        "teacher": "DeepSeek specialist models (specialist distillation into the generalist)",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 85,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 82.4,
        "humaneval_or_swe": 73.1,
        "coding_benchmark": "SWE-bench Verified",
        "aime": 93.1,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "MIT",
        "weights": "open",
        "hosting": "Self-host / third-party",
        "releaseDate": "2025-12",
        "note": "Trains domain specialists from the base checkpoint and distils their outputs back into one generalist. No longer on the first-party DeepSeek API.",
        "source": "https://arxiv.org/html/2512.02556",
        "benchmarkSource": "https://arxiv.org/html/2512.02556",
        "latencySource": null
      },
      {
        "model": "DeepSeek-R1 (0528)",
        "vendor": "DeepSeek",
        "family": "DeepSeek R1",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": 685,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 85,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 81,
        "humaneval_or_swe": 73.3,
        "coding_benchmark": "LiveCodeBench",
        "aime": 87.5,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": 64,
        "license": "MIT",
        "weights": "open",
        "hosting": "Self-host / third-party",
        "releaseDate": "2025-05-28",
        "note": "The teacher behind the six R1-Distill students. MIT licence explicitly permits distillation. Original Jan-2025 R1: GPQA 71.5, AIME24 79.8, MMLU-Pro 84.0.",
        "source": "https://huggingface.co/deepseek-ai/DeepSeek-R1-0528",
        "benchmarkSource": "https://huggingface.co/deepseek-ai/DeepSeek-R1-0528",
        "latencySource": null
      },
      {
        "model": "DeepSeek-R1-Distill-Llama-70B",
        "vendor": "DeepSeek / Meta base",
        "family": "R1-Distill",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "DeepSeek-R1",
        "params_b": 70,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 65.2,
        "humaneval_or_swe": 57.5,
        "coding_benchmark": "LiveCodeBench",
        "aime": 70,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "MIT (weights) over Llama 3.3 Community License base",
        "weights": "open",
        "hosting": "Self-host / third-party",
        "releaseDate": "2025-01-20",
        "note": "Highest-scoring R1 student: 91% of teacher GPQA, MATH-500 94.5. Dual-licence stack is the usual legal snag.",
        "source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "benchmarkSource": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "latencySource": null
      },
      {
        "model": "DeepSeek-R1-Distill-Qwen-32B",
        "vendor": "DeepSeek / Qwen base",
        "family": "R1-Distill",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "DeepSeek-R1",
        "params_b": 32,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 62.1,
        "humaneval_or_swe": 57.2,
        "coding_benchmark": "LiveCodeBench",
        "aime": 72.6,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "MIT (weights), Qwen2.5-32B base under Apache 2.0",
        "weights": "open",
        "hosting": "Self-host (1x A100/H100)",
        "releaseDate": "2025-01-20",
        "note": "Beat o1-mini on AIME 2024 (72.6 vs 63.6), MATH-500 (94.3 vs 90.0) and GPQA (62.1 vs 60.0) at zero marginal licence cost. The row that changed enterprise procurement conversations in 2025.",
        "source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "benchmarkSource": "https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B",
        "latencySource": null
      },
      {
        "model": "DeepSeek-R1-Distill-Qwen-14B",
        "vendor": "DeepSeek / Qwen base",
        "family": "R1-Distill",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "DeepSeek-R1",
        "params_b": 14,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 59.1,
        "humaneval_or_swe": 53.1,
        "coding_benchmark": "LiveCodeBench",
        "aime": 69.7,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "MIT (weights), Qwen2.5-14B base under Apache 2.0",
        "weights": "open",
        "hosting": "Self-host (single 24-48GB GPU)",
        "releaseDate": "2025-01-20",
        "note": "MATH-500 93.9. Fits on one L40S/A10G-class card.",
        "source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "benchmarkSource": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "latencySource": null
      },
      {
        "model": "DeepSeek-R1-Distill-Llama-8B",
        "vendor": "DeepSeek / Meta base",
        "family": "R1-Distill",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "DeepSeek-R1",
        "params_b": 8,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 49,
        "humaneval_or_swe": 39.6,
        "coding_benchmark": "LiveCodeBench",
        "aime": 50.4,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "MIT (weights) over Llama 3.1 Community License base",
        "weights": "open",
        "hosting": "Self-host / edge",
        "releaseDate": "2025-01-20",
        "note": "MATH-500 89.1.",
        "source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "benchmarkSource": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "latencySource": null
      },
      {
        "model": "DeepSeek-R1-Distill-Qwen-7B",
        "vendor": "DeepSeek / Qwen base",
        "family": "R1-Distill",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "DeepSeek-R1",
        "params_b": 7,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 49.1,
        "humaneval_or_swe": 37.6,
        "coding_benchmark": "LiveCodeBench",
        "aime": 55.5,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "MIT (weights), Qwen2.5-Math-7B base under Apache 2.0",
        "weights": "open",
        "hosting": "Self-host / edge",
        "releaseDate": "2025-01-20",
        "note": "MATH-500 92.8. Built on a maths-specialised base, so general knowledge lags the AIME score.",
        "source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "benchmarkSource": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "latencySource": null
      },
      {
        "model": "DeepSeek-R1-Distill-Qwen-1.5B",
        "vendor": "DeepSeek / Qwen base",
        "family": "R1-Distill",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "DeepSeek-R1",
        "params_b": 1.5,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 33.8,
        "humaneval_or_swe": 16.9,
        "coding_benchmark": "LiveCodeBench",
        "aime": 28.9,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "MIT (weights), Qwen2.5-Math-1.5B base under Apache 2.0",
        "weights": "open",
        "hosting": "Self-host / CPU / on-device",
        "releaseDate": "2025-01-20",
        "note": "MATH-500 83.9 from a 1.5B model, but GPQA collapses to 33.8 — the clearest illustration that distillation transfers narrow skills far better than broad knowledge.",
        "source": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "benchmarkSource": "https://huggingface.co/deepseek-ai/DeepSeek-R1",
        "latencySource": null
      },
      {
        "model": "Qwen3.8-27B",
        "vendor": "Alibaba",
        "family": "Qwen3.8",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 27,
        "active_params_b": null,
        "input_per_mtok_usd": 0.5,
        "output_per_mtok_usd": 3,
        "blended_per_mtok_usd": 1.125,
        "cost_per_1m_requests_usd": 500,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 89.2,
        "humaneval_or_swe": 61.7,
        "coding_benchmark": "SWE-bench Pro",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 262,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Groq / Alibaba Model Studio / self-host",
        "releaseDate": "2026-08",
        "note": "89.2 GPQA Diamond from a 27B Apache-2.0 model, LiveCodeBench v6 90.3, Terminal-Bench 2.1 73.0. Price shown is Alibaba Model Studio's list rate ($0.50/$3.00); Groq hosts the same weights at $0.80/$4.00 and serves it at ~450 tok/s. Context extensible to 1M.",
        "source": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "benchmarkSource": "https://huggingface.co/Qwen/Qwen3.8-27B",
        "latencySource": null,
        "hostingPriceOptions": [
          {
            "host": "Alibaba Model Studio",
            "input_per_mtok_usd": 0.5,
            "output_per_mtok_usd": 3,
            "source": "https://www.alibabacloud.com/help/en/model-studio/model-pricing"
          },
          {
            "host": "Groq",
            "input_per_mtok_usd": 0.8,
            "output_per_mtok_usd": 4,
            "source": "https://console.groq.com/docs/models"
          }
        ]
      },
      {
        "model": "qwen3.8-max",
        "vendor": "Alibaba",
        "family": "Qwen3.8",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 2,
        "output_per_mtok_usd": 6,
        "blended_per_mtok_usd": 3,
        "cost_per_1m_requests_usd": 1400,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": null,
        "maxOutputK": null,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Alibaba Model Studio",
        "releaseDate": "2026",
        "note": "Closed flagship above the open Qwen line.",
        "source": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "benchmarkSource": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "latencySource": null
      },
      {
        "model": "qwen3.8-flash",
        "vendor": "Alibaba",
        "family": "Qwen3.8",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.15,
        "output_per_mtok_usd": 0.47,
        "blended_per_mtok_usd": 0.23,
        "cost_per_1m_requests_usd": 107,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": null,
        "maxOutputK": null,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Alibaba Model Studio",
        "releaseDate": "2026",
        "note": "Hosted cheap tier; note that data residency is PRC/Singapore depending on region.",
        "source": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "benchmarkSource": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "latencySource": null
      },
      {
        "model": "qwen-turbo",
        "vendor": "Alibaba",
        "family": "Qwen",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.05,
        "output_per_mtok_usd": 0.2,
        "blended_per_mtok_usd": 0.0875,
        "cost_per_1m_requests_usd": 40,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": null,
        "maxOutputK": null,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Alibaba Model Studio",
        "releaseDate": "2025",
        "note": "Cheapest hosted tier in this table alongside GPT-5 nano.",
        "source": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "benchmarkSource": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "latencySource": null
      },
      {
        "model": "Qwen3-4B-Instruct-2507",
        "vendor": "Alibaba",
        "family": "Qwen3",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "Qwen3-32B / Qwen3-235B-A22B (off-policy + on-policy strong-to-weak distillation)",
        "params_b": 4,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 69.6,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 62,
        "humaneval_or_swe": 35.1,
        "coding_benchmark": "LiveCodeBench v6",
        "aime": 47.4,
        "latency_ttft_ms": null,
        "contextK": 262,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Self-host / edge",
        "releaseDate": "2025-07",
        "note": "The Qwen3 technical report documents strong-to-weak distillation for the 0.6B/1.7B/4B/8B/14B dense models and 30B-A3B, and states it beat RL on both performance and training cost.",
        "source": "https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507",
        "benchmarkSource": "https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507",
        "latencySource": null
      },
      {
        "model": "qwen3-8b (hosted)",
        "vendor": "Alibaba",
        "family": "Qwen3",
        "role": "distilled",
        "isDistilled": true,
        "teacher": "Qwen3-32B / Qwen3-235B-A22B (strong-to-weak distillation)",
        "params_b": 8,
        "active_params_b": null,
        "input_per_mtok_usd": 0.18,
        "output_per_mtok_usd": 0.7,
        "blended_per_mtok_usd": 0.31,
        "cost_per_1m_requests_usd": 142,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Alibaba Model Studio / self-host",
        "releaseDate": "2025-04",
        "note": "One of the five dense Qwen3 sizes the technical report names as strong-to-weak distillation targets.",
        "source": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "benchmarkSource": "https://arxiv.org/pdf/2505.09388",
        "latencySource": null
      },
      {
        "model": "Phi-4 (14B)",
        "vendor": "Microsoft",
        "family": "Phi",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 14,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 70.4,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 56.1,
        "humaneval_or_swe": 82.6,
        "coding_benchmark": "HumanEval",
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 16,
        "maxOutputK": null,
        "license": "MIT",
        "weights": "open",
        "hosting": "Self-host / Azure AI Foundry",
        "releaseDate": "2024-12",
        "note": "MMLU 84.8, MATH 80.4. Trained on synthetic \"textbook-like\" data; Microsoft describes the recipe as synthetic-data curation, not distillation. 16K context is the practical limit for enterprise use.",
        "source": "https://huggingface.co/microsoft/phi-4",
        "benchmarkSource": "https://huggingface.co/microsoft/phi-4",
        "latencySource": null
      },
      {
        "model": "Phi-4-mini-instruct (3.8B)",
        "vendor": "Microsoft",
        "family": "Phi",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 3.8,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": 52.8,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 25.2,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "MIT",
        "weights": "open",
        "hosting": "Self-host / edge / Azure AI Foundry",
        "releaseDate": "2025-02",
        "note": "MMLU 67.3, GSM8K 88.6, MATH 64.0. MIT licence makes it the least legally encumbered small model here.",
        "source": "https://huggingface.co/microsoft/Phi-4-mini-instruct",
        "benchmarkSource": "https://huggingface.co/microsoft/Phi-4-mini-instruct",
        "latencySource": null
      },
      {
        "model": "Mistral Medium 3.5",
        "vendor": "Mistral AI",
        "family": "Mistral",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 1.5,
        "output_per_mtok_usd": 7.5,
        "blended_per_mtok_usd": 3,
        "cost_per_1m_requests_usd": 1350,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": null,
        "maxOutputK": null,
        "license": "Modified MIT",
        "weights": "open",
        "hosting": "Mistral API / self-host",
        "releaseDate": "2026-04",
        "note": "Frontier-class multimodal model for agentic and coding work; EU-headquartered vendor with EU data residency.",
        "source": "https://mistral.ai/pricing/api",
        "benchmarkSource": "https://docs.mistral.ai/getting-started/models/models_overview/",
        "latencySource": null
      },
      {
        "model": "Mistral Large 3",
        "vendor": "Mistral AI",
        "family": "Mistral",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.5,
        "output_per_mtok_usd": 1.5,
        "blended_per_mtok_usd": 0.75,
        "cost_per_1m_requests_usd": 350,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": null,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Mistral API / self-host",
        "releaseDate": "2025-12",
        "note": "Apache 2.0 large model at $0.50/$1.50 — the cheapest large open-weight API in this table.",
        "source": "https://mistral.ai/pricing/api",
        "benchmarkSource": "https://docs.mistral.ai/getting-started/models/models_overview/",
        "latencySource": null
      },
      {
        "model": "Mistral Small 4",
        "vendor": "Mistral AI",
        "family": "Mistral",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.15,
        "output_per_mtok_usd": 0.6,
        "blended_per_mtok_usd": 0.2625,
        "cost_per_1m_requests_usd": 120,
        "mmlu": 78,
        "mmlu_variant": "MMLU-Pro",
        "gpqa": 71.2,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": null,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Mistral API / self-host",
        "releaseDate": "2026-03-16",
        "note": "Hybrid instruct + reasoning + coding model under Apache 2.0 at GPT-4o-mini prices. The EU-sovereignty default.",
        "source": "https://mistral.ai/pricing/api",
        "benchmarkSource": "https://openrouter.ai/mistralai/mistral-small-2603",
        "latencySource": null
      },
      {
        "model": "Ministral 3 14B Instruct",
        "vendor": "Mistral AI",
        "family": "Ministral 3",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 14,
        "active_params_b": null,
        "input_per_mtok_usd": 0.2,
        "output_per_mtok_usd": 0.2,
        "blended_per_mtok_usd": 0.2,
        "cost_per_1m_requests_usd": 100,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 71.2,
        "humaneval_or_swe": 64.6,
        "coding_benchmark": "LiveCodeBench",
        "aime": 85,
        "latency_ttft_ms": null,
        "contextK": 256,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Mistral API / self-host",
        "releaseDate": "2025-12-02",
        "note": "13.5B language model + 0.4B vision encoder, 256K context, AIME25 85.0, MATH 90.4. Symmetric $0.20 in/out pricing. MMLU is not separately published for the Instruct variant (the model card's 79.4 is the base model), so it is left undisclosed here.",
        "source": "https://mistral.ai/pricing/api",
        "benchmarkSource": "https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512",
        "latencySource": null
      },
      {
        "model": "Ministral 3 8B",
        "vendor": "Mistral AI",
        "family": "Ministral 3",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 8,
        "active_params_b": null,
        "input_per_mtok_usd": 0.15,
        "output_per_mtok_usd": 0.15,
        "blended_per_mtok_usd": 0.15,
        "cost_per_1m_requests_usd": 75,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 256,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Mistral API / self-host",
        "releaseDate": "2025-12-02",
        "note": "Mid size of the Ministral 3 edge family.",
        "source": "https://mistral.ai/pricing/api",
        "benchmarkSource": "https://docs.mistral.ai/getting-started/models/models_overview/",
        "latencySource": null
      },
      {
        "model": "Ministral 3 3B",
        "vendor": "Mistral AI",
        "family": "Ministral 3",
        "role": "open",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": 3,
        "active_params_b": null,
        "input_per_mtok_usd": 0.1,
        "output_per_mtok_usd": 0.1,
        "blended_per_mtok_usd": 0.1,
        "cost_per_1m_requests_usd": 50,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 256,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Mistral API / self-host / edge",
        "releaseDate": "2025-12-02",
        "note": "Apache 2.0 edge model with a 256K window — unusual combination.",
        "source": "https://mistral.ai/pricing/api",
        "benchmarkSource": "https://docs.mistral.ai/getting-started/models/models_overview/",
        "latencySource": null
      },
      {
        "model": "Amazon Nova Premier",
        "vendor": "Amazon",
        "family": "Nova",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1000,
        "maxOutputK": 10,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Amazon Bedrock",
        "releaseDate": "2025-04",
        "note": "AWS documents it as \"best teacher for distilling custom models\" — teacher to Pro, Lite and Micro in Bedrock Model Distillation. Cannot itself be fine-tuned.",
        "source": "https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html",
        "benchmarkSource": "https://docs.aws.amazon.com/nova/latest/userguide/what-is-nova.html",
        "latencySource": null
      },
      {
        "model": "Amazon Nova Pro",
        "vendor": "Amazon",
        "family": "Nova",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.8,
        "output_per_mtok_usd": 3.2,
        "blended_per_mtok_usd": 1.4,
        "cost_per_1m_requests_usd": 640,
        "mmlu": 85.9,
        "mmlu_variant": "MMLU",
        "gpqa": 46.9,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 300,
        "maxOutputK": 10,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Amazon Bedrock",
        "releaseDate": "2024-12",
        "note": "AWS supports Nova Pro as a distillation student (and as a teacher to smaller Nova tiers) in Bedrock Model Distillation. That is a capability of the service, not a statement about how the shipped weights were trained: Nova Pro shipped 2024-12, four months before Nova Premier. Training provenance undisclosed.",
        "source": "https://ucstrategies.com/news/amazon-nova-pro-aws-bedrock-model-guide-specs-pricing-2026/",
        "benchmarkSource": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf",
        "latencySource": null,
        "distillationServiceRole": "Supported student (and teacher) in Amazon Bedrock Model Distillation"
      },
      {
        "model": "Amazon Nova Lite",
        "vendor": "Amazon",
        "family": "Nova",
        "role": "distilled",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.06,
        "output_per_mtok_usd": 0.24,
        "blended_per_mtok_usd": 0.105,
        "cost_per_1m_requests_usd": 48,
        "mmlu": 80.5,
        "mmlu_variant": "MMLU",
        "gpqa": 42,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 300,
        "maxOutputK": 10,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Amazon Bedrock",
        "releaseDate": "2024-12",
        "note": "94% of Nova Pro MMLU at 7.5% of the input price. Multimodal with a 300K window. AWS supports it as a distillation student in Bedrock; the shipped weights' training provenance is undisclosed.",
        "source": "https://pricepertoken.com/pricing-page/model/amazon-nova-lite-v1",
        "benchmarkSource": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf",
        "latencySource": null,
        "distillationServiceRole": "Supported student in Amazon Bedrock Model Distillation"
      },
      {
        "model": "Amazon Nova Micro",
        "vendor": "Amazon",
        "family": "Nova",
        "role": "distilled",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.035,
        "output_per_mtok_usd": 0.14,
        "blended_per_mtok_usd": 0.0613,
        "cost_per_1m_requests_usd": 28,
        "mmlu": 77.6,
        "mmlu_variant": "MMLU",
        "gpqa": 40,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": 10,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Amazon Bedrock",
        "releaseDate": "2024-12",
        "note": "Cheapest per-token model in this entire table ($0.035 in / $0.14 out). Text-only, 128K. AWS supports it as a distillation student in Bedrock; the shipped weights' training provenance is undisclosed.",
        "source": "https://pricepertoken.com/pricing-page/model/amazon-nova-micro-v1",
        "benchmarkSource": "https://assets.amazon.science/96/7d/0d3e59514abf8fdcfafcdc574300/nova-tech-report-20250317-0810.pdf",
        "latencySource": null,
        "distillationServiceRole": "Supported student in Amazon Bedrock Model Distillation"
      },
      {
        "model": "Amazon Nova 2 Lite",
        "vendor": "Amazon",
        "family": "Nova 2",
        "role": "small-sibling",
        "isDistilled": false,
        "teacher": "undisclosed",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 0.3,
        "output_per_mtok_usd": 2.5,
        "blended_per_mtok_usd": 0.85,
        "cost_per_1m_requests_usd": 370,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 1000,
        "maxOutputK": null,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "Amazon Bedrock",
        "releaseDate": "2025-12",
        "note": "Adds extended thinking with three intensity levels, web grounding and a code interpreter; supports SFT and RFT. ~149 output tok/s.",
        "source": "https://pricepertoken.com/pricing-page/model/amazon-nova-2-lite-v1",
        "benchmarkSource": "https://docs.aws.amazon.com/nova/latest/nova2-userguide/whats-new.html",
        "latencySource": null
      },
      {
        "model": "SmolLM3-3B",
        "vendor": "Hugging Face",
        "family": "SmolLM",
        "role": "open",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": 3,
        "active_params_b": null,
        "input_per_mtok_usd": null,
        "output_per_mtok_usd": null,
        "blended_per_mtok_usd": null,
        "cost_per_1m_requests_usd": null,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": 41.7,
        "humaneval_or_swe": 30.48,
        "coding_benchmark": "HumanEval+",
        "aime": 36.7,
        "latency_ttft_ms": null,
        "contextK": 128,
        "maxOutputK": null,
        "license": "Apache 2.0",
        "weights": "open",
        "hosting": "Self-host / edge",
        "releaseDate": "2025-07",
        "note": "Fully open: weights, data mixture and training configs. GPQA 35.7 without thinking, 41.7 with. The reference point for buyers who need a reproducible supply chain rather than the best score.",
        "source": "https://huggingface.co/HuggingFaceTB/SmolLM3-3B",
        "benchmarkSource": "https://huggingface.co/HuggingFaceTB/SmolLM3-3B",
        "latencySource": null
      },
      {
        "model": "grok-4.6",
        "vendor": "xAI",
        "family": "Grok",
        "role": "teacher",
        "isDistilled": false,
        "teacher": "n/a",
        "params_b": null,
        "active_params_b": null,
        "input_per_mtok_usd": 2,
        "output_per_mtok_usd": 6,
        "blended_per_mtok_usd": 3,
        "cost_per_1m_requests_usd": 1400,
        "mmlu": null,
        "mmlu_variant": null,
        "gpqa": null,
        "humaneval_or_swe": null,
        "coding_benchmark": null,
        "aime": null,
        "latency_ttft_ms": null,
        "contextK": 200,
        "maxOutputK": null,
        "license": "Proprietary API",
        "weights": "closed",
        "hosting": "xAI API",
        "releaseDate": "2026-08",
        "note": "Prices shown are the <200K-context tier; longer contexts cost more.",
        "source": "https://docs.x.ai/docs/models",
        "benchmarkSource": "https://docs.x.ai/docs/models",
        "latencySource": null
      }
    ],
    "decisionGuide": [
      {
        "useCase": "High-volume classification, routing, tagging, PII redaction",
        "recommendation": "Amazon Nova Micro or GPT-5 nano; Ministral 3 3B if you must self-host",
        "why": "At 400 input + 100 output tokens a request, Nova Micro costs $28 per million requests and GPT-5 nano $60. Neither needs reasoning. Nova Micro is a documented Bedrock distillation student, so you can also distil it further on your own labels. Ministral 3 3B (Apache 2.0, 256K context) is the equivalent when the data cannot leave your network.",
        "models": [
          "Amazon Nova Micro",
          "GPT-5 nano",
          "qwen-turbo",
          "Ministral 3 3B",
          "Gemma 4 E4B"
        ]
      },
      {
        "useCase": "Customer-facing chat and support deflection",
        "recommendation": "GPT-5.6 Luna at low effort, with Gemini 3.5 Flash-Lite as the latency fallback",
        "why": "Luna measures 1.70s time-to-first-token at low effort, holds 87.0 GPQA Diamond, and costs $200 per million requests. Gemini 3.5 Flash-Lite runs at 391 output tokens/s if perceived typing speed matters more than depth. Anthropic’s own worked example puts 10,000 support conversations on Claude Haiku 4.5 at about $37, which is the same order of magnitude — pick on latency and evals, not price.",
        "models": [
          "GPT-5.6 Luna",
          "Gemini 3.5 Flash-Lite",
          "Claude Haiku 4.5",
          "Amazon Nova 2 Lite"
        ]
      },
      {
        "useCase": "Retrieval-augmented Q&A over a large private corpus",
        "recommendation": "GPT-5.6 Luna or Gemini 3.1 Flash-Lite, with aggressive prompt caching",
        "why": "Both carry 1M+ context at $0.20–$0.25 per million input tokens, and RAG is input-heavy, so the input rate dominates. GPT-5.6 Luna is the stronger of the two on knowledge benchmarks (87.0 GPQA Diamond vs Gemini 3.1 Flash-Lite’s 72.2), so prefer Luna where answer quality over technical corpora matters and Flash-Lite where throughput and price dominate. AWS measures under 2% accuracy loss on RAG-style tasks from distillation, so this is the workload class where cheap tiers are safest. Cache reads at 10% of input beat any further tier reduction.",
        "models": [
          "GPT-5.6 Luna",
          "Gemini 3.1 Flash-Lite",
          "DeepSeek-V4-Flash",
          "Claude Haiku 4.5"
        ]
      },
      {
        "useCase": "Autonomous coding agent that must finish multi-step tasks unsupervised",
        "recommendation": "Claude Opus 5; drop to Claude Sonnet 5 only after measuring your own retry rate",
        "why": "This is the one workload where the retention curve is genuinely steep: Opus 5 scores 96.0 on SWE-bench Verified, Sonnet 5 85.2 and Haiku 4.5 73.3. A 20% failure delta on a long-horizon task compounds into far more than 20% extra cost once retries and human review are counted. Gemini 3.1 Pro (80.6 SWE-bench Verified) and DeepSeek-V4-Pro (80.6, MIT) are the credible alternatives.",
        "models": [
          "Claude Opus 5",
          "Claude Sonnet 5",
          "Gemini 3.1 Pro (Preview)",
          "DeepSeek-V4-Pro",
          "GPT-5.6 Sol"
        ]
      },
      {
        "useCase": "Code completion and inline suggestions in an IDE",
        "recommendation": "Qwen3.8-27B on Groq, or gpt-oss-20b if you need the weights",
        "why": "Latency is the product here. Qwen3.8-27B posts LiveCodeBench v6 90.3 and SWE-bench Pro 61.7 under Apache 2.0 and is served at roughly 450 tokens/s; gpt-oss-20b runs at about 1000 tokens/s and fits in 16GB. Neither needs a frontier tier because the human reviews every suggestion.",
        "models": [
          "Qwen3.8-27B",
          "gpt-oss-20b",
          "Ministral 3 14B Instruct",
          "Mistral Small 4"
        ]
      },
      {
        "useCase": "Regulated workload that cannot leave your infrastructure",
        "recommendation": "Gemma 4 31B or 26B A4B (Apache 2.0); Phi-4-mini (MIT) at the small end",
        "why": "Gemma 4 31B posts MMLU-Pro 85.2 and GPQA Diamond 84.3 under Apache 2.0 with a 256K window — self-hostable quality that did not exist a year ago. The 26B MoE gets 82.3 GPQA with 3.8B active parameters, so it serves cheaply. Phi-4-mini is MIT with a 128K window for the smallest footprint. Avoid Gemma 3 (custom licence) and Llama (700M MAU clause plus naming obligations) if legal review is the bottleneck.",
        "models": [
          "Gemma 4 31B",
          "Gemma 4 26B A4B (MoE)",
          "Phi-4-mini-instruct (3.8B)",
          "Mistral Small 4",
          "SmolLM3-3B"
        ]
      },
      {
        "useCase": "EU data-sovereignty requirement",
        "recommendation": "Mistral Small 4 or Ministral 3 14B, both Apache 2.0 from an EU vendor",
        "why": "Mistral Small 4 posts MMLU-Pro 78.0 and GPQA Diamond 71.2 at $0.15/$0.60; Ministral 3 14B posts GPQA Diamond 71.2, AIME25 85.0 and MATH 90.4 with a 256K context at symmetric $0.20 pricing. If you must stay on a US hyperscaler, Anthropic’s inference_geo:\"us\" and Bedrock/Vertex regional endpoints carry a 10% premium.",
        "models": [
          "Mistral Small 4",
          "Ministral 3 14B Instruct",
          "Mistral Medium 3.5",
          "Gemma 4 12B Unified"
        ]
      },
      {
        "useCase": "On-device or offline assistant (laptop, handset, vehicle)",
        "recommendation": "Gemma 4 E4B, with Llama 3.2 3B as the proven-in-production alternative",
        "why": "Gemma 4 E4B runs at 4.5B effective parameters with MMLU-Pro 69.4 and a 128K window — roughly GPT-4o mini class, entirely offline, Apache 2.0. Llama 3.2 3B is the reference documented distillation (logits from Llama 3.1 8B and 70B as token-level targets after pruning) and has the widest edge-runtime support. Expect broad-knowledge accuracy to degrade much faster than format compliance at this size.",
        "models": [
          "Gemma 4 E4B",
          "Llama 3.2 3B Instruct",
          "Ministral 3 3B",
          "Phi-4-mini-instruct (3.8B)",
          "SmolLM3-3B"
        ]
      },
      {
        "useCase": "Maths, scientific reasoning and quantitative analysis at low cost",
        "recommendation": "DeepSeek-V4-Flash, or R1-Distill-Qwen-32B if the weights must be local",
        "why": "V4-Flash keeps 97.8% of V4-Pro’s GPQA Diamond (88.1 vs 90.1) at a third of the price under MIT. For self-hosting, R1-Distill-Qwen-32B fits on one 80GB GPU and beat o1-mini on AIME 2024, MATH-500 and GPQA Diamond. Gemma 4 31B’s AIME 2026 score of 89.2 is the Apache 2.0 alternative if MIT-from-a-PRC-vendor is a procurement problem.",
        "models": [
          "DeepSeek-V4-Flash",
          "DeepSeek-R1-Distill-Qwen-32B",
          "Gemma 4 31B",
          "Ministral 3 14B Instruct"
        ]
      },
      {
        "useCase": "You want to build your own distilled model on your own task data",
        "recommendation": "Amazon Bedrock Model Distillation for a managed path; DeepSeek V4 or Qwen3.8 as the teacher if you build it yourself",
        "why": "Bedrock takes only your prompts, generates teacher responses and fine-tunes the student, with Nova Premier/Pro, Claude and Llama 3.3 70B / 3.2 1B / 3.2 3B in the supported graph; AWS claims up to 500% faster and 75% cheaper inference with under 2% accuracy loss on RAG. If you build the pipeline yourself, the teacher must be one whose licence permits it: OpenAI’s Services Agreement and Anthropic’s Commercial Terms D.4 both prohibit training competing models on their outputs, while DeepSeek (MIT) and Qwen3.8/Gemma 4 (Apache 2.0) do not.",
        "models": [
          "Amazon Nova Premier",
          "Amazon Nova Pro",
          "Llama 3.3 70B Instruct",
          "DeepSeek-V4-Pro",
          "Qwen3.8-27B"
        ]
      }
    ],
    "buyerChecklist": [
      "Pin the effort/reasoning level in any latency SLA — the same model ID measures 1.70s and 19.87s to first token depending on it.",
      "Price the workload at 400 input + 100 output tokens per request before comparing tiers; input-heavy RAG and output-heavy generation rank models differently.",
      "Turn on prompt caching before changing model tier: cache reads cost 10% of input on most Claude models and 2.5% on Fable 5.1.",
      "Use the batch API for anything not user-facing — 50% off input and output at both OpenAI and Anthropic.",
      "Audit for pinned legacy model IDs. GPT-4o still lists at 2024 prices; Claude Sonnet 4.6 costs 50% more than the better Sonnet 5; Gemini 3.5 Flash costs 2x Gemini 3.8 Flash.",
      "Check the anti-distillation clause before choosing a teacher if you plan to train anything on its outputs.",
      "For open weights, check the base-model licence as well as the released weights licence — R1-Distill-Llama-70B is MIT weights over a Llama Community License base.",
      "Confirm the context window on the cheap tier specifically. Claude Haiku 4.5 is 200K while every other current Claude is 1M.",
      "Keep a second vendor integrated. In 2026 alone OpenAI cut prices twice, Anthropic cancelled an announced increase, and DeepSeek raised prices more than tenfold.",
      "Evaluate small models on your own broad-knowledge tasks, not just format compliance — that is where distillation retention drops fastest."
    ]
  }
}
