Trends

What’s actually changed in AI models

Every figure on this page is computed from the same catalogue the consultant uses — 316 models with published list prices, specifications and third-party benchmarks, dated by release. Nothing is smoothed, fitted or forecast.

9.3×

Frontier intelligence

6.8 → 63.1 since 2023-05

109×

Cheaper per capability

Index 35+ for $0.032/M, was $3.44/M

3.4

Open-weights gap

59.7 open vs 63.1 proprietary

198

Released in 12 months

162 models carry a benchmark score

Frontier intelligence, over time

Each step is a model that beat everything released before it. The line holds flat between steps because that is what actually happened — the record stood until something took it.

From OpenAI: GPT-4 at 6.8 to Claude Opus 5 at 63.1: a 9.3× increase, and the steps have been getting closer together rather than further apart.

Frontier intelligence index

Running maximum across all scored models, by release date

0204060Jan '24Jan '25Jan '26Intelligence index

162 models publish an intelligence index; 18 of them set a record. Hover any point for the model that held it.

What a fixed level of capability costs

The frontier chart shows capability rising. This one shows the same capability getting cheaper: for each threshold, the running minimum blended price among models that clear it.

An intelligence index of 35 cost $3.44/M when OpenAI: GPT-5 first reached it. It now costs $0.032/M — roughly 109× less in 11 months. This is the single most useful trend on the page: a capability you priced last year is almost certainly cheaper now.

Cheapest model at each intelligence threshold

Running minimum blended price ($/M tokens, 3:1 mix) — log scale

  • Index 45+
  • Index 35+
  • Index 25+
$0.1$1$10Jul '25Jan '26Jul '26Blended $/MIndex 45+Index 35+Index 25+

A line begins when a model first reached that threshold, so a late start means the capability did not exist earlier — not that data is missing. Log scale: each gridline is a tenfold change.

Open weights versus proprietary

Both frontiers, tracked separately. Open-weights models peak at 59.7 against 63.1 for proprietary — a gap of 3.4 points. Whether that gap is closing is the question this chart is for; read the vertical distance between the lines, not their absolute height.

Frontier intelligence by licence type

Running maximum within each group, by release date

  • Proprietary
  • Open weights
0204060Jan '24Jan '25Jan '26Intelligence indexProprietaryOpen weights

140 of 316 catalogue models ship open weights; 15 of them set an open-weights record.

Every scored model, by release date

The frontier lines above trace the top edge of this cloud. What the cloud adds is spread: at any given moment there is a wide range of capability on sale, and the floor rises more slowly than the ceiling. A model released today is not automatically better than one from six months ago.

Intelligence index vs. release date

One dot per model that publishes a score

  • Proprietary
  • Open weights
0204060Jan '24Jan '25Jan '26Release dateIntelligence indexClaude Opus 5OpenAI: GPT-4

162 models plotted. Models without a published intelligence index are absent, not at zero.

Price by capability tier

Median blended price each quarter, split by capability band. The bands deflate independently — a tier does not get cheaper because better models arrived above it, but because more entrants crowded into it. Bands appear only once at least two models occupied them in a quarter.

Median price within each intelligence band

Blended $/M tokens at a 3:1 input:output mix, by quarter — log scale

  • Index 50+
  • Index 35–50
  • Index under 35
$1$10Jan '24Jan '25Jan '26Blended $/MIndex 50+Index 35–50Index under 35

Quarters with fewer than two models in a band are omitted rather than plotted, so a gap is thin evidence rather than a real absence.

Context windows

Median context window by quarter, split by licence. Proprietary models led for most of the period; the medians have since converged as million-token windows became ordinary rather than exceptional.

Median context window

Median across models released each quarter

  • Proprietary
  • Open weights
0500K1MJan '24Jan '25Jan '26Context tokensProprietaryOpen weights

Median, not maximum — one outlier with a huge window does not move this line.

Output speed

Median measured throughput across models released each quarter. Speed has improved, but far less dramatically than price has fallen — if latency is your binding constraint, it is the constraint that has moved least.

Median output throughput

Tokens per second, across models released each quarter

0 tok/s20 tok/s40 tok/s60 tok/sJan '24Jan '25Jan '26Tokens/sec

Throughput is a point measurement and varies by provider, region and load. Treat the trend as directional, not as a service-level figure.

Release volume

Models entering the catalogue each quarter, and how many shipped open weights. 198 arrived in the last twelve months alone — which is the practical case for a tool like this one. Evaluating the field by hand stopped being feasible some time ago.

Models released per quarter

Total entering the catalogue, and the open-weights share

  • All models
  • Open weights
Q2 '23
Q3 '23
Q4 '23
Q1 '24
Q2 '24
Q3 '24
Q4 '24
Q1 '25
Q2 '25
Q3 '25
Q4 '25
Q1 '26
Q2 '26
Q3 '26

Counted by the release date recorded in the catalogue. The final quarter is partial.

Where each lab stands

The strongest scored model from each of the ten leading labs. This is a snapshot of the top of each range, not a measure of depth — a lab with one strong model and a lab with forty both appear once.

Best model by lab

Highest intelligence index achieved by each maker

Anthropic · Claude Opus 5
63.1
OpenAI · OpenAI: GPT-5.6 Sol
60.9
Moonshot AI · MoonshotAI: Kimi K3
59.7
Qwen · Qwen: Qwen3.8 Max
58.1
Meta · Meta: Muse Spark 1.2
56.8
xAI · SpaceXAI: Grok 4.5
55.8
Z.ai · Z.ai: GLM 5.2
52.6
Google · Google: Gemini 3.5 Flash
52
DeepSeek · DeepSeek: DeepSeek V4 Flash 0731
51.8
MiniMax · MiniMax: MiniMax M3
45.4

Only labs with at least one benchmark-scored model appear.

Capability against model size

For the open-weights models that publish a parameter count, capability plotted against size. The relationship is loose — the spread at any given size is wider than the trend across sizes, which is another way of saying architecture and training matter more than scale alone at this point.

Intelligence index vs. total parameters

Open-weights models that publish a parameter count — log scale

0204060101001KTotal parameters (billions, log scale)Intelligence indexMoonshotAI: Kimi K3Qwen: Qwen3 VL 8B Instruct

71 models publish both a parameter count and an intelligence score. Total parameters, not active — the catalogue does not distinguish them, so mixture-of-experts models look larger than they run.

How to read this

What the numbers are

Release dates come from the catalogue’s own timestamps. Intelligence, coding and agentic indices are third-party results from Artificial Analysis. Prices are list prices; “blended” means a 3:1 input:output token mix. Throughput is a measured point value.

What we do not do

  • No smoothing, curve fitting or forecasting. Every vertex is a real model on a real date.
  • No interpolation across gaps. A quarter with fewer than two qualifying models is omitted rather than estimated.
  • No zero-filling. A model without a published benchmark is absent from that chart, never plotted at zero.
  • Batch and free-tier endpoints are excluded so a single model is not counted twice as a release.

Where this is weakest

Benchmark coverage thins as you go back: 162 of 316 models carry an intelligence index, and the early quarters have only a handful each. Early points are therefore more sensitive to a single model than recent ones. The most recent quarter is partial by definition.

The catalogue is also a snapshot. It reflects models known at the time it was compiled, and it under-represents labs that do not submit to public benchmarks at all.

What this can’t tell you

Training compute, training-token counts, active-versus-total parameters, model architecture and country of origin are not in this dataset, so the charts that would need them are absent rather than approximated. Benchmark scores are also a proxy for behaviour on your workload, not a substitute for testing it.