Trends
What’s actually changed in AI models
Every figure on this page is computed from the same catalogue the consultant uses — 316 models with published list prices, specifications and third-party benchmarks, dated by release. Nothing is smoothed, fitted or forecast.
9.3×
Frontier intelligence
6.8 → 63.1 since 2023-05
109×
Cheaper per capability
Index 35+ for $0.032/M, was $3.44/M
3.4
Open-weights gap
59.7 open vs 63.1 proprietary
198
Released in 12 months
162 models carry a benchmark score
Frontier intelligence, over time
Each step is a model that beat everything released before it. The line holds flat between steps because that is what actually happened — the record stood until something took it.
From OpenAI: GPT-4 at 6.8 to Claude Opus 5 at 63.1: a 9.3× increase, and the steps have been getting closer together rather than further apart.
Running maximum across all scored models, by release date
162 models publish an intelligence index; 18 of them set a record. Hover any point for the model that held it.
What a fixed level of capability costs
The frontier chart shows capability rising. This one shows the same capability getting cheaper: for each threshold, the running minimum blended price among models that clear it.
An intelligence index of 35 cost $3.44/M when OpenAI: GPT-5 first reached it. It now costs $0.032/M — roughly 109× less in 11 months. This is the single most useful trend on the page: a capability you priced last year is almost certainly cheaper now.
Running minimum blended price ($/M tokens, 3:1 mix) — log scale
- Index 45+
- Index 35+
- Index 25+
A line begins when a model first reached that threshold, so a late start means the capability did not exist earlier — not that data is missing. Log scale: each gridline is a tenfold change.
Open weights versus proprietary
Both frontiers, tracked separately. Open-weights models peak at 59.7 against 63.1 for proprietary — a gap of 3.4 points. Whether that gap is closing is the question this chart is for; read the vertical distance between the lines, not their absolute height.
Running maximum within each group, by release date
- Proprietary
- Open weights
140 of 316 catalogue models ship open weights; 15 of them set an open-weights record.
Every scored model, by release date
The frontier lines above trace the top edge of this cloud. What the cloud adds is spread: at any given moment there is a wide range of capability on sale, and the floor rises more slowly than the ceiling. A model released today is not automatically better than one from six months ago.
One dot per model that publishes a score
- Proprietary
- Open weights
162 models plotted. Models without a published intelligence index are absent, not at zero.
Price by capability tier
Median blended price each quarter, split by capability band. The bands deflate independently — a tier does not get cheaper because better models arrived above it, but because more entrants crowded into it. Bands appear only once at least two models occupied them in a quarter.
Blended $/M tokens at a 3:1 input:output mix, by quarter — log scale
- Index 50+
- Index 35–50
- Index under 35
Quarters with fewer than two models in a band are omitted rather than plotted, so a gap is thin evidence rather than a real absence.
Context windows
Median context window by quarter, split by licence. Proprietary models led for most of the period; the medians have since converged as million-token windows became ordinary rather than exceptional.
Median across models released each quarter
- Proprietary
- Open weights
Median, not maximum — one outlier with a huge window does not move this line.
Output speed
Median measured throughput across models released each quarter. Speed has improved, but far less dramatically than price has fallen — if latency is your binding constraint, it is the constraint that has moved least.
Tokens per second, across models released each quarter
Throughput is a point measurement and varies by provider, region and load. Treat the trend as directional, not as a service-level figure.
Release volume
Models entering the catalogue each quarter, and how many shipped open weights. 198 arrived in the last twelve months alone — which is the practical case for a tool like this one. Evaluating the field by hand stopped being feasible some time ago.
Total entering the catalogue, and the open-weights share
- All models
- Open weights
Counted by the release date recorded in the catalogue. The final quarter is partial.
Where each lab stands
The strongest scored model from each of the ten leading labs. This is a snapshot of the top of each range, not a measure of depth — a lab with one strong model and a lab with forty both appear once.
Highest intelligence index achieved by each maker
Only labs with at least one benchmark-scored model appear.
Capability against model size
For the open-weights models that publish a parameter count, capability plotted against size. The relationship is loose — the spread at any given size is wider than the trend across sizes, which is another way of saying architecture and training matter more than scale alone at this point.
Open-weights models that publish a parameter count — log scale
71 models publish both a parameter count and an intelligence score. Total parameters, not active — the catalogue does not distinguish them, so mixture-of-experts models look larger than they run.
How to read this
What the numbers are
Release dates come from the catalogue’s own timestamps. Intelligence, coding and agentic indices are third-party results from Artificial Analysis. Prices are list prices; “blended” means a 3:1 input:output token mix. Throughput is a measured point value.
What we do not do
- No smoothing, curve fitting or forecasting. Every vertex is a real model on a real date.
- No interpolation across gaps. A quarter with fewer than two qualifying models is omitted rather than estimated.
- No zero-filling. A model without a published benchmark is absent from that chart, never plotted at zero.
- Batch and free-tier endpoints are excluded so a single model is not counted twice as a release.
Where this is weakest
Benchmark coverage thins as you go back: 162 of 316 models carry an intelligence index, and the early quarters have only a handful each. Early points are therefore more sensitive to a single model than recent ones. The most recent quarter is partial by definition.
The catalogue is also a snapshot. It reflects models known at the time it was compiled, and it under-represents labs that do not submit to public benchmarks at all.
What this can’t tell you
Training compute, training-token counts, active-versus-total parameters, model architecture and country of origin are not in this dataset, so the charts that would need them are absent rather than approximated. Benchmark scores are also a proxy for behaviour on your workload, not a substitute for testing it.