Using the consultant

Reading the charts

What the bars, the scatter and the percentile ranks actually mean.

3 min read

Fit score

A 0–100 composite of percentile ranks on the benchmarks relevant to your task. A coding question weights the coding index most heavily, then intelligence, then agentic; a design question weights Design Arena Elo. It is a ranking aid, not a published benchmark — treat a 3-point gap as noise and a 15-point gap as real.

Weights are renormalised over the metrics each model actually reports, so a model missing one benchmark isn't quietly penalised. A model with no relevant coverage at all is left out of the ranking rather than scored badly.

Percentile ranks

“92nd percentile on intelligence” means the model scores higher than 92% of the models that publish an intelligence index. Percentiles are computed against distinct models only — batch and free-tier endpoints don't get a vote on where the distribution sits.

Blended price

Input and output tokens are priced separately, which makes two models hard to compare at a glance. Blended price collapses them at a 3:1 input:output mix — a typical ratio for chat and retrieval workloads.

The price-performance scatter

Price runs along a logarithmic x-axis because the catalogue spans four orders of magnitude, from free to $150 per million input tokens. On a linear axis every affordable model would collapse into one pile at the left edge. Up and to the left is better: more capability for less money.

Table view

Every chart has a Table button in its top-right corner. It shows the same numbers as rows, which is the accessible path and the one to use when you want to copy figures out.

Empty cells

A dash means the catalogue has no published value for that model and metric. It never means zero. If a comparison leaves a column mostly blank, that metric isn't a sound basis for the decision.