Account

Troubleshooting

Answers that look wrong, charts that look odd, and what to do about it.

2 min read

That's usually correct rather than wrong — the catalogue includes far more models than any one team tracks, and smaller or open-weights models often score well on a narrow task. Open the shortlist table to see how it compares, and check the trade-off section for what you'd give up.

The answer excluded a model I expected

Three common reasons: it lacks the benchmark coverage to be scored for that task, it fails one of your constraints, or it's a batch/free-tier endpoint of a model that's already in the list. The Thinking panel shows how many models were filtered out at each step.

A chart bar sits at zero

In a head-to-head comparison, a zero-length bar means that model publishes no result for that dimension. It is not a score of zero. Use the table view to confirm.

The cost projection doesn't match my bill

Projections use list prices and assume 3,000 input and 700 output tokens per request unless you say otherwise. Caching, batch pricing, committed-use discounts and a different token mix all move the number. Give the consultant your real volume and token counts for a closer estimate.

Something is broken

Send us the question you asked, word for word. Because answers are deterministic, the exact wording is usually enough for us to reproduce it immediately.

Still stuck? Contact us