Account
Troubleshooting
Answers that look wrong, charts that look odd, and what to do about it.
2 min read
It recommended a model I've never heard of
That's usually correct rather than wrong — the catalogue includes far more models than any one team tracks, and smaller or open-weights models often score well on a narrow task. Open the shortlist table to see how it compares, and check the trade-off section for what you'd give up.
The answer excluded a model I expected
Three common reasons: it lacks the benchmark coverage to be scored for that task, it fails one of your constraints, or it's a batch/free-tier endpoint of a model that's already in the list. The Thinking panel shows how many models were filtered out at each step.
A chart bar sits at zero
In a head-to-head comparison, a zero-length bar means that model publishes no result for that dimension. It is not a score of zero. Use the table view to confirm.
The cost projection doesn't match my bill
Projections use list prices and assume 3,000 input and 700 output tokens per request unless you say otherwise. Caching, batch pricing, committed-use discounts and a different token mix all move the number. Give the consultant your real volume and token counts for a closer estimate.
Something is broken
Send us the question you asked, word for word. Because answers are deterministic, the exact wording is usually enough for us to reproduce it immediately.
Still stuck? Contact us