It makes the call
Every answer opens with a recommendation and the score behind it — not a table for you to interpret. Ask for the cheapest model that still handles agentic coding and you get one name, with the competence bar it had to clear.
316 models · 51 makers · free to use
AI Windes is a consultant, not a leaderboard. Ask in plain language and get a recommendation, the pricing and benchmark evidence behind it, and the trade-off you’d otherwise discover in production.
Sign in with an email link. No password, no card, nothing to cancel.
Best model for agentic coding under $5 per million tokens
GPT-5.6 Terra (OpenAI) is my pick for coding under $5/M blended. It scores 97/100 on my composite fit index — a coding index of 76.7, 98th percentile.
Blended price
$2.25/M
$1 in · $6 out
Intelligence index
56.6
96th percentile
Context window
1.1M
128K max output
Fit for coding
Composite benchmark index, 0–100
Trade-off: Qwen3.8 Max is 4 points behind but costs 25% more. The cheapest shortlisted option is 15 points down — a real capability step, not a rounding error.
316
Models compared
384 catalogue entries
51
Makers covered
frontier labs to open-weights
44%
Open weights
140 self-hostable models
$0.749/M
Median blended price
8 priced at $0
Frontier capability has risen 9× in three years while the price of a fixed capability level has collapsed. Ten charts, no smoothing.
Every answer opens with a recommendation and the score behind it — not a table for you to interpret. Ask for the cheapest model that still handles agentic coding and you get one name, with the competence bar it had to clear.
Pricing, context limits, throughput, latency, licence and third-party benchmarks — rendered as charts and tables inline. Every figure is computed from the catalogue; nothing is asserted without a number under it.
The runner-up, the capability you give up, the batch endpoint that's half price but asynchronous, and the benchmarks that are simply missing. A recommendation with no downside is a sales pitch.
Enter your email, click the link we send. No password to invent, no card, no trial that quietly ends.
“Cheapest model with vision and a 200K context window.” “Compare Claude Opus 5 and GPT-5.2.” “What does 2M requests a month cost?” Plain language is the interface.
A named recommendation, the charts that justify it, the trade-off against the runner-up, and an honest note on where the data thins out.
Pricing
No plans, no seats, no usage meter, no card on file. Sign in with an email link and ask as many questions as you like.
Try AI WindesFrom the blog
The ones worth answering before you sign in. Everything else lives in the help docs.
Actually free. No paid tier, no usage cap, no card, no trial that converts. Signing in exists so your chats stay attached to you between visits and so we can tell you when the catalogue changes materially.
If we ever add paid features, nothing you already use starts charging without notice first.
No. Answers are computed directly from the catalogue by a deterministic analyst — no model provider is called and nothing is generated. That buys you three things:
Your constraints filter the catalogue, then the remaining models are scored on a composite of the benchmarks relevant to your task — coding index for a coding question, Design Arena Elo for a design one, and so on. Weights are renormalised over the metrics each model actually reports.
When you ask for the cheapest model that does something, candidates must clear a competence bar before the price sort runs — otherwise the cheapest answer is a model that technically has the capability flag and can’t do the job. Full method.
It’s excluded, never treated as zero. A model with no relevant coverage is left out of the ranking rather than scored badly, and the answer tells you how many models qualified.
This matters more than it sounds: only 213 of 384 entries publish an intelligence index and 117 publish a coding index. Newer and smaller open-weights models are the most likely to be thin. Where the data ends.
Trust it to take you from hundreds of options to two worth evaluating properly — that is the job it does well. Then test those two against your own workload.
Benchmarks measure benchmark performance, which is a proxy for how a model behaves on your task, not a substitute for finding out. Prices are list prices and exclude committed-use discounts. The consultant states these assumptions in every answer rather than burying them.
It only answers questions it can ground in the model catalogue. “How do I write Python code to reverse a list” mentions coding but is a request for code, not a model-selection question — so it declines. Rephrased as “which model is best at Python code generation”, it answers.
A confident answer built on no data is worse than a refusal. More on the scope rule.
384 entries covering list pricing, context and output limits, capability flags, measured throughput and latency, serving providers, licensing and knowledge cutoffs — plus third-party benchmark results from Artificial Analysis and Design Arena. We don’t run our own evaluations, take payment for placement, or weight any maker preferentially.
Send sign-in links and occasional product updates. It’s the only account identifier we hold — there’s no password, no profile, and your chats are never stored on our servers. Ask us and we’ll delete it the same day. Privacy policy.