What do you need?

Pick the job, not the model. The profiles mirror how AI is actually used in 2026 — coding agents, support, writing, research — and everything below is recomputed in your browser.

One pick per mode. Each mode states a different preference; none of them is a universal score.

Economy

GPT-5.6 Luna (high)

OpenAI

ECST
$0.041
ETST
120s
Success
High
Capability
71.4

Evidence: High · coverage 86%

F1

Balanced

Claude Sonnet 5

Anthropic

ECST
$0.062
ETST
94s
Success
Very high
Capability
74.8

Evidence: High · coverage 91%

F1

Frontier

Gemini 3.1 Pro (max)

Google

ECST
$0.19
ETST
230s
Success
High
Capability
81.2

Evidence: Medium · coverage 62%

F2

Why?

Select a row or a recommendation to see how that pick was reached.

Nothing selected yet. Pick a system in the cards above or the chart below, and this panel will show the capability interval, the success prior, the budget constraints it passed, and the systems it robustly dominates.

Frontier

X = expected cost per successful task (log). Y = safe capability. Bubble size ranks throughput. All three point the same way: up, left and larger are better.

$0.01 $0.05 $0.20 $1.00 85 75 65 55
On the robust frontier (F1) Within 15% of a frontier system (F2) Behind, but inside the noise (F3) Recommended for a mode Bubble size ranks throughput — bigger is faster

All systems

239 systems match the current filters. A row is one system configuration, not a model family.

Capability is a 0–100 index with cohort mean 50 and standard deviation 15 — it is not a percentage. ECST and ETST are expected cost and time per successful task. Success is a modelled band; a row below the selected profile's success floor says so.
Model Org Capability Success (modelled) ECST ETST $/task Time Evidence Frontier
Claude Sonnet 5 Anthropic 74.8 Very high $0.062 94s $0.058 87 s B F1
GPT-5.6 Luna (high) OpenAI 71.4 High $0.041 120s $0.036 105 s B F1
Gemini 3.1 Pro (max) Google 81.2 High $0.194 230s $0.186 222 s C F2
Llama 5 405B Meta 63.9 Mixed $0.088 180s $0.062 125 s D F3
Mistral Large 4 Mistral 58.3 Below floor $0.147 320s $0.076 165 s D D

Methodology

Every number on this page, and how it was produced.

How the task profiles were chosen

The twelve profiles are not benchmark categories restated — they mirror how AI is actually used as of August 2026. OpenAI's usage study (How People Use ChatGPT, NBER w34255) finds writing is the most common work task and grounded question-answering dominates consumer usage; Anthropic's Economic Index reports show software engineering at roughly half of enterprise API traffic, with “modifying software to correct errors” the single most common task. The profile order follows that observed workload share: coding agents first, then support and operations, content, and knowledge work.

Each profile declares what one task means (“one resolved ticket reply”, “one merged fix with tests passing”), a weighting over the nine measured capability dimensions, a context floor, and a minimum acceptable success rate. Those declarations are judgment calls, published so you can disagree with them.

What a row is

Every row is a system, which means a model plus a reasoning level plus a serving route — not a model family. GPT-5.6 Luna (max) and GPT-5.6 Luna (high) are two different rows, because they cost different amounts, take different amounts of time, and succeed at different rates.

This matters because the thing you actually deploy is a configuration, not a brand. A benchmark table that collapses every effort level into one line hides the choice you are really making. Where a capability score was measured at a different effort level than the row it appears on, the row is graded lower and says so.

Why expected cost per successful task

The headline metric is ECST: the price per task divided by the probability that the task actually succeeds. Its sibling ETST does the same for wall-clock time. A model that costs half as much but fails twice as often is not cheaper, and this arithmetic says so without any tuning.

The previous version of this index ranked on Q^5 / cost^0.4. Those exponents were a value judgment written into a formula: they decided, for everyone, how much capability is worth. ECST needs no exponent. Failures pay for themselves in the numerator, so a cheap unreliable system rises in expected cost on its own.

Why capability is an index, not a percentage

Raw scores from different benchmarks are not comparable to each other. Within each dimension, a system's raw score is converted to its rank within the group of systems that have data for that dimension, then mapped onto a scale with mean 50 and standard deviation 15.

So 65 does not mean “65% correct”. It means roughly one standard deviation above the cohort. Never read these numbers as percentages. Uncertainty is carried alongside every value as a 90% interval, and router decisions use the lower bound of that interval — the safe capability — rather than the midpoint.

What the evidence grades mean

Every displayed number carries a grade, so you can see how much of the answer is measurement and how much is inference.

  • A Direct, reproducible measurement of this exact system.
  • B Trusted first-party or high-quality third-party source, configuration matches.
  • C Cross-source derived — for example a family match tested at a different effort level.
  • D Inferred or proxied — for example a task time filled in from a peer group.
  • E Synthetic assumption.

No grade-E number is allowed to drive a headline recommendation. Coverage is reported separately from the statistical interval, because “we did not measure this” and “we measured this imprecisely” are different problems and collapsing them would hide which one you are looking at.

Why the energy metric was removed

The earlier index computed energy as a fixed 350 watts multiplied by task time, then ranked on intelligence per kilojoule. Since task time was the only input that varied, that ranking was just task time restated in different units. It said nothing about how efficiently any model uses power.

It is a grade-E assumption, so it is gone from every ranking and every chart. The raw energy fields are still carried through the data file for traceability, marked E, and are never shown as a ranking column.

What this is not

The success probabilities on this site are declared priors, not calibrated measurements. Each task profile carries a difficulty threshold and a sensitivity constant, and the probability comes out of a curve fitted to those declared values. There is no production telemetry behind them yet. Calibrating them against real outcomes is a later stage of this project.

Service-level failure rates — API outages, protocol errors, broken tool calls — have no data source here. They are fixed at 1.0 and reported as not measured rather than quietly folded into the numbers. There is no trained router, no LLM-as-judge success signal, no subscription-effective pricing, and no universal score. This is a transparent frontier, not a leaderboard.