Economy
GPT-5.6 Luna (high)
OpenAI
- ECST
- $0.041
- ETST
- 120s
- Success
- High
- Capability
- 71.4
Evidence: High · coverage 86%
F1Expected cost and time per successful task, on a real Pareto frontier.
Artificial Analysis as of 2026-08-26
LiveBench as of 2026-06-25
Pick the job, not the model. The profiles mirror how AI is actually used in 2026 — coding agents, support, writing, research — and everything below is recomputed in your browser.
One pick per mode. Each mode states a different preference; none of them is a universal score.
Economy
OpenAI
Evidence: High · coverage 86%
F1Balanced
Anthropic
Evidence: High · coverage 91%
F1Frontier
Evidence: Medium · coverage 62%
F2Select a row or a recommendation to see how that pick was reached.
Nothing selected yet. Pick a system in the cards above or the chart below, and this panel will show the capability interval, the success prior, the budget constraints it passed, and the systems it robustly dominates.
X = expected cost per successful task (log). Y = safe capability. Bubble size ranks throughput. All three point the same way: up, left and larger are better.
239 systems match the current filters. A row is one system configuration, not a model family.
| Model | Org | Capability | Success (modelled) | ECST | ETST | $/task | Time | Evidence | Frontier |
|---|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5 | Anthropic | 74.8 | Very high | $0.062 | 94s | $0.058 | 87 s | B | F1 |
| GPT-5.6 Luna (high) | OpenAI | 71.4 | High | $0.041 | 120s | $0.036 | 105 s | B | F1 |
| Gemini 3.1 Pro (max) | 81.2 | High | $0.194 | 230s | $0.186 | 222 s | C | F2 | |
| Llama 5 405B | Meta | 63.9 | Mixed | $0.088 | 180s | $0.062 | 125 s | D | F3 |
| Mistral Large 4 | Mistral | 58.3 | Below floor | $0.147 | 320s | $0.076 | 165 s | D | D |
Every number on this page, and how it was produced.
The twelve profiles are not benchmark categories restated — they mirror how AI is actually used as of August 2026. OpenAI's usage study (How People Use ChatGPT, NBER w34255) finds writing is the most common work task and grounded question-answering dominates consumer usage; Anthropic's Economic Index reports show software engineering at roughly half of enterprise API traffic, with “modifying software to correct errors” the single most common task. The profile order follows that observed workload share: coding agents first, then support and operations, content, and knowledge work.
Each profile declares what one task means (“one resolved ticket reply”, “one merged fix with tests passing”), a weighting over the nine measured capability dimensions, a context floor, and a minimum acceptable success rate. Those declarations are judgment calls, published so you can disagree with them.
Every row is a system, which means a model plus a reasoning level plus a
serving route — not a model family. GPT-5.6 Luna (max) and
GPT-5.6 Luna (high) are two different rows, because they cost different
amounts, take different amounts of time, and succeed at different rates.
This matters because the thing you actually deploy is a configuration, not a brand. A benchmark table that collapses every effort level into one line hides the choice you are really making. Where a capability score was measured at a different effort level than the row it appears on, the row is graded lower and says so.
The headline metric is ECST: the price per task divided by the probability that the task actually succeeds. Its sibling ETST does the same for wall-clock time. A model that costs half as much but fails twice as often is not cheaper, and this arithmetic says so without any tuning.
The previous version of this index ranked on Q^5 / cost^0.4. Those exponents
were a value judgment written into a formula: they decided, for everyone, how much
capability is worth. ECST needs no exponent. Failures pay for themselves in the numerator,
so a cheap unreliable system rises in expected cost on its own.
Raw scores from different benchmarks are not comparable to each other. Within each dimension, a system's raw score is converted to its rank within the group of systems that have data for that dimension, then mapped onto a scale with mean 50 and standard deviation 15.
So 65 does not mean “65% correct”. It means roughly one standard deviation above the cohort. Never read these numbers as percentages. Uncertainty is carried alongside every value as a 90% interval, and router decisions use the lower bound of that interval — the safe capability — rather than the midpoint.
Every displayed number carries a grade, so you can see how much of the answer is measurement and how much is inference.
No grade-E number is allowed to drive a headline recommendation. Coverage is reported separately from the statistical interval, because “we did not measure this” and “we measured this imprecisely” are different problems and collapsing them would hide which one you are looking at.
The earlier index computed energy as a fixed 350 watts multiplied by task time, then ranked on intelligence per kilojoule. Since task time was the only input that varied, that ranking was just task time restated in different units. It said nothing about how efficiently any model uses power.
It is a grade-E assumption, so it is gone from every ranking and every chart. The raw energy fields are still carried through the data file for traceability, marked E, and are never shown as a ranking column.
The success probabilities on this site are declared priors, not calibrated measurements. Each task profile carries a difficulty threshold and a sensitivity constant, and the probability comes out of a curve fitted to those declared values. There is no production telemetry behind them yet. Calibrating them against real outcomes is a later stage of this project.
Service-level failure rates — API outages, protocol errors, broken tool calls — have no data source here. They are fixed at 1.0 and reported as not measured rather than quietly folded into the numbers. There is no trained router, no LLM-as-judge success signal, no subscription-effective pricing, and no universal score. This is a transparent frontier, not a leaderboard.