Model capabilities

Compare models across eight dimensions, from intelligence to real-world cost and responsiveness.

01

Overall intelligence

Compare overall capability across reasoning, coding and knowledge work.

© Artificial Analysis · 2026-09-27Intelligence Index v4.3.2
Higher is better. The index combines 10 evaluations; it is a comparison score, not a percentage of tasks passed.

Chart interpretation

Opus 5.5 leads this selection at 58; Fable 5.1 and GPT-6 Astra both round to 53. Reasoning settings remain part of each result.

02

Specialist evaluations

Explore coding, reasoning, factual knowledge, long-context and visual reasoning separately.

Explore 19 benchmark charts19 evaluations · separate scales

Chart interpretation

Leaders change by task. Use Terminal-Bench and SciCode for coding, AA-LCR for long-context reasoning, and MMMU-Pro for visual reasoning.

03

Knowledge-work agents

Assess whether an agent can turn a complex brief into a useful deliverable.

© Artificial Analysis · 2026-09-27AA-Briefcase v1.1 · Elo
Higher is better. AA-Briefcase combines rubric pass rate, analytical quality and presentation. Agent tools and execution settings affect results.

Chart interpretation

Opus 5.5 reaches 1822 Elo, followed by Fable 5.1 at 1678 and Grok 4.7 at 1657 in this selection.

04

Intelligence and cost

Find the capability you need at a sustainable cost per task.

© Artificial Analysis · 2026-09-27Intelligence Index · USD/task · log scale
Costs are weighted benchmark task costs, including token usage. The horizontal axis is logarithmic. These are not this site’s API prices.

Chart interpretation

The upper-left region combines stronger capability with lower cost. The dotted frontier helps identify efficient choices in the selected models.

05

Token efficiency

See how much reasoning and answer output a benchmark task consumes.

© Artificial Analysis · 2026-09-27Weighted output tokens/task
Lower usage can reduce cost and waiting time, but does not establish better quality. Compare this chart with intelligence and token prices.

Chart interpretation

Opus 5.5 uses about 119k output tokens per task, versus 31k for GPT-6 Sol. Reasoning tokens account for a large share of usage.

06

Context capacity

Compare the capacity available for documents, code and conversation history.

© Artificial Analysis · 2026-09-27Context window · tokens
A larger window is a capacity limit, not proof of better comprehension. Check AA-LCR results and the provider’s input/output limits for your task.

Chart interpretation

Many models in this selection offer around 1M tokens of context; Grok 4.7 is shown at 500k and Mistral Medium 3.5 at 256k.

07

Output speed

Compare how quickly text arrives once generation has started.

© Artificial Analysis · 2026-09-27Output tokens/second · 10k input · single query
Higher is better. Generation speed excludes the wait before output begins; it is not the total time to finish a task or a throughput guarantee for this site.

Chart interpretation

Gemini 3.5 Flash-Lite reaches 337 tokens/s and Gemini 3.8 Flash at high effort reaches 285 tokens/s in this selection. Faster output can improve long-response delivery.

08

Time to first answer

Measure the wait before the user receives the first answer token.

© Artificial Analysis · 2026-09-27Seconds · includes thinking · 10k input · single query
Lower is better. This metric includes reasoning time and differs from time to first token. Compare like-for-like effort settings and API providers.

Chart interpretation

Grok 4.7 at xhigh effort is shown at 54.4 s, while GPT-6 Astra at max effort takes 292.2 s. Thinking time can dominate the perceived wait.

Charts preserve the source’s model selection and reasoning settings. Missing models or measurements are not zero scores.

Original charts: Artificial Analysis. Interpretations on this page are editorial guidance; verify important choices on your own tasks.