THE BENCHMARK DESK

Best at what?

Reasoning, writing, building websites and generating media are different jobs. Here’s how the leaders compare in each.

Published third-party results, not tests run by Next Interrupt. We check for updates daily when this section is visited. Retrieval dates describe our snapshot, not when every model was tested. A failed refresh leaves a dated saved snapshot.

Artificial Analysis

Reasoning & knowledge

Retrieved 2026-10-04

A composite of reasoning and knowledge evaluations. Higher scores are better; these are index points, not percent correct.

  1. Claude Opus 5.5 (max with fallback)Intelligence Index
    58
  2. Claude Sonnet 5.5 (max with fallback)Intelligence Index
    56
  3. Claude Opus 5.5 (xhigh with fallback)Intelligence Index
    56
  4. Claude Opus 5.5 (high with fallback)Intelligence Index
    54
  5. Claude Fable 5.1 (max with fallback)Intelligence Index
    53

Effort settings and fallback runs affect results. An aggregate score can hide weaknesses on your particular task.

Full results & methodology at Artificial Analysis ↗

Arena

Writing & conversation

Retrieved 2026-10-04

People compare anonymous model responses and choose the answer they prefer.

  1. gemini-4-argon-highPreliminary · 4,932 votes
    1525±9
  2. claude-opus-4-6-high77,636 votes
    1505±3
  3. claude-fable-5-high38,387 votes
    1504±4
  4. claude-opus-5.5-high4,552 votes
    1504±9
  5. claude-opus-4-7-high64,946 votes
    1501±4

Preference is not a factual-accuracy score. Gemini 4 Argon has a restricted rollout; a leaderboard listing does not guarantee access. Uncertainty ranges overlap for many entries; their order is not proof of a meaningful difference.

Full results & methodology at Arena ↗

Arena

Web development

Retrieved 2026-10-04

Head-to-head preferences for generated web applications.

  1. claude-opus-5.5-max2,062 votes
    1815+16/-16
  2. gpt-6-astra-max6,123 votes
    1788+10/-10
  3. claude-sonnet-5.5-xhigh1,531 votes
    1786+18/-18
  4. gpt-6.1-sol-max1,620 votes
    1758+17/-17
  5. claude-fable-5.1-max6,318 votes
    1749+10/-10

This measures web-building output, not every kind of coding, security or maintenance work. Uncertainty ranges overlap for many entries; their order is not proof of a meaningful difference.

Full results & methodology at Arena ↗

Arena

Image generation

Retrieved 2026-10-04

Human preferences between images generated from prompts.

  1. gpt-image-2.5-sunburstPreliminary · 10,884 votes
    1424±8
  2. gpt-image-2.5-flarePreliminary · 10,058 votes
    1401±8
  3. gpt-image-2 (medium)88,744 votes
    1383±4
  4. mai-image-2.618,625 votes
    1335±6
  5. grok-imagine-image-2.0 (canvas)Preliminary · 3,087 votes
    1335±12

A high rating does not guarantee accurate text, product details or permission to use an image commercially. Uncertainty ranges overlap for many entries; their order is not proof of a meaningful difference.

Full results & methodology at Arena ↗

Arena

Video generation

Retrieved 2026-10-04

Human preferences between generated video clips.

  1. gemini-omni-1.1-flash1,784 votes
    1516±15
  2. gemini-omni-flash26,576 votes
    1513±9
  3. flux-3-videoPreliminary · 1,302 votes
    1493±17
  4. grok-imagine-video-1.5-agentPreliminary · 1,218 votes
    1492±18
  5. dreamina-seedance-2.0-720p56,318 votes
    1479±8

Compare duration, resolution and audio separately. Early results can move substantially as votes accumulate. Uncertainty ranges overlap for many entries; their order is not proof of a meaningful difference.

Full results & methodology at Arena ↗

Read the result, not just the rank.

Keep settings attached

“High,” “max,” tool access and fallback runs can change both results and cost. They are part of the tested configuration.

Watch uncertainty

A small rating gap with overlapping ranges is weak evidence of a clear winner. Preliminary entries may have fewer votes.

Compare within a board

An Arena rating of 1,500 cannot be compared with an Intelligence Index of 50. The scales and underlying tasks differ.

Try your own workload

Use known answers, count retries and compare time to a usable result. Long context is capacity, not a guarantee of accurate recall.

Compare access and API costs → · A practical benchmark guide →