Keep settings attached
“High,” “max,” tool access and fallback runs can change both results and cost. They are part of the tested configuration.
THE BENCHMARK DESK
Reasoning, writing, building websites and generating media are different jobs. Here’s how the leaders compare in each.
Published third-party results, not tests run by Next Interrupt. We check for updates daily when this section is visited. Retrieval dates describe our snapshot, not when every model was tested. A failed refresh leaves a dated saved snapshot.
Artificial Analysis
A composite of reasoning and knowledge evaluations. Higher scores are better; these are index points, not percent correct.
Effort settings and fallback runs affect results. An aggregate score can hide weaknesses on your particular task.
Full results & methodology at Artificial Analysis ↗Arena
People compare anonymous model responses and choose the answer they prefer.
Preference is not a factual-accuracy score. Gemini 4 Argon has a restricted rollout; a leaderboard listing does not guarantee access. Uncertainty ranges overlap for many entries; their order is not proof of a meaningful difference.
Full results & methodology at Arena ↗Arena
Head-to-head preferences for generated web applications.
This measures web-building output, not every kind of coding, security or maintenance work. Uncertainty ranges overlap for many entries; their order is not proof of a meaningful difference.
Full results & methodology at Arena ↗Arena
Human preferences between images generated from prompts.
A high rating does not guarantee accurate text, product details or permission to use an image commercially. Uncertainty ranges overlap for many entries; their order is not proof of a meaningful difference.
Full results & methodology at Arena ↗Arena
Human preferences between generated video clips.
Compare duration, resolution and audio separately. Early results can move substantially as votes accumulate. Uncertainty ranges overlap for many entries; their order is not proof of a meaningful difference.
Full results & methodology at Arena ↗“High,” “max,” tool access and fallback runs can change both results and cost. They are part of the tested configuration.
A small rating gap with overlapping ranges is weak evidence of a clear winner. Preliminary entries may have fewer votes.
An Arena rating of 1,500 cannot be compared with an Intelligence Index of 50. The scales and underlying tasks differ.
Use known answers, count retries and compare time to a usable result. Long context is capacity, not a guarantee of accurate recall.
Compare access and API costs → · A practical benchmark guide →