Maisa Korhonen
← Back to glossary

Benchmark

A standardised test used to score AI models, so they can be compared on the same tasks.

When a new model launches, the announcement always includes charts: our model scored this on maths, that on coding, this much on reasoning. Those scores come from benchmarks, shared test sets that every model can be run against. Leaderboards rank models by them, and "state of the art" claims usually mean "top score on some benchmark".

Treat the charts with the same scepticism you would treat any vendor's own slides. Models can be tuned to shine on famous benchmarks (the AI version of teaching to the test), tests leak into training data, and a win on a graduate-level maths benchmark says little about whether the model writes good product copy.

Why you keep hearing it

Because benchmarks are how the industry keeps score, and every release cycle produces a new round of charts, arguments about fairness, and claims of leapfrogging.

What it means for you

Benchmarks answer "which model is impressive". Your question is "which model is good at my work", and no public leaderboard measures your work. The practical move: build a tiny benchmark of your own, five real tasks from your week, and run them through any model you are considering. Twenty minutes, and more informative than every chart in the launch post.

Updated 12 July 2026