AE43E69367F40FE1DB5F244507269D43

Benchmarks Are Not Progress

A number that improves every quarter is reassuring and largely decorative. Measuring what a system is good at is much harder than ranking it.

AI benchmark dashboard on a laptop showing rising model scores and rankings, contrasted with notes emphasizing real-world usefulness, reliability, safety, and impact.
Benchmarks Are Not Progress. Higher AI benchmark scores may look impressive, but they don't always mean AI systems are becoming more useful in the real world. The harder question is: What is the system actually good at?

Benchmark scores are the currency of AI progress reporting, for the understandable reason that they are the only comparable numbers available. They also mislead in ways that are well known to the people producing them and rarely mentioned in the coverage.

Contamination

Benchmarks are published on the internet. Models are trained on the internet. A test set that has leaked into training data measures memorisation, and nobody outside the training organisation can verify whether it has. This is not a hypothetical concern; it is a structural feature of how these systems are built.

Optimising for the measure

When a benchmark becomes the way capability is reported, effort flows towards it. That is not cheating, it is what measurement does to any field. It does mean that a rising score can reflect increasing attention to the test rather than increasing general capability — and the two are indistinguishable from the outside.

The tasks are not the work

Multiple-choice questions and self-contained coding puzzles are chosen because they can be graded automatically, not because they resemble what anyone actually does. Real tasks are ambiguous, involve incomplete information, span long contexts, and are judged by whether a person was helped. Almost none of that is captured by anything on a leaderboard.

What would be better

Held-out evaluations nobody can train on. Measurement on tasks that people genuinely perform, graded by the people who perform them. Reporting the distribution of outcomes rather than an average, since the tail is where harm lives. And, for anyone deploying, an evaluation set built from their own data — which will tell them more than every public benchmark combined.

Get the next one by email

AI in Africa

AI Didn't Create Africa's Fraud Problem. It Sped It Up

INTERPOL says AI is now linked to 55% of reported cybercrime across Africa, with losses more than doubling to $484 million in a single year. The real story isn't that AI is new to fraud — it's that everything else was already too slow to keep up with it.

3 min read