How to Read a Model Release Without Being Sold To

Every launch comes with a chart showing the new model winning. Here is what the chart leaves out, and which parts of the announcement are worth your attention.

A model release follows a template now. There is a name, a chart with bars, a table of benchmark scores against competitors, and a set of examples chosen by the people who built it. Read enough of them and the shape becomes predictable — which is itself useful, because it tells you where to look.

The benchmark table is the least informative part

Public benchmarks are the only common currency available, so everyone reports them. They also leak into training data over time, they are chosen by the party reporting them, and a small difference in score frequently reflects prompt formatting rather than capability. A model that scores two points higher on a reasoning benchmark may be no better at anything you actually want to do.

What is worth reading instead: the context window, the cost per token, the latency, and whether the weights are available. Those are unambiguous, they are hard to present misleadingly, and they decide what you can build.

"Open" is doing a lot of work

The word covers everything from a permissive licence with published weights to a bespoke agreement that restricts commercial use above a certain size, forbids training other models, and does not include the training data. All get called open. The licence text is short and it is the only thing that tells you what you may actually do.

The system card is where the caveats live

Most serious releases ship with a document describing evaluations, known failure modes and limitations. It is the least-read part of any launch and by some distance the most useful, because it is the only place where the organisation states plainly what its model is bad at.

The question to end on

Not "is this the best model" but "does this change what I can build". Most releases do not. The ones that do usually change a constraint — price, context length, an ability to run on hardware you own — rather than a leaderboard position.

Get the next one by email