Twelve months of releases, and the honest summary is narrower than the announcements suggest. Context windows are dramatically longer. Inference costs have fallen far enough to change what is economically viable to build. On the reasoning benchmarks that vendors lead their launches with, the gains are real but incremental, and the gap between the leading models has narrowed rather than widened.
The more consequential shift is not capability at the frontier. It is that models good enough for most production work are now cheap and, in several cases, open-weight. A team that needed a frontier API eighteen months ago can often now run something adequate themselves.
Where the numbers mislead
Benchmark scores continue to be reported without error bars and without contamination checks, which makes small differences between models largely meaningless. Several widely cited evaluations have appeared in training data. Treat a two-point difference on a leaderboard as noise unless the methodology is published alongside it.
What actually got better
Tool use and structured output improved markedly, which matters more for real deployments than raw reasoning scores. Long-context retrieval got more reliable, though degradation over very long inputs remains real and under-reported. Multilingual performance improved for high-resource languages and barely moved for everything else — a gap covered in our reporting on African languages.