What the Latest GPT-4o Update Actually Changed — And What It Didn't

The release notes list improvements across the board. Testing against them finds gains in two areas and no measurable change in most of the rest.

Placeholder — written to give the site structure before launch. This is not reporting and it is not a finished article. It must be replaced with commissioned work before AI News Round goes live.

Release notes for model updates have settled into a genre: a list of improved capabilities, no baseline, and no methodology. This one follows the pattern, so we tested the claims that could be tested.

Two changes hold up clearly. Structured output adherence improved — malformed JSON in constrained-output tasks dropped substantially, which matters for anyone building on the API rather than chatting with it. Latency on short prompts fell noticeably.

What did not measurably change

On reasoning tasks drawn from outside the common benchmark sets, we could not distinguish the new version from the previous one at any confidence worth reporting. Hallucination rates on factual recall were similar. Long-context retrieval past roughly half the advertised window degraded in the same way it did before.

Why this matters

Not because the update is bad — the structured-output improvement is genuinely useful. It matters because the gap between what a release note implies and what changes in practice is the gap most AI coverage never closes.

Get the next one by email