Release notes for model updates have settled into a genre: a list of improved capabilities, no baseline, and no methodology. This one follows the pattern, so we tested the claims that could be tested.
Two changes hold up clearly. Structured output adherence improved — malformed JSON in constrained-output tasks dropped substantially, which matters for anyone building on the API rather than chatting with it. Latency on short prompts fell noticeably.
What did not measurably change
On reasoning tasks drawn from outside the common benchmark sets, we could not distinguish the new version from the previous one at any confidence worth reporting. Hallucination rates on factual recall were similar. Long-context retrieval past roughly half the advertised window degraded in the same way it did before.
Why this matters
Not because the update is bad — the structured-output improvement is genuinely useful. It matters because the gap between what a release note implies and what changes in practice is the gap most AI coverage never closes.