Every model launch arrives with the same breathless takes, so we did the boring thing instead: we pointed the newest one at a week of the work we actually do and kept notes.
The honest read is that it's meaningfully better at two things we care about, and no different at a third that everyone claims it fixed. We'll take the two. This is a working theory, and next month's model may move the line again, but here's what held up after the novelty wore off.
What got better
Longer context that it actually uses, and fewer confident wrong answers on the kind of structured business data we hand it every day. Both matter more to a real workflow than the demo everyone shared.
What didn't
It still needs a person deciding what "right" looks like. That hasn't changed, and we'd argue it isn't supposed to.