
Two flagship models arrived within days of each other, and every channel we follow spent the week crowning a winner. The question that reaches us from business owners is simpler and more practical: which one should we be paying for?
So here is the honest version — not a benchmark table, but what we actually reach for, job by job, and why the answer matters less than the volume of argument suggests.
What we actually use
For hard thinking, Claude Fable 5.1 is the brain. Anything that needs a problem held whole — architecture, a tricky piece of judgement, writing that has to carry a voice — starts there. We still reach for Opus 5, and sometimes Opus 4.8, depending on the work; the older model is not obsolete just because a newer one shipped.
For the daily grind, GPT-5.6 Sol does an enormous amount of work here, every single day. It is particularly good at operating a computer — moving between applications, handling the mundane writing that piles up, working through the small tasks nobody enjoys — and it is a strong strategist and programmer besides. GPT-6 Astra is excellent, and we reach for it when a job genuinely needs the frontier; Sol is what carries the week.
For directed programming where cost matters, Grok 4.6 earns its place. It follows instructions closely and gets a lot done. It is not the most creative of the group, and for that kind of work that is exactly the point.
And when a job doesn't need frontier-level thinking at all, Claude Sonnet is still very good. A surprising share of day-to-day work falls into that bucket, and using a smaller, faster model for it means not burning through your allowance on tasks that never needed the expensive option.
The honest test isn't which model wins a benchmark. It's which one does your Tuesday work with less checking.
The differences are real, and narrower than the headlines
We are not going to pretend the models are interchangeable. They differ in how they hold a long problem, how they write, how carefully they follow an instruction, and how gracefully they admit uncertainty. Anyone using them daily notices.
But those differences show up at the edges of hard work. On the tasks most businesses actually need — summarising a document, drafting a reply, pulling structure out of a mess, writing a first version of something — every frontier model available today is good enough, and the gap between them is smaller than the gap between a business that has built the habit and one that hasn't.
Why the wrapper still decides the result
The reason the choice matters less than the arguments suggest is that the model is only the engine. What decides whether AI does useful work for your business is what it can see, what it is allowed to do and who checks its output — which is why we wrote that the harness matters more than the model.
A strong model with no access to your documents will guess. A modest model with the right context, a couple of tools and a check on its work will quietly get things right. Swapping the engine rarely fixes a problem that lives in the plumbing.
How to choose without reading a single review
Pick two. Give them both the same week of your real work — not puzzles, not prompts from a video, the actual jobs on your desk. Then compare on three things: how often you had to ask again to get what you wanted, how often it stated something confidently that turned out to be wrong, and whether the output needed rewriting before anyone else saw it.
That test takes a week and tells you more about your business than any leaderboard will. Run it again in six months, because the answer will have changed.
One last piece of advice: whatever you choose, keep the ability to change your mind. Build your habits and your systems so the model underneath can be swapped when a better one arrives, rather than wiring your business into one provider's engine — the same reasoning as owning the systems your business runs on. The same restraint applies to the money side: it is early for paying to appear inside an assistant, just as it is early to bet a workflow on one model's particular quirks.
If you want a second opinion on which tools fit the work your business actually does, tell us what you are trying to get done.


