AI· · 2 min read

GPT-5.6 Sol is winning? Claude's Fable 5 stuck at a velvet rope.

The 2-minute version
Black-and-white editorial drawing of two frontier AI rockets at a velvet rope
Two frontier models. One velvet rope.

Two frontier models walked into July with names that sound like interplanetary airlines.

OpenAI’s Sol landed across ChatGPT, Codex and the API. Anthropic’s Fable landed, disappeared, returned, then moved behind a velvet rope: included for Max and Team Premium at 50% of usage limits, while Pro users reach it through credits.

One launch said, “Start building.”
The other said, “Are you on the list?”

On OpenAI’s scoreboard, Sol leads Fable 52.7% to 40.5% on Agents’ Last Exam and 80.0 to 77.2 on the Artificial Analysis Coding Agent Index. Sol also costs $5/$30 per million input/output tokens, against Fable’s $10/$50.

However, Fable evades and throws a clean counterpunch 🥊

It leads SWE-Bench Pro 80.0% to 64.6% and narrowly edges Sol on GDPval-AA v2. Claude is not losing its intelligence. It is losing the convenience round.

The new additions make the rivalry more fun. Sol brings adjustable reasoning, a million-token context window, programmatic tool calling, a single unified weekly limit and multi-agent support. Fable is built for ambitious, days-long work, but some cybersecurity and biology requests can be rerouted to Opus 4.8.

Both models also arrive with fresh quirks waiting to be discovered.

A NeurIPS study found that some apparent “emergent abilities” can vanish when the measurement changes. Nature found that accuracy-focused evaluations can reward confident guessing.

So keep the team jerseys on. Keep the benchmark tattoos temporary.

Fable may be the better driver on certain roads. Sol is cheaper, easier to reach and already parked inside the tools where work happens. Right now, GPT is winning the product race.

Claude is still very much in it & dominant.

It just needs to remove the velvet rope.

Two frontier models walked into July with names that sound like interplanetary airlines. One was called Sol. The other was called Fable. If you say them quickly enough, the only thing missing is a boarding gate and someone in a blazer asking whether your bag is cabin-approved.

OpenAI’s Sol landed with the easy confidence of a product team that had already booked the billboards: ChatGPT, Codex and the API, all in one swing. Anthropic’s Fable had a more dramatic entrance. It landed, disappeared, returned, and then found itself behind a velvet rope: Max and Team Premium got it at 50% of their usual usage limits; Pro users could reach it through credits.

That contrast is the whole story. One launch said, “Start building.” The other said, “Are you on the list?” Neither line tells you which model is smarter. Both tell you which product is easier to put into a real week of work.

Start with the scorecards, because this is where the AI internet immediately starts tattooing numbers on its forearm. On OpenAI’s published comparison table, Sol leads Fable 52.7% to 40.5% on Agents’ Last Exam. It also leads 80.0 to 77.2 on the Artificial Analysis Coding Agent Index. Those are not tiny gaps. They are the sort of margins that make a founder move a default model, make a developer re-open their settings, and make every other person on X post a screenshot with “it’s over” in the caption.

Then there is the price. Sol is listed at $5 per million input tokens and $30 per million output tokens. Fable is $10 and $50. That does not mean every task costs exactly half as much; real cost depends on prompts, tool calls, caching and how much the model decides to think out loud. It does mean the direction is not subtle. If you are running lots of work, a cheaper frontier model changes the kind of work you are willing to try. The experiment that felt like a luxury becomes a button you press before lunch.

But Fable is not the model sitting in the corner, pretending not to hear the applause. It throws a clean counterpunch 🥊 On SWE-Bench Pro it leads 80.0% to Sol’s 64.6%. It also narrowly leads on GDPval-AA v2. That is a useful reminder that “best model” is a sentence missing its final clause. Best at what? Best under which harness? Best after how much steering? Best when you are fixing a gnarly repository at 11:47 p.m. and want the model to stay stubborn for one more hour?

Fable was built for that last category: ambitious, long-running work. Anthropic describes it as a model for days-long, asynchronous projects, with planning across stages, sub-agent delegation and self-checking. That is not a small claim. It is the actual dream of this category: not a chatbot that answers the next question, but an operator that notices there is a next five questions and starts doing the boring middle.

Sol’s answer is less romantic and more product-shaped. Adjustable reasoning lets you decide how much room the model gets to think. Programmatic tool calling lets it process the middle of a workflow without dragging every intermediate scrap back through the conversation. Multi-agent support turns a difficult task into parallel workstreams. A million-token context window means the room is much bigger before somebody has to repeat themselves. This is the frontier model as a well-stocked workshop: more tools on the wall, less ceremony to pick one up.

That is why I think GPT is winning the product race right now. Not because Claude forgot how to reason. Not because every benchmark points in one direction. Because Sol is cheaper, easier to reach and already parked inside the tools where a lot of people are doing the work. Distribution is a capability too. A brilliant engine behind a closed rope is still making people wait in the lobby.

The rope also has a safety-shaped knot in it. Anthropic says Fable can route some cybersecurity and biology prompts to Opus 4.8 when its safeguards flag them. There is a perfectly serious reason for that. Frontier capability in those domains comes with real risk. But users do not experience a policy document. They experience a model swap in the middle of a job. That friction matters, even when it is responsible friction. It becomes part of the product comparison, just like price or latency.

Now, before anyone orders matching jackets, remember what the research says about the things we call sudden miracles. A NeurIPS paper made the uncomfortable case that some apparent “emergent abilities” can dissolve when the measurement changes. The model may have improved; the theatrical jump may belong to the metric. A Nature paper makes a related point about accuracy-focused evaluations: if guessing is rewarded and abstaining is punished, confidence can start to look like knowledge.

That does not make benchmarks useless. It makes them more like a map than a coronation. They show terrain. They do not tell you which road you are about to drive, whether it is raining, or whether your team needs a truck, a scooter or someone who can simply stop hallucinating in a spreadsheet.

So yes: Sol is winning the convenience round. It has the reach, the price, the surfaces and enough benchmark wins to make the conversation loud. Fable still has roads where it is the better driver, including one of the most practical coding tests in the room. Claude is still very much in it & dominant.

It just needs to remove the velvet rope.