Is Claude Opus 5.5 better than Fable 5.1?

Claude Opus 5.5 came out yesterday afternoon, and the number that matters isn't a score. It's a price. A million tokens (those bits of text that models bill by the piece) goes from $5 to $4 for input, from $25 to $20 for output, and Anthropic says it costs 40% less to run over an ordinary workday.

This isn't the first price cut of the year, I talk about it here almost every month. But this one doesn't come from the same place as the others. Usually, it's the smaller models that get sold off cheap to attract people. Here, it's Anthropic's top-of-the-line model arriving cheaper than the model it replaces, which hadn't happened in a long time.

And that same afternoon, OpenAI was releasing GPT-6 Sol and GPT-6 Luna, two models presented as stripped-down versions of Astra with prices cut in half. Two labs, two price cuts, the same Tuesday. It looks like two supermarkets putting up the same promotion on the same day without having coordinated.

The four tests, and what they tell us

Anthropic publishes its in-house chart. I redrew it with the figures exactly as they are, keeping the three models that matter today: Opus 5.5, Fable 5.1, and Opus 5, the one it replaces.

Horizontal bar chart comparing Claude Opus 5.5, Fable 5.1 and Claude Opus 5 on four tests

Four tests, three models, all figures as percentages. And they all come from the same page, the manufacturer's page

Translation, because the names of the tests don't mean anything to anyone. Terminal-Bench is finding your way around a command line, with no interface and nobody there to catch you. CursorBench is writing code in an editor when the project already exists. FrontierCode is the same thing on a large repository, the kind where you have to take everything else into account before touching a line. AutomationBench is chaining tasks together in place of a human: open a piece of software, click, fill out a form, start over.

Bottom line in one sentence: Opus 5.5 beats Fable 5.1 everywhere in this chart, and it crushes its predecessor where I work every day, 66.4% versus 52.3% on the terminal. On chaining tasks together in place of a human, the gap is in the same ballpark, 40% versus 26.9%.

Except there are two lines in Anthropic's chart that nobody highlights. OpenAI's GPT-6 Astra is still ahead on two tests: 64.6% versus 58.7% on a scientific version of command-line work, and 41.4% versus 40.0% precisely on chaining tasks together. Anthropic loses there, and says so in its own chart, at the bottom. It's more honest than what most manufacturers do, but it's still at the bottom.

What it looks like on a real project

The tests are all very nice. What I look at first is what the model does when you give it a big job and go to bed. Anthropic says 680,000 lines of code migrated in less than a day, and 39 successes out of 40 on a project to reduce loading times. These are company figures, nobody else measured them, I give them for what they are.

There's a more interesting figure in their internal audit, and it's almost invisible: about 85% fewer attempts to bypass instructions compared with Opus 5, across nearly 2,000 tested scenarios. In other words, the model tries much less often to do what you didn't ask it to do. It doesn't look like much. It's exactly what changes the night of someone who leaves an agent running on its own while they sleep.

A laptop left open on a kitchen table in the middle of the night, screen on in a sleeping house

Three in the morning, the house is asleep, and there's a project carrying on by itself on this laptop. This is the night when the 85% figure pays for itself

The setting that can no longer be set

A change went unnoticed, and I think it deserves a line. Thinking mode can no longer be disabled. Before, you could ask the model to answer without thinking first, to go faster on a simple question. That button no longer exists.

Alongside that, there's now a fast mode, billed at a higher rate: $8 for input and $40 for output, advertised as up to 2.5 times faster. The message is clear, speed costs money. I don't know what it really feels like in use yet, and I'd rather say so than sell you a comfort I haven't tried.

What it changes in your day

Three things, and the first is immediate.

If you pay for a subscription to an assistant, you pay the same price for a better engine. Your bill doesn't change, it's the contents of the plan that change, and it feels good to see things going in that direction for once.

If it's your company that's paying, it matters more than it seems. The “artificial intelligence” line in budgets is becoming a line people keep an eye on, like the electricity bill. A top-tier model that costs 40% less to run ends up being noticed somewhere, and not in a bad way.

And there's something I didn't see coming when I started running three coding assistant sessions in parallel. It's not the price that's worrying me, it's the time I spend on it. The percentage of quota dropping in my statusline keeps me on the edge of my seat like a video game counter: “I've got 12% left before it resets, come on, full speed ahead, I'll start one last project before the reset!” I should learn to close the window, one of these days.

What you shouldn't believe

The figures in the chart come from the manufacturer's page, and it's a vendor's chart. That doesn't mean they're false, it means nobody else has reproduced them yet. The only figure of the day that you can check yourself in thirty seconds is the price, and it's public.

And no, this isn't the end of small models, quite the opposite. Anthropic is announcing Sonnet 5.5 and Haiku 5.5 in the coming weeks. Those are the ones most of us use without knowing it, in a browser tab or in the phone app. The interesting fight this quarter will be there, not at the top of the chart.

A word about the price differences, by the way, because I feel like we're getting used to them too quickly. A model that announces almost the top score for a fraction of the price, I did the calculation a month and a half ago, and the question hasn't changed: what matters isn't the score at the top of the chart, it's what you get for what you pay. And two months ago, I was already saying that July was going off the rails on that front. Since then, nobody has slowed down.

Bar chart of the price per million tokens of the three models, for input and output

The only chart of the day that you can check yourself in thirty seconds, and the only one where you don't have to take my word for it

What I take away from this week is the pace. Grok 4.7 and MiMo the day before yesterday, Opus 5.5 and two OpenAI models yesterday, and it's only Wednesday. The model I find extraordinary this morning will be the normal model next month, and I'll have to learn to choose instead of waiting. Good luck to those keeping a comparison up to date, they haven't finished sweating.

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙