Claude Opus 4.5, and why evaluating new LLMs is increasingly difficult
24th November 2025
Anthropic released Claude Opus 4.5 this morning, which they call âbest model in the world for coding, agents, and computer useâ. This is their attempt to retake the crown for best coding model after significant challenges from OpenAIâs GPT-5.1-Codex-Max and Googleâs Gemini 3, both released within the past week!
The core characteristics of Opus 4.5 are a 200,000 token context (same as Sonnet), 64,000 token output limit (also the same as Sonnet), and a March 2025 âreliable knowledge cutoffâ (Sonnet 4.5 is January, Haiku 4.5 is February).
The pricing is a big relief: $5/million for input and $25/million for output. This is a lot cheaper than the previous Opus at $15/$75 and keeps it a little more competitive with the GPT-5.1 family ($1.25/$10) and Gemini 3 Pro ($2/$12, or $4/$18 for >200,000 tokens). For comparison, Sonnet 4.5 is $3/$15 and Haiku 4.5 is $1/$5.
The Key improvements in Opus 4.5 over Opus 4.1 document has a few more interesting details:
- Opus 4.5 has a new effort parameter which defaults to high but can be set to medium or low for faster responses.
- The model supports enhanced computer use, specifically a
zoomtool which you can provide to Opus 4.5 to allow it to request a zoomed in region of the screen to inspect. - "Thinking blocks from previous assistant turns are preserved in model context by default"âapparently previous Anthropic models discarded those.
I had access to a preview of Anthropicâs new model over the weekend. I spent a bunch of time with it in Claude Code, resulting in a new alpha release of sqlite-utils that included several large-scale refactoringsâOpus 4.5 was responsible for most of the work across 20 commits, 39 files changed, 2,022 additions and 1,173 deletions in a two day period. Hereâs the Claude Code transcript where I had it help implement one of the more complicated new features.
Itâs clearly an excellent new model, but I did run into a catch. My preview expired at 8pm on Sunday when I still had a few remaining issues in the milestone for the alpha. I switched back to Claude Sonnet 4.5 and... kept on working at the same pace Iâd been achieving with the new model.
With hindsight, production coding like this is a less effective way of evaluating the strengths of a new model than I had expected.
Iâm not saying the new model isnât an improvement on Sonnet 4.5âbut I canât say with confidence that the challenges I posed it were able to identify a meaningful difference in capabilities between the two.
This represents a growing problem for me. My favorite moments in AI are when a new model gives me the ability to do something that simply wasnât possible before. In the past these have felt a lot more obvious, but today itâs often very difficult to find concrete examples that differentiate the new generation of models from their predecessors.
Googleâs Nano Banana Pro image generation model was notable in that its ability to render usable infographics really does represent a task at which previous models had been laughably incapable.
The frontier LLMs are a lot harder to differentiate between. Benchmarks like SWE-bench Verified show models beating each other by single digit percentage point margins, but what does that actually equate to in real-world problems that I need to solve on a daily basis?
And honestly, this is mainly on me. Iâve fallen behind on maintaining my own collection of tasks that are just beyond the capabilities of the frontier models. I used to have a whole bunch of these but theyâve fallen one-by-one and now Iâm embarrassingly lacking in suitable challenges to help evaluate new models.
I frequently advise people to stash away tasks that models fail at in their notes so they can try them against newer models later onâa tip I picked up from Ethan Mollick. I need to double-down on that advice myself!
Iâd love to see AI labs like Anthropic help address this challenge directly. Iâd like to see new model releases accompanied by concrete examples of tasks they can solve that the previous generation of models from the same provider were unable to handle.
âHereâs an example prompt which failed on Sonnet 4.5 but succeeds on Opus 4.5â would excite me a lot more than some single digit percent improvement on a benchmark with a name like MMLU or GPQA Diamond.
In the meantime, Iâm just gonna have to keep on getting them to draw pelicans riding bicycles. Hereâs Opus 4.5 (on its default âhighâ effort level):

It did significantly better on the new more detailed prompt:

Hereâs that same complex prompt against Gemini 3 Pro and against GPT-5.1-Codex-Max-xhigh.
Still susceptible to prompt injection
From the safety section of Anthropicâs announcement post:
With Opus 4.5, weâve made substantial progress in robustness against prompt injection attacks, which smuggle in deceptive instructions to fool the model into harmful behavior. Opus 4.5 is harder to trick with prompt injection than any other frontier model in the industry:
On the one hand this looks great, itâs a clear improvement over previous models and the competition.
What does the chart actually tell us though? It tells us that single attempts at prompt injection still work 1/20 times, and if an attacker can try ten different attacks that success rate goes up to 1/3!
I still donât think training models not to fall for prompt injection is the way forward here. We continue to need to design our applications under the assumption that a suitably motivated attacker will be able to find a way to trick the models.
More recent articles
- Claude Haiku 5.5 - 7th October 2026
- We're going to need default hard budget caps on pretty much everything - 3rd October 2026
- OpenAI DevDay 2026 live blog - 29th September 2026
