When the obvious route closes, it finds another
Anthropic is pitching a near-frontier working model that checks, adapts and keeps going when a task stops being straightforward.

In Anthropic’s FreeCAD evaluation, Opus 5 was asked to rebuild a machine part from a drawing it could not directly see. Rather than stop there, Anthropic says the model built a computer-vision pipeline, extracted the geometry and completed the model.
That episode captures the launch proposition neatly: Opus 5 is meant to find another route, inspect what it produces and keep working towards a finished handoff. It is Anthropic’s evaluation, not independent proof, but it gives us a more useful question than whether a model can produce an impressive first response.
Unfinished work costs more
A strong opening answer loses its shine when retries, repairs and supervision still sit with us.

Token price tells us what generation costs. It does not tell us what the task costs once we include failed approaches, repeated prompts, manual corrections and the time spent checking whether the result is actually usable.
That is why Anthropic’s emphasis on verification matters. If Opus 5 completes more of the loop itself, the saving may appear in fewer repairs rather than a cheaper token. Whether that happens consistently is something we still need to test on our own work.
Price is not task cost
More reasoning may cost more tokens while still reducing the total cost of reaching a usable result.


Anthropic has kept Opus 5 at the same base API price as Opus 4.8: $5 per million input tokens and $25 per million output tokens. The effort setting then lets us trade faster, cheaper runs for deeper reasoning without changing models.
Higher effort can consume more tokens, so it is not automatically the economical choice. The useful measure is the cost to finish the job: model spend, elapsed time, corrections and human intervention together. Fast mode adds another trade-off, with around 2.5 times the speed at twice the base price.
Anthropic’s evidence sets the boundary
The comparison is promising, but it remains a collection of named evaluations and early-access reports.

Anthropic describes Opus 5 as a thoughtful, proactive model that approaches Fable 5 intelligence at a lower task cost in selected comparisons. That is narrower than saying it is universally half the price, and more useful because the claim can be tested against named workloads.
The supporting material combines Anthropic-run evaluations with reports from early-access customers. It can show us where to look, but it cannot establish how often the same gains will survive different prompts, tools, data and review standards.
Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds.
Anthropic — Introducing Claude Opus 5
The benchmark shows a trade-off
The chart compares score with task cost inside one coding evaluation; it does not promise the same result everywhere.

The strongest reading is specific: on Anthropic’s Frontier-Bench v0.1 presentation, Opus 5 moves beyond Opus 4.8 on both performance and cost per completed task. CursorBench 3.2 adds a similar supplier-published signal at higher effort settings.
Neither chart gives us a universal productivity rate. They are task-bound comparisons, and the Frontier-Bench presentation includes Anthropic’s own run conditions and fallback handling. We can treat the result as a reason to test long coding work, not as a settled verdict on every workload.
Progress includes checking the work
The promised improvement is a longer loop: plan, act, inspect, revise and hand back something usable.

The FreeCAD example gives this loop a concrete shape. Faced with a drawing it could not directly inspect, the model created another way to extract the information it needed. In a separate package-manager evaluation, Anthropic says Opus 5 traced a bug to its root cause and found an edge case missed by an existing patch.
These are reported episodes rather than guarantees. Their value is in the behaviour they illustrate: the model does not have to treat a blocked route as the end of the task, and it can use inspection and revision as part of the work rather than waiting for us to notice every gap.
Plan
Inspect
Revise
Test the whole working session
Our first run should measure completion and correction, not simply whether the opening answer sounds capable.

We can start with one task that normally needs several tools, decisions and checks. Before the run, we define what finished means. During it, we record the effort setting, model spend, elapsed time, failed routes and each correction we still have to make.
The final review remains ours. Cyber safeguards may restrict or reroute some requests, including fallback to Opus 4.8 in covered cases, and a confident handoff can still be wrong. The point of the trial is to see whether Opus 5 reduces repair work without hiding cost or risk.
What to record
- Completion
- Corrections
- Latency
- Total task cost
Finished work is the value test
Near-frontier intelligence matters most when it survives the route from request to checked result.
The useful model is not the one that looks smartest at the start. It is the one that helps us reach a checked finish with less repair.
Creator Broadcast synthesis, based on Anthropic’s published evidence
Anthropic is not presenting Opus 5 as its absolute frontier. It is presenting a model close to that frontier, available at everyday Opus pricing and designed to spend effort where a longer task demands it.
That proposition becomes valuable only if the whole session improves. Completion, corrections, latency, safeguards and total cost belong in the same comparison. A real, repeatable task will tell us more than a headline score alone.
Test the cost of finished work
Act as a senior analyst completing a difficult, multi-step brief. First state a concise plan. Then produce the brief, inspect it against every stated constraint, identify any gaps, revise the work, and return both the finished version and a short verification note. Do not claim success where evidence is missing. Report the number of corrections, elapsed working time, and any points that still need human review so we can compare completion quality and total task cost across model runs.Ready to copy
Measure the checked finish
Give Opus 5 and our current model the same long-horizon task, use the same completion standard, and compare corrections, latency, total task cost and the quality of the final human-reviewed handoff.
Try the promptWhat could change the verdict
- Independent evidence
- Effort economics
- Safeguard behaviour
- Fast mode