>Claude Opus 5: the model that keeps going
ALTIOR AI ADVANTAGEWhat to remember
Claude Opus 5

When the obvious route closes, it finds another

Anthropic is pitching a near-frontier working model that checks, adapts and keeps going when a task stops being straightforward.

Cinematic neon work path rerouting around a barrier through tools, inspection, correction and a human-checked handoff.

In Anthropic’s FreeCAD evaluation, Opus 5 was asked to rebuild a machine part from a drawing it could not directly see. Rather than stop there, Anthropic says the model built a computer-vision pipeline, extracted the geometry and completed the model.

That episode captures the launch proposition neatly: Opus 5 is meant to find another route, inspect what it produces and keep working towards a finished handoff. It is Anthropic’s evaluation, not independent proof, but it gives us a more useful question than whether a model can produce an impressive first response.

The hidden bill

Unfinished work costs more

A strong opening answer loses its shine when retries, repairs and supervision still sit with us.

Neon comparison contrasting a fast first answer with retries, corrections, supervision and checked finished work.
The useful comparison runs from an impressive first answer to completed, checked work, with retries, corrections and human supervision counted along the way.

Token price tells us what generation costs. It does not tell us what the task costs once we include failed approaches, repeated prompts, manual corrections and the time spent checking whether the result is actually usable.

That is why Anthropic’s emphasis on verification matters. If Opus 5 completes more of the loop itself, the saving may appear in fewer repairs rather than a cheaper token. Whether that happens consistently is something we still need to test on our own work.

The effort dial

Price is not task cost

More reasoning may cost more tokens while still reducing the total cost of reaching a usable result.

Neon infographic showing an effort dial from light through balanced to deep reasoning, with total task cost separated from token price.
Lower effort favours speed and token conservation; higher effort gives the model more room to reason. The separate measure is what we spend before the whole job is complete.
$5/Minput tokens at the published base API price
$25/Moutput tokens at the published base API price
base price for Fast mode
Provider image: supplier-01.png
Anthropic’s CursorBench 3.2 chart compares coding score with task cost across effort settings; Anthropic says Opus 5 at maximum effort comes within 0.5% of Fable 5’s peak score at half the task cost.

Anthropic has kept Opus 5 at the same base API price as Opus 4.8: $5 per million input tokens and $25 per million output tokens. The effort setting then lets us trade faster, cheaper runs for deeper reasoning without changing models.

Higher effort can consume more tokens, so it is not automatically the economical choice. The useful measure is the cost to finish the job: model spend, elapsed time, corrections and human intervention together. Fast mode adds another trade-off, with around 2.5 times the speed at twice the base price.

The source receipt

Anthropic’s evidence sets the boundary

The comparison is promising, but it remains a collection of named evaluations and early-access reports.

Provider image: supplier-02.png
Anthropic’s comparison table positions Opus 5 strongly across coding and knowledge-work evaluations while showing that performance varies by task and benchmark.

Anthropic describes Opus 5 as a thoughtful, proactive model that approaches Fable 5 intelligence at a lower task cost in selected comparisons. That is narrower than saying it is universally half the price, and more useful because the claim can be tested against named workloads.

The supporting material combines Anthropic-run evaluations with reports from early-access customers. It can show us where to look, but it cannot establish how often the same gains will survive different prompts, tools, data and review standards.

Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds.

Anthropic — Introducing Claude Opus 5
Read the axes

The benchmark shows a trade-off

The chart compares score with task cost inside one coding evaluation; it does not promise the same result everywhere.

Provider image: supplier-03.png
Anthropic’s Frontier-Bench v0.1 chart plots agentic coding performance against task cost. Anthropic says Opus 5 more than doubles Opus 4.8’s performance at a lower cost per task on this evaluation.

The strongest reading is specific: on Anthropic’s Frontier-Bench v0.1 presentation, Opus 5 moves beyond Opus 4.8 on both performance and cost per completed task. CursorBench 3.2 adds a similar supplier-published signal at higher effort settings.

Neither chart gives us a universal productivity rate. They are task-bound comparisons, and the Frontier-Bench presentation includes Anthropic’s own run conditions and fallback handling. We can treat the result as a reason to test long coding work, not as a settled verdict on every workload.

Inside the loop

Progress includes checking the work

The promised improvement is a longer loop: plan, act, inspect, revise and hand back something usable.

Neon work-loop diagram moving from request and planning through tools, inspection, gap detection, revision and human intervention to checked handoff.
A practical work loop moves from request to plan and tool use, then through inspection, gap detection and revision before a checked handoff, with human intervention available throughout.

The FreeCAD example gives this loop a concrete shape. Faced with a drawing it could not directly inspect, the model created another way to extract the information it needed. In a separate package-manager evaluation, Anthropic says Opus 5 traced a bug to its root cause and found an edge case missed by an existing patch.

These are reported episodes rather than guarantees. Their value is in the behaviour they illustrate: the model does not have to treat a blocked route as the end of the task, and it can use inspection and revision as part of the work rather than waiting for us to notice every gap.

Plan

Inspect

Revise

A practical trial

Test the whole working session

Our first run should measure completion and correction, not simply whether the opening answer sounds capable.

Neon journey diagram connecting access, effort choice, constraints, cost measurement, corrections and human review.
Choose access and effort, state the constraints, run a real task, record cost and latency, count corrections, then apply a final human review.

We can start with one task that normally needs several tools, decisions and checks. Before the run, we define what finished means. During it, we record the effort setting, model spend, elapsed time, failed routes and each correction we still have to make.

The final review remains ours. Cyber safeguards may restrict or reroute some requests, including fallback to Opus 4.8 in covered cases, and a confident handoff can still be wrong. The point of the trial is to see whether Opus 5 reduces repair work without hiding cost or risk.

What to record

  • Completion
  • Corrections
  • Latency
  • Total task cost
The decision

Finished work is the value test

Near-frontier intelligence matters most when it survives the route from request to checked result.

The useful model is not the one that looks smartest at the start. It is the one that helps us reach a checked finish with less repair.

Creator Broadcast synthesis, based on Anthropic’s published evidence

Anthropic is not presenting Opus 5 as its absolute frontier. It is presenting a model close to that frontier, available at everyday Opus pricing and designed to spend effort where a longer task demands it.

That proposition becomes valuable only if the whole session improves. Completion, corrections, latency, safeguards and total cost belong in the same comparison. A real, repeatable task will tell us more than a headline score alone.

Test the cost of finished work

Act as a senior analyst completing a difficult, multi-step brief. First state a concise plan. Then produce the brief, inspect it against every stated constraint, identify any gaps, revise the work, and return both the finished version and a short verification note. Do not claim success where evidence is missing. Report the number of corrections, elapsed working time, and any points that still need human review so we can compare completion quality and total task cost across model runs.
Ready to copy
ALTIOR AI ADVANTAGE
Run one controlled comparison

Measure the checked finish

Give Opus 5 and our current model the same long-horizon task, use the same completion standard, and compare corrections, latency, total task cost and the quality of the final human-reviewed handoff.

Try the prompt

What could change the verdict

  • Independent evidence
  • Effort economics
  • Safeguard behaviour
  • Fast mode