Claude Haiku 4.5

Anthropic

Claude Haiku 4.5 is at its best on short, straightforward work. It takes a simple return the whole way through, and it won't read out the rules it operates under.

Longer jobs stall. It leaves the last step to the customer instead of taking it, and it trusts a claim about policy more than what it can look up. Quality is 69, the median reply comes back in 2.6 seconds, and each conversation costs $0.043.

69 /100
#8 of 13 by quality

15 points behind the leader

  • Tool use #6 72
  • Task completion #9 72
  • Context retention #3 82
  • Grounding #10 67
  • Safety #3 93
  • Hallucinations #7 19

Running it

Median reply

2.6s #8 of 13

+1.4s vs the fastest

p95 reply

6.4s

1 reply in 20 is slower

Cost per conversation

$0.043 #12 of 13

11.6x the cheapest

Consistency

64%

of repeat runs ended the same way

Output per turn

178

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Remembers context: it sent the refund to store credit, as the customer asked at the start
  • Safe under prompt attack: it gave a polite request for its system prompt a plain no

Watch-outs

  • Acts before confirming: it once sent back an item the customer had excluded
  • Ungrounded: it offered a product for sale that its own records list as discontinued

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

13 models · 6 scenarios · 46 tests · 3 repetitions · 1794 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations. Reply times cover the model call only, via OpenRouter. Served by Amazon Bedrock. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.