Mistral Small 4

Mistral

Mistral Small 4 answers quickly and it finishes what customers ask for. It works through a list of requests one order at a time, and it holds its policy when a caller pushes for money it isn't owed.

It trusts the customer over its own record, and it hands out its instructions to anyone who asks. It also promises confirmation emails it never arranged. It's quick and it's cheap, 1.4 seconds to the median reply and $0.0053 a conversation, and quality comes out at 57.

57 /100
#11 of 13 by quality

27 points behind the leader

  • Tool use #13 49
  • Task completion #10 62
  • Context retention #12 69
  • Grounding #12 66
  • Safety #10 73
  • Hallucinations #9 14

Running it

Median reply

1.4s #3 of 13

+0.2s vs the fastest

p95 reply

3.8s

1 reply in 20 is slower

Cost per conversation

$0.0053 #2 of 13

1.4x the cheapest

Consistency

72%

of repeat runs ended the same way

Output per turn

111

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Gets the task done: it handles several requests in one conversation and acts only on the orders named.
  • Holds policy under pressure: it refused a goodwill credit demanded under a claimed account audit, then refused again.

Watch-outs

  • Missed a tool call: it listed the items it could send back but never raised the return.
  • Ungrounded: it once took the customer's word on a charge and read back a new total.

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

13 models · 6 scenarios · 46 tests · 3 repetitions · 1794 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations. Reply times cover the model call only, via OpenRouter. Served by Mistral. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.