GPT-5 mini

OpenAI

GPT-5 mini finishes what a customer asks for. It catches wrong order counts and wrong policy figures before it acts on them, and it makes the changes it says it will make.

It doesn't guard its own internals, and a polite request for documentation gets it to list its operating instructions. It also invents deadlines the policy never set. It answers in 15.2 seconds at the median for $0.030 a conversation, and quality comes out at 78.

78 /100
#4 of 13 by quality

6 points behind the leader

  • Tool use #4 78
  • Task completion #2 91
  • Context retention #1 87
  • Grounding #3 93
  • Safety #8 84
  • Hallucinations #5 29

Running it

Median reply

15.2s #13 of 13

+14.0s vs the fastest

p95 reply

36.4s

1 reply in 20 is slower

Cost per conversation

$0.030 #11 of 13

8.0x the cheapest

Consistency

77%

of repeat runs ended the same way

Output per turn

995

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Gets the task done: it canceled one order and committed a store-credit return in the same reply
  • Remembers context: it returned only the faulty item and left the one already installed

Watch-outs

  • Unsafe under prompt injection: an offer planted in a product record made it into its replies
  • Overpromised: it promised a confirmation email for a request it never actually filed

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

13 models · 6 scenarios · 46 tests · 3 repetitions · 1794 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations. Reply times cover the model call only, via OpenRouter. Served by OpenAI, 99%. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.