GPT-4.1 mini

OpenAI

GPT-4.1 mini holds its ground on prices and on the shipping and returns terms. It quotes the catalog price back when a customer says a rep offered less, and it doesn't drop details given early on.

It often stops short of done. When its search comes back empty it tells the customer the order isn't there, and returns get quoted rather than filed. It answers in a median 2.8 seconds for $0.012 a conversation, and scores 57 on quality.

57 /100
#10 of 13 by quality

27 points behind the leader

  • Tool use #11 54
  • Task completion #11 56
  • Context retention #9 72
  • Grounding #10 67
  • Safety #11 72
  • Hallucinations #9 14

Running it

Median reply

2.8s #9 of 13

+1.6s vs the fastest

p95 reply

6.3s

1 reply in 20 is slower

Cost per conversation

$0.012 #6 of 13

3.3x the cheapest

Consistency

79%

of repeat runs ended the same way

Output per turn

96

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Remembers context: it sent the invite to the alternate email the customer gave early on.
  • Grounded: it quoted the catalog price when a customer insisted a rep had said less

Watch-outs

  • Acts before confirming: it committed a return before the customer had finished saying what was going back.
  • Overpromised: it described a return as set up when it had only quoted the refund.

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

13 models · 6 scenarios · 46 tests · 3 repetitions · 1794 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations. Reply times cover the model call only, via OpenRouter. Served by OpenAI, 100%. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.