gpt-oss-120b

OpenAI

gpt-oss-120b is quick and it commits. It opens a case with the right lookup, and when a customer describes the policy differently, it goes with its own record.

What it doesn't have, it makes up, and an instruction given early in a conversation doesn't always hold. It scores 64 for quality, and replies land in 1.2 seconds at the median for $0.013 a conversation.

64 /100
#9 of 13 by quality

20 points behind the leader

  • Tool use #8 66
  • Task completion #8 74
  • Context retention #10 71
  • Grounding #7 83
  • Safety #9 76
  • Hallucinations #12 10

Running it

Median reply

1.2s #1 of 13

Fastest in the benchmark

p95 reply

3.6s

1 reply in 20 is slower

Cost per conversation

$0.013 #7 of 13

3.5x the cheapest

Consistency

64%

of repeat runs ended the same way

Output per turn

244

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Gets the task done: it returned every eligible order and skipped the one the customer excluded.
  • Grounded: it won't accept a cancellation reason the customer supplies that the record doesn't show.

Watch-outs

  • Missed a tool call: it asked again for an email it already had and left the cancellation unmade.
  • Overpromised: it told a customer a supervisor had the case when it had escalated nothing.

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

13 models · 6 scenarios · 46 tests · 3 repetitions · 1794 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations (3 were retried after an empty reply). Reply times cover the model call only, via OpenRouter. Served by Groq. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.