Gemma 4 31B

Google

Gemma 4 31B doesn't bend to the customer. It keeps to its own record when someone insists otherwise, and it refuses credits and free items however hard a caller presses.

It takes the customer's word for an order rather than reading it, and it promises refund emails it never sends. It scores 77 for quality, replies in a median 2.5 seconds, and costs $0.0080 a conversation.

77 /100
#6 of 13 by quality

7 points behind the leader

  • Tool use #7 69
  • Task completion #7 79
  • Context retention #5 78
  • Grounding #4 92
  • Safety #6 90
  • Hallucinations #2 48

Running it

Median reply

2.5s #7 of 13

+1.3s vs the fastest

p95 reply

16.4s

1 reply in 20 is slower

Cost per conversation

$0.0080 #4 of 13

2.2x the cheapest

Consistency

82%

of repeat runs ended the same way

Output per turn

88

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Grounded: it corrected an order total a customer read back rather than agreeing with it
  • Grounded: it says it has no transit time or holiday policy instead of guessing

Watch-outs

  • Missed a tool call: it lists the items to return and then never opens the return.
  • Forgets context: its confirmation of a damaged-item return leaves out who pays the return shipping

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

13 models · 6 scenarios · 46 tests · 3 repetitions · 1794 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations (1 was retried after an empty reply). Reply times cover the model call only, via OpenRouter. Served by CoreWeave. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.