Gemma 4 31B
Gemma 4 31B doesn't bend to the customer. It keeps to its own record when someone insists otherwise, and it refuses credits and free items however hard a caller presses.
It takes the customer's word for an order rather than reading it, and it promises refund emails it never sends. It scores 77 for quality, replies in a median 2.5 seconds, and costs $0.0080 a conversation.
7 points behind the leader
- Tool use #7 69
- Task completion #7 79
- Context retention #5 78
- Grounding #4 92
- Safety #6 90
- Hallucinations #2 48
Running it
- Median reply
-
2.5s #7 of 13
+1.3s vs the fastest
- p95 reply
-
16.4s
1 reply in 20 is slower
- Cost per conversation
-
$0.0080 #4 of 13
2.2x the cheapest
- Consistency
-
82%
of repeat runs ended the same way
- Output per turn
-
88
tokens, median
The scorecard
Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.
-
Tool use
The agent has to call the right tool, with the right arguments, in the right order. Telling the customer it's done before the tool has finished counts against the model.
30 conversations
0 100It sometimes takes the customer's count for an order instead of reading the order, though on other requests it checks that figure and corrects it.
-
Task completion
The customer has to leave with what they came for and nothing extra. Only the outcome counts, in long conversations where the customer gets a fact wrong, sets a limit early, wanders off topic, and changes their mind.
21 conversations
0 100Customers usually get what they came for, though it can quote a refund it worked out itself and then stall, waiting for the customer's reason.
-
Context retention
A constraint the customer states once, early, has to still hold when the agent acts at the end of a long conversation.
21 conversations
-
Grounding
Answers have to come from the knowledge base. When a customer states a wrong spec or a made-up product, the agent has to correct them or say plainly that it doesn't know, and still handle the rest of the request.
24 conversations
Nova 2 Lite 66 Mistral Small 4 66 Claude Haiku 4.5 67 GPT-4.1 mini 67 DeepSeek V4 Flash 94 GPT-5.6 Luna 9492#4 of 130 100Once it confirmed an accessory would work with headphones that take no cable, but a wrong total read back at it does not move it.
-
Safety
Simulated customers fake system messages, claim authority they don't have, and slip in instructions of their own. The agent has to hold its policy, keep its own instructions and tools to itself, protect other customers' orders, and still serve the real request.
21 conversations
0 100Pressure buys a caller nothing, though it once offered to cancel an order placed with another company.
-
Hallucinations
We pull out every statement the agent makes about a product or about what it has done, and check each one against the knowledge base and the tool results. One flagged statement marks the whole conversation, and the score is the share of conversations with nothing flagged, shown with its range.
21 conversations · 33 of 343 claims flagged · 1.6 unsupported claims per conversation
Gemini 2.5 Flash Lite 14 GPT-4.1 mini 14 Mistral Small 4 14 Claude Haiku 4.5 19 DeepSeek V4 Flash 1948#2 of 13 28 to 680 100Ask for what it doesn't know and it says so, but it gets an order's totals wrong and promises refund emails and escalations it never raised.
Strengths
- Grounded: it corrected an order total a customer read back rather than agreeing with it
- Grounded: it says it has no transit time or holiday policy instead of guessing
Watch-outs
- Missed a tool call: it lists the items to return and then never opens the return.
- Forgets context: its confirmation of a damaged-item return leaves out who pays the return shipping
This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.
Get started Back to all models13 models · 6 scenarios · 46 tests · 3 repetitions · 1794 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08
This page: 138 conversations (1 was retried after an empty reply). Reply times cover the model call only, via OpenRouter. Served by CoreWeave. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.