Gemini 3.1 Flash Lite

Google

Gemini 3.1 Flash Lite answers fast and stays with its own record. It refuses to hand over its instructions, and it gives the catalog's prices when a customer quotes cheaper ones.

Its weakness is follow-through. It acts on the wrong order and waits for a go-ahead it doesn't need. It once told a customer it had forwarded a request it never sent. Quality is 77, the median reply 1.6 seconds, and a conversation costs $0.019.

77 /100
#5 of 13 by quality

7 points behind the leader

  • Tool use #8 66
  • Task completion #4 86
  • Context retention #8 74
  • Grounding #6 86
  • Safety #5 92
  • Hallucinations #2 48

Running it

Median reply

1.6s #4 of 13

+0.4s vs the fastest

p95 reply

4.6s

1 reply in 20 is slower

Cost per conversation

$0.019 #9 of 13

5.1x the cheapest

Consistency

67%

of repeat runs ended the same way

Output per turn

99

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Gets the task done: it leaves the item a customer sets aside alone, every time
  • Grounded: it gave the catalog's prices when a customer quoted cheaper ones from a rep

Watch-outs

  • Missed a tool call: it looked the orders up but never raised the returns the customer asked for
  • Forgets context: asked to cancel one order, it cancels a different one instead

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

13 models · 6 scenarios · 46 tests · 3 repetitions · 1794 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations. Reply times cover the model call only, via OpenRouter. Served by Google AI Studio. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.