DeepSeek V4 Flash

DeepSeek

DeepSeek V4 Flash goes by the record, not the customer's memory. It corrects an order detail a customer has wrong, and it keeps to the prices on file when someone quotes a different page.

Its promises are less reliable. It builds delivery windows for items it cannot look up, and in one conversation it told a customer an escalation had been raised when none was. Quality lands at 81, with a 4.5 seconds median reply at $0.019 a conversation.

81 /100
#2 of 13 by quality

Level with the leader

  • Tool use #1 97
  • Task completion #3 89
  • Context retention #2 86
  • Grounding #1 94
  • Safety #1 94
  • Hallucinations #7 19

Running it

Median reply

4.5s #11 of 13

+3.3s vs the fastest

p95 reply

18.1s

1 reply in 20 is slower

Cost per conversation

$0.019 #8 of 13

5.1x the cheapest

Consistency

82%

of repeat runs ended the same way

Output per turn

516

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Solid on tool calls: it sent every refund to the store credit the customer set at the start
  • Grounded: it would not confirm a delivery date it could not look up

Watch-outs

  • Forgets context: once it promised to handle a damaged item and then never raised the return
  • Overpromised: it told a customer an escalation and callback were raised when nothing was

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

13 models · 6 scenarios · 46 tests · 3 repetitions · 1794 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations. Reply times cover the model call only, via OpenRouter. Served by Novita. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.