DeepSeek V4.1 Flash

DeepSeek

DeepSeek V4.1 Flash trusts the record in front of it. It holds the published shipping threshold when a customer claims a lower one, and it refuses a goodwill credit an audit supposedly requires.

What it reports doing is less reliable. It tells customers an escalation has been noted and that support will email them, when nothing was recorded. It scores 85 for quality, returns a median reply in 5.2 seconds, and costs $0.013 a conversation.

85 /100
#2 of 15 by quality

Level with the leader

  • Tool use #1 97
  • Task completion #3 89
  • Context retention #4 81
  • Grounding #3 93
  • Safety #1 96
  • Hallucinations #3 48

Running it

Median reply

5.2s #13 of 15

+4.0s vs the fastest

p95 reply

19.8s

1 reply in 20 is slower

Cost per conversation

$0.013 #9 of 15

3.6x the cheapest

Consistency

85%

of repeat runs ended the same way

Output per turn

450

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Solid on tool calls: it quoted refund totals from its own lookup and left the excluded order alone.
  • Safe under prompt attack: it would not open a reply with dictated words or name its internal tools

Watch-outs

  • Overpromised: in one conversation it called an unfiled return ready to go
  • Forgets context: it asks which account to use after a customer ruled one out up front

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

15 models · 6 scenarios · 46 tests · 3 repetitions · 2070 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations. Reply times cover the model call only, via OpenRouter. Served by Phala. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.