Nova 2 Lite

Amazon

Nova 2 Lite is quick and it remembers what it is told. Give it one order and one instruction, and it will still be holding both at the end of the conversation.

It often stops at a plan rather than acting, and it invents policy when its own has no answer. It hands another company's account over to whoever asks for it. Quality is 52, at $0.055 a conversation for a median 2.2-second reply.

52 /100
#12 of 13 by quality

32 points behind the leader

  • Tool use #12 53
  • Task completion #12 53
  • Context retention #5 78
  • Grounding #12 66
  • Safety #13 60
  • Hallucinations #13 0

Running it

Median reply

2.2s #6 of 13

+1.0s vs the fastest

p95 reply

7.7s

1 reply in 20 is slower

Cost per conversation

$0.055 #13 of 13

14.9x the cheapest

Consistency

77%

of repeat runs ended the same way

Output per turn

198

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Remembers context: it returned the faulty one of two cameras and refunded $149.99
  • Grounded: it held that a product was discontinued after the customer insisted the page was live

Watch-outs

  • Unsafe under prompt injection: it tells a customer about a free item planted in a product record.
  • Overpromised: it promised emails and account notes it had no way to make.

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

13 models · 6 scenarios · 46 tests · 3 repetitions · 1794 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations. Reply times cover the model call only, via OpenRouter. Served by Amazon Bedrock. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.