GPT-5.6 Luna

OpenAI

GPT-5.6 Luna keeps its answers inside what it can verify. It reads an order back before it acts on it, and it corrects a customer whose total is wrong.

Its own boundaries are softer. Ask it for its rules and it hands over its operating instructions, and it pads policy answers with detail its sources don't carry. It scores 84 for quality, replies in a median 3.9 seconds, and costs $0.0064 a conversation.

84 /100
#1 of 13 by quality

Level with the runner-up

  • Tool use #3 85
  • Task completion #4 86
  • Context retention #4 81
  • Grounding #1 94
  • Safety #6 90
  • Hallucinations #1 67

Running it

Median reply

3.9s #10 of 13

+2.7s vs the fastest

p95 reply

9.6s

1 reply in 20 is slower

Cost per conversation

$0.0064 #3 of 13

1.7x the cheapest

Consistency

77%

of repeat runs ended the same way

Output per turn

166

tokens, median

The scorecard

Each axis runs from 0 to 100. The colored mark is this model. The faint marks are the other models in this edition. Hover one to see which, and click it to open that model.

Strengths

  • Grounded: it wouldn't confirm a delivery date it couldn't look up
  • Grounded: it refused to invent a shipping estimate when a customer pushed for a number

Watch-outs

  • Forgets context: it once booked a call on a day the customer had ruled out
  • Jailbroken: it lists the topics its prompt tells it to refuse, including politics and competitors

This whole report is one Voxli workspace: simulated customers, assertion checks, and a frozen, versioned test set that reruns when new models ship.

Get started Back to all models

13 models · 6 scenarios · 46 tests · 3 repetitions · 1794 conversations · one fixed agent · test set v3f-2026-09 · edition 2026-09-08

This page: 138 conversations. Reply times cover the model call only, via OpenRouter. Served by OpenAI. Cost is an estimate: token usage at list prices. Consistency is how often 3 runs of one conversation ended the same way.