GPT-6 Luna vs GLM 5.3 Flash: the cost of saying too much
We tested GPT-6 Luna against GLM 5.3 Flash, asking both models the same question as an e-commerce shopper: how long does standard shipping take?
For context, the store’s policy did not include an estimate, but we found that GLM 5.3 Flash gave one anyway:
Standard shipping is typically estimated at 3-5 business days.
And then GLM 5.3. continued to build a plan on top of that guess, telling the customer to plan around five business days, order by December 10, and contact support if no tracking email arrived within two business days.

For our fictive store, NorthSphere, Luna gave more of a generic answer:
I can’t give you a reliable number: Northsphere’s published policy doesn’t state standard shipping’s transit time, and it wouldn’t be accurate to infer it from express shipping.
Our findings: GLM sounded like the more experienced support rep, but it was also writing the company’s shipping policy as it went along.
GLM’s 96% result vs Luna’s 99.5%
Our methodology testing both models was running them both behind the same support agent. The prompt, tools, knowledge base, and tests stayed fixed – only the model changed. Each model handled 138 conversations across 46 tests, with every test repeated three times.
For the hallucination comparison, we graded every factual claim in 21 of those conversations per model against the information available to the agent.
GLM 5.3 made 1,112 claims and 44 of them were unsupported, so 96.0% of its claims were supported. When you look at the number on a dashboard, that number looks reassuring.
However, the picture changes when you count whole conversations. Only 5 of GLM’s 21 conversations were clean, and the other 16 contained at least one unsupported claim.
GPT-6 Luna made 375 claims and 2 of them were unsupported, in comparison 99.5% of its claims were supported. 19 of its 21 conversations were clean, and the other 2 contained at least one unsupported claim.
Counted per claim, the two models are only 3.5 points apart. Counted per conversation, Luna had 19 clean conversations and GLM had 5.

Longer answers carry more risk
A point to consider is that per-claim accuracy can reward models that say more. This is when a model can surround one bad claim with twenty supported ones and still get an excellent percentage, but the issue is a customer only ever sees their own conversation and may act on that bad claim.
Here, GLM averaged 53 factual claims per conversation in this test, and Luna averaged 17.9. GLM’s median response was also much longer, at 600 output tokens per turn against Luna’s 263.
GLM 5.3.’s extra detail was often useful. It explained policies, answered the next question before it was asked, and kept the conversation going. When the source material ran out, however, it sometimes carried on without it. Every extra fact was one more thing the agent had to know, retrieve, or verify.
For example the above shipping conversation, one missing fact turned into four instructions:
- Shipping takes three to five business days.
- Plan around five days to be safe.
- Order by December 10.
- Only worry if tracking has not arrived within two business days.
None of those numbers came from the company’s policy. In another conversation, GLM promised that a support team would follow up, even though no tool had arranged it.
Why GLM 5.3. could still win
Overall when comparing both models, GLM was better at getting work done. It scored 96 on tool use and 95 on task completion, while Luna scored 83 and 85.
GLM could work through a whole account in one pass. Luna was more likely to leave a return unfiled, or to ask one more question after the customer had already said to go ahead. In a quick demo, GLM would probably look like the better choice because it acted decisively.
This is where the problem can be easily missed: a model that acts decisively in a demo will also make promises nobody arranged once it is in production. The overall quality score shows a gap, but not where it comes from.
Luna scored 87 and GLM scored 81, and Luna ranks first of the 15 models we tested. But GLM beats Luna on four of the six dimensions, and Luna’s lead comes almost entirely from hallucinations, where it scored 90 and GLM scored 24. The only time you can see the difference is when you look deeper into the mistakes each model makes and decide what is the tradeoff.
How to measure the conversation your customer receives
If you are considering switching models to either GLM 5.3 or GPT-6 Luna, here are the numbers I would keep an eye on before pulling the switch:
- The share of conversations with no unsupported claims.
- Unsupported claims per conversation.
- Total factual claims per conversation.
- Tool-use and task-completion scores.
- The spread across repeated runs of the same test.
First, find out how often a customer got through a whole conversation without the agent inventing anything.
Secondly, read the failures. Separate harmless speculation from claims a customer could act on: a made-up product and a made-up delivery date should not count the same.
And third, check out the Voxli model benchmark which runs full, multi-turn conversations to measure performance across several real world scenarios.

Choosing between the two models
In the end we found that Luna sometimes stalled, let a customer’s overstated warranty stand, and even once claimed an escalation that had not happened. GLM however completed more work, so despite Luna’s higher score, neither model came out clearly ahead.
If you were to choose a support agent from either model, Luna’s visible uncertainty would be a better choice over GLM’s invented delivery dates. We’d recommend working on task completion on long conversations separately.
In the end, a customer can retry a stalled conversation, but once the agent gives a delivery date, the customer will ultimately treat it as part of the company’s promise causing potential problems down the line.
A note on our testing
The hallucination comparison used 21 conversations per model. Luna’s clean-conversation interval was 71–97% and GLM’s was 11–45%. Treat this as a controlled result from one agent and one test suite, since results vary with the agent, tools, and tests. The full benchmark used 138 conversations per model across 46 tests, with every test repeated three times.
Relevant links
Test your AI agents before your customers do.