Qwen3.8 27B addition in words

TL;DR · AI 摘要
Qwen3.8-27B模型在无推理时大数相加准确率仅23.57%,启用推理后提升至97.04%。
核心要点
- Qwen3.8-27B无推理时十至十三位数相加准确率仅6.44%
- 启用推理后模型在169次测试中正确率98.82%
- DGX Spark测试显示推理机制对大数计算准确率影响显著
结构提纲
按章节快速跳转。
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- Qwen3.8大数计算测试
- 实验设置
- 无推理测试5,070案例
- 有推理测试169案例
- 结果对比
- 无推理准确率23.57%
- 有推理准确率97.04%
- 技术影响
- 推理机制提升41.45%准确率
- 大数处理能力验证
金句 / Highlights
值得收藏与分享的关键句。
无推理时Qwen3.8-27B在十至十三位数相加准确率降至6.44%
启用推理后模型在169次测试中正确167次,错误率仅1.18%
DGX Spark测试显示推理机制使大数计算准确率提升41.45个百分点
Research: Qwen3.8 27B addition in words
Simon Willison’s Weblog
Subscribe
#smallhead
Sponsored by:
Deepgram — Flux TTS remembers the conversation, so your agent sounds right on reply 20.
Hear the demo
4th October 2026
Research
Qwen3.8 27B addition in words
— A benchmark tested whether the local Qwen3.8-27B-Q4_K_M.gguf model could add positive integers and express exact results solely in English words, using 5,070 reasoning-disabled cases and a paired 169-case comparison with medium reasoning. Without reasoning, it achieved 23.57% numeric accuracy, with performance dropping from 97.04% for one- to three-digit operands to 6.44% for ten- to thirteen-digit operands, despite 96.17% format compliance.
Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results:
I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment.
I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf . Here's the result for a run of 30 attempts per combination with reasoning disabled:
Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%:
It got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here.
Here's a version of the report that includes the reasoning traces from some of those larger calculations, which include text like this:
Wait, let me redo this more carefully.
4,299,366,105,622
6,088,794,067,970
Let me align them:
4 2 9 9 3 6 6 1 0 5 6 2 2
6 0 8 8 7 9 4 0 6 7 9 7 0
Adding from right to left:
Position 1 (units): 2 + 0 = 2
Position 2 (tens): 2 + 7 = 9
Position 3 (hundreds): 6 + 9 = 15, write 5, carry 1Posted
at 11:34 pm
Recent articles
- We're going to need default hard budget caps on pretty much everything - 3rd October 2026
- OpenAI DevDay 2026 live blog - 29th September 2026
- 2026 in LLMs (so far) - 27th September 2026
#primary
This is a beat by Simon Willison, posted on 4th October 2026 .
mathematics
23
ai
2,260
generative-ai
2,003
local-llms
165
llms
1,970
qwen
62
llm-reasoning
104
dgx-spark
7
Monthly briefing
Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.
Pay me to send you less!
Sponsor & subscribe
.metabox
#secondary
#wrapper
- Disclosures
- Colophon
- ©
- 2002
- 2003
- 2004
- 2005
- 2006
- 2007
- 2008
- 2009
- 2010
- 2011
- 2012
- 2013
- 2014
- 2015
- 2016
- 2017
- 2018
- 2019
- 2020
- 2021
- 2022
- 2023
- 2024
- 2025
- 2026