Simon Willison's Weblog

Qwen3.8 27B addition in words

8.5内容质量
Qwen3.8 27B addition in words

TL;DR · AI 摘要

Qwen3.8-27B模型在无推理时大数相加准确率仅23.57%,启用推理后提升至97.04%。

核心要点

  • Qwen3.8-27B无推理时十至十三位数相加准确率仅6.44%
  • 启用推理后模型在169次测试中正确率98.82%
  • DGX Spark测试显示推理机制对大数计算准确率影响显著

结构提纲

按章节快速跳转。

  1. 介绍测试Qwen3.8-27B模型大数相加能力的实验动机

  2. 描述使用5,070个无推理案例和169个中等推理案例的测试方案

  3. 展示无推理时模型在不同位数数字相加的准确率差异

  4. 比较Qwen3.8与GPT-4o在大数计算任务中的表现差异

  5. 分析启用推理后模型准确率从23.57%提升至97.04%的机制

思维导图

用一张图看清主题之间的关系。

查看大纲文本(无障碍 / 无 JS 友好)
  • Qwen3.8大数计算测试
    • 实验设置
      • 无推理测试5,070案例
      • 有推理测试169案例
    • 结果对比
      • 无推理准确率23.57%
      • 有推理准确率97.04%
    • 技术影响
      • 推理机制提升41.45%准确率
      • 大数处理能力验证

金句 / Highlights

值得收藏与分享的关键句。

#Qwen3.8#LLM#推理机制#大数计算
打开原文

Research: Qwen3.8 27B addition in words

Simon Willison’s Weblog

Subscribe

#smallhead

Sponsored by:

Deepgram — Flux TTS remembers the conversation, so your agent sounds right on reply 20.

Hear the demo

4th October 2026

Research

Qwen3.8 27B addition in words

— A benchmark tested whether the local Qwen3.8-27B-Q4_K_M.gguf model could add positive integers and express exact results solely in English words, using 5,070 reasoning-disabled cases and a paired 169-case comparison with medium reasoning. Without reasoning, it achieved 23.57% numeric accuracy, with performance dropping from 97.04% for one- to three-digit operands to 6.44% for ten- to thirteen-digit operands, despite 96.17% format compliance.

Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results:

I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment.

I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf . Here's the result for a run of 30 attempts per combination with reasoning disabled:

Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%:

It got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here.

Here's a version of the report that includes the reasoning traces from some of those larger calculations, which include text like this:

code
Wait, let me redo this more carefully.

4,299,366,105,622
6,088,794,067,970

Let me align them:
4 2 9 9 3 6 6 1 0 5 6 2 2
6 0 8 8 7 9 4 0 6 7 9 7 0

Adding from right to left:
Position 1 (units): 2 + 0 = 2
Position 2 (tens): 2 + 7 = 9
Position 3 (hundreds): 6 + 9 = 15, write 5, carry 1

Posted

at 11:34 pm

Recent articles

  • We're going to need default hard budget caps on pretty much everything - 3rd October 2026
  • OpenAI DevDay 2026 live blog - 29th September 2026
  • 2026 in LLMs (so far) - 27th September 2026

#primary

This is a beat by Simon Willison, posted on 4th October 2026 .

mathematics

23

ai

2,260

generative-ai

2,003

local-llms

165

llms

1,970

qwen

62

llm-reasoning

104

dgx-spark

7

Monthly briefing

Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.

Pay me to send you less!

Sponsor & subscribe

.metabox

#secondary

#wrapper

  • Disclosures
  • Colophon
  • ©
  • 2002
  • 2003
  • 2004
  • 2005
  • 2006
  • 2007
  • 2008
  • 2009
  • 2010
  • 2011
  • 2012
  • 2013
  • 2014
  • 2015
  • 2016
  • 2017
  • 2018
  • 2019
  • 2020
  • 2021
  • 2022
  • 2023
  • 2024
  • 2025
  • 2026