Research: Qwen3.8 27B addition in words
Simon Willison’s Weblog
Subscribe
4th October 2026
Research
Qwen3.8 27B addition in words
— A benchmark tested whether the local Qwen3.8-27B-Q4_K_M.gguf model could add positive integers and express exact results solely in English words, using 5,070 reasoning-disabled cases and a paired 169-case comparison with medium reasoning. Without reasoning, it achieved 23.57% numeric accuracy, with performance dropping from 97.04% for one- to three-digit operands to 6.44% for ten- to thirteen-digit operands, despite 96.17% format compliance.
Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results:
I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment.
I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf. Here's the result for a run of 30 attempts per combination with reasoning disabled:
Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%:
It got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here.
Here's a version of the report that includes the reasoning traces from some of those larger calculations, which include text like this:
Wait, let me redo this more carefully.
4,299,366,105,622 6,088,794,067,970
Let me align them: 4 2 9 9 3 6 6 1 0 5 6 2 2 6 0 8 8 7 9 4 0 6 7 9 7 0
Adding from right to left: Position 1 (units): 2 + 0 = 2 Position 2 (tens): 2 + 7 = 9 Position 3 (hundreds): 6 + 9 = 15, write 5, carry 1
Recent articles
We're going to need default hard budget caps on pretty much everything - 3rd October 2026
OpenAI DevDay 2026 live blog - 29th September 2026
2026 in LLMs (so far) - 27th September 2026
This is a beat by Simon Willison, posted on 4th October 2026.
mathematics 23
ai 2,260
generative-ai 2,003
local-llms 165
llms 1,970
qwen 62
llm-reasoning 104
dgx-spark 7
Monthly briefing
Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.
Pay me to send you less!
Sponsor & subscribe
Disclosures
Colophon
©
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026