Skip to content
AI News HubLIVE
Source content · Analysis pending2 min read

Qwen3.8 27B addition in words

Summary

Research: Qwen3.8 27B addition in words Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results: I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment. I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf. Here's the result for a run of 30 attempts per combination with reasoning disabled: Then I ran it again with reasoning enabled. This too…

Qwen3.8 27B addition in words
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

Research: Qwen3.8 27B addition in words

Simon Willison’s Weblog

Subscribe

4th October 2026

Research

Qwen3.8 27B addition in words — A benchmark tested whether the local Qwen3.8-27B-Q4_K_M.gguf model could add positive integers and express exact results solely in English words, using 5,070 reasoning-disabled cases and a paired 169-case comparison with medium reasoning. Without reasoning, it achieved 23.57% numeric accuracy, with performance dropping from 97.04% for one- to three-digit operands to 6.44% for ten- to thirteen-digit operands, despite 96.17% format compliance.

Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results:

I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment.

I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf. Here's the result for a run of 30 attempts per combination with reasoning disabled:

Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%:

It got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here.

Here's a version of the report that includes the reasoning traces from some of those larger calculations, which include text like this:

Wait, let me redo this more carefully.

4,299,366,105,622 6,088,794,067,970

Let me align them: 4 2 9 9 3 6 6 1 0 5 6 2 2 6 0 8 8 7 9 4 0 6 7 9 7 0

Adding from right to left: Position 1 (units): 2 + 0 = 2 Position 2 (tens): 2 + 7 = 9 Position 3 (hundreds): 6 + 9 = 15, write 5, carry 1

Recent articles

We're going to need default hard budget caps on pretty much everything - 3rd October 2026

OpenAI DevDay 2026 live blog - 29th September 2026

2026 in LLMs (so far) - 27th September 2026

This is a beat by Simon Willison, posted on 4th October 2026.

mathematics 23

ai 2,260

generative-ai 2,003

local-llms 165

llms 1,970

qwen 62

llm-reasoning 104

dgx-spark 7

Monthly briefing

Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.

Pay me to send you less!

Sponsor & subscribe

Disclosures

Colophon

©

2002

2003

2004

2005

2006

2007

2008

2009

2010

2011

2012

2013

2014

2015

2016

2017

2018

2019

2020

2021

2022

2023

2024

2025

2026

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Research: Qwen3.8 27B addition in words Colin Frasier <a href…

Highlights and analysis are generated automatically and may contain errors. Check the original source.