Release: ttok 1.0
Simon Willison’s Weblog
Newsletter
9th October 2026
Release
ttok 1.0 — Count and truncate text based on tokens
I released ttok 0.4, ran uv tool upgrade ttok, piped a file into the new version... and realized that it was defaulting to the GPT-4 tokenizer when it should very clearly default to GPT-5/GPT-6 instead!
I figured switching the default was a reasonable excuse to finally ship a 1.0.
OpenAI haven't actually confirmed that GPT-6 uses the same tokenizer as the GPT-5 family yet - there's an angry issue about it - but I found this commit by William Liu which reports on an experiment he ran confirming that the tokenizers are likely the same:
All seven GPT models (5.5, 5.6 Sol/Terra/Luna, 6 Astra/Sol/Luna) report 44,794 tokens and match each other on every one of the 31 fixtures. GPT-6 introduces no input-count change on this corpus.
Recent articles
Claude Haiku 5.5 - 7th October 2026
We're going to need default hard budget caps on pretty much everything - 3rd October 2026
OpenAI DevDay 2026 live blog - 29th September 2026
This is a beat by Simon Willison, posted on 9th October 2026.
projects 556
ai 2,271
openai 474
generative-ai 2,013
llms 1,979
tokenization 15
Monthly briefing
Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.
Pay me to send you less!
Sponsor & subscribe
Disclosures
Colophon
©
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026