A quote from Anthropic Frontier Red Team
Simon Willison’s Weblog
Subscribe
29th September 2026
We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.
— Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities
Recent articles
OpenAI DevDay 2026 live blog - 29th September 2026
2026 in LLMs (so far) - 27th September 2026
Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war - 22nd September 2026
This is a quotation collected by Simon Willison, posted on 29th September 2026.
ai 2,256
generative-ai 2,000
llms 1,967
anthropic 343
ai-in-china 109
glm 10
ai-security-research 45
Disclosures
Colophon
©
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026