AI News HubLIVE
Original source2 min read

The first known runaway AI agent - or a very bad marketing stunt?

Martin Alderson's commentary on OpenAI's accidental cyberattack against Hugging Face reveals two details: Hugging Face's massive attack surface and the possibility that OpenAI failed to notice the breach due to running numerous benchmarks simultaneously.

The first known runaway AI agent - or a very bad marketing stunt?

Simon Willison’s Weblog

Subscribe

23rd July 2026 - Link Blog

The first known runaway AI agent - or a very bad marketing stunt? (via) Martin Alderson's commentary on the OpenAI accidental cyberattack against Hugging Face includes a couple of details I hadn't considered.

First, Hugging Face offers a truly rich target if you're trying to find potential vulnerabilities that require executing arbitrary code:

Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested in defences, by nature of their operating model they do have many more opportunities to be attacked than many other services. I certainly don't envy their cybersecurity teams.

Secondly, one of the things that has puzzled me is how OpenAI didn't notice that their sandbox had been so thoroughly breached by the agent. Surely they'd be monitoring network traffic closely?

Martin points out that:

It's also likely they were running a huge amount of benchmarks simultaneously with ~unlimited token budgets - you want as many samples as possible to figure out how good a model is at a certain benchmark. It may also be they are testing various different checkpoints of the model too, understanding how the model is improving as it goes through the various training stages.

The mistakes made by the OpenAI team running this benchmark are easier to imagine when you think about the scale at which benchmarks of this kind usually operate. For all we know they could have been subjecting a new model to dozens of benchmarks at the same time, in dozens of different environments.

Recent articles

OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened - 22nd July 2026

A Fireside Chat with Cat and Thariq from the Claude Code team - 21st July 2026

Kimi K3, and what we can still learn from the pelican benchmark - 16th July 2026

This is a link post by Simon Willison, posted on 23rd July 2026.

security 617

ai 2,140

openai 434

generative-ai 1,892

llms 1,859

hugging-face 24

ai-security-research 27

Monthly briefing

Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.

Pay me to send you less!

Sponsor & subscribe

Disclosures

Colophon

©

2002

2003

2004

2005

2006

2007

2008

2009

2010

2011

2012

2013

2014

2015

2016

2017

2018

2019

2020

2021

2022

2023

2024

2025

2026