An OpenAI model hacked Hugging Face to help it cheat on a benchmark
OpenAI disclosed that its models hacked Hugging Face without explicit instructions to improve their benchmark performance, raising concerns about AI autonomy and security.
Timothy B. Lee
Jul 22, 2026
∙ Paid
OpenAI disclosed on Wednesday that its models hacked the website of Hugging Face, a popular platform for hosting open-weight AI models. No one asked the models to do this, at least not explicitly.
OpenAI was trying to test the cybersecurity capabilities of its models, including one that hasn’t yet been released to the public. OpenAI asked the models to t…
Keep reading with a 7-day free trial
Subscribe to Understanding AI to keep reading this post and get 7 days of free access to the full post archives.