AI News HubLIVE
Original source2 min read

Are AI labs pelicanmaxxing?

Dylan Castillo conducted a rigorous study testing 7 AI models on drawing various animals riding vehicles, investigating whether AI labs deliberately train models to draw pelicans on bicycles. The results show no evidence of 'pelicanmaxxing.'

Are AI labs pelicanmaxxing?

Simon Willison’s Weblog

Subscribe

22nd July 2026 - Link Blog

Are AI labs pelicanmaxxing? (via) Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my deeply unscientific benchmark.

I've been randomly spot-checking this in the past by testing models against other animals riding other types of vehicle, but never with anything close to the diligence of Dylan's methodology here.

Dylan took 8 animals × 6 vehicles = 48 prompts and ran them three times each through 7 different models ( GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro). He then used GPT-5.6 Luna and Gemini 3.1 Flash-Lite to help evaluate the results.

There's a neat filter view for exploring the results:

For the models he tested he could find no evidence of pelimaxxing:

The pelicans on bicycles don’t look any better

Labs are not better at drawing pelicans

Labs are not better at drawing bicycles

Labs are not better at drawing pelicans on bicycles, even adjusting for difficulty

The pelican-bicycle scenes don’t look memorized [...]

Pelicans aren’t drawn any better than other animals. Bicycles aren’t drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict. GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn’t put too much weight on it.

Recent articles

A Fireside Chat with Cat and Thariq from the Claude Code team - 21st July 2026

Kimi K3, and what we can still learn from the pelican benchmark - 16th July 2026

The new GPT-5.6 family: Luna, Terra, Sol - 9th July 2026

This is a link post by Simon Willison, posted on 22nd July 2026.

ai 2,137

generative-ai 1,889

llms 1,856

evals 44

pelican-riding-a-bicycle 129

Monthly briefing

Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.

Pay me to send you less!

Sponsor & subscribe

Disclosures

Colophon

©

2002

2003

2004

2005

2006

2007

2008

2009

2010

2011

2012

2013

2014

2015

2016

2017

2018

2019

2020

2021

2022

2023

2024

2025

2026