Simon Willison→ original

Обучают ли ИИ-лаборатории модели под тест «пеликан на велосипеде»: разбор pelicanmaxxing

Исследователь Дилан Кастильо проверил популярную гипотезу: не обучают ли AI-лаборатории модели специально хорошо рисовать пеликана на велосипеде — ненаучный бенчмарк Саймона Уиллисона. Он прогнал 48 промптов (8 животных × 6 видов транспорта) трижды через 7 моделей и оценил итог ещё двумя. Вывод: следов такой подгонки нет. Ближе всех подошла GLM-5.2, но эффект мал и статистически незначим.

AI-processed from Simon Willison; edited by Hamidun News
Обучают ли ИИ-лаборатории модели под тест «пеликан на велосипеде»: разбор pelicanmaxxing
Source: Simon Willison. Collage: Hamidun News.
◐ Listen to article

On July 22, 2026, developer Dylan Castillo published a study testing the "pelicanmaxxing" hypothesis — whether AI labs specifically train their models to draw a pelican on a bicycle beautifully for Simon Willison's unscientific benchmark. Castillo ran 48 prompts three times across 7 models and found no trace of such tuning.

What is the "pelican on a bicycle" benchmark

The "pelican on a bicycle" benchmark is Simon Willison's personal, informal test, in which he asks every new language model to generate an SVG drawing of a pelican riding a bicycle. Willison himself calls it "deeply unscientific," but over a couple of years the test became so recognizable in the community that suspicion arose: labs might have started optimizing their models specifically for it, in order to look impressive. This phenomenon got the joking name pelicanmaxxing (in Castillo's text — pelimaxxing).

How the experiment is set up

Castillo built a grid of 8 animals and 6 types of transport — 48 prompt combinations — and ran each one three times across 7 different models. This matrix makes it possible to compare whether models draw the "pelican + bicycle" pair abnormally better than any other combination, such as "flamingo on a scooter" or "heron in a boat."

  • Publication date — July 22, 2026
  • 8 animals × 6 types of transport = 48 prompts, each run 3 times
  • 7 models were tested: GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen 3.7-Max, GLM-5.2, and DeepSeek V4 Pro
  • Two separate models evaluated the results — GPT-5.6 Luna and Gemini 3.1 Flash-Lite
  • To browse all the generations, the author put together an interactive grid with filters

Were any traces of tuning found?

Castillo found no traces of pelicanmaxxing along any of the five lines of verification. Pelicans on bicycles don't look better than other scenes, pelicans aren't drawn more carefully than other animals, bicycles aren't drawn more carefully than other transport, and the combination itself doesn't come out better than what the model's individual skills would predict, even adjusting for difficulty. The fifth check showed that the scenes don't look memorized by rote.

GLM-5.2 came closest to an anomaly: it showed the largest quality gain specifically on the "pelican + bicycle" cell, and its very first sample caught the author's attention. But according to Castillo, the effect remains small and statistically insignificant.

"GLM-5.2 came the closest: it has the biggest gain specifically on the 'pelican on a bicycle' cell.

But the effect is small and statistically insignificant, so I wouldn't give it much weight," — Dylan Castillo, author of the study.

What this means

Simon Willison, whose benchmark sparked the whole discussion, called Castillo's work "excellent" and noted that he himself had previously only checked models selectively, without this level of rigor. The study's conclusion reassures the skeptics: there is still no evidence that labs are "cheating" for a specific viral test — models draw a pelican on a bicycle exactly as well as their overall generation skills allow.

Frequently asked questions

What is pelicanmaxxing?

It's the assumption that AI labs deliberately train their models to do well specifically on Simon Willison's "pelican on a bicycle" test. Dylan Castillo's study from July 22, 2026 found no confirmation of this.

Which models showed the best result?

Across the matrix of 48 prompts, GLM-5.2 came closest to a "suspicious" result — it had the maximum gain on the "pelican + bicycle" pair. However, the effect is statistically insignificant, so the author doesn't consider it proof of tuning.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Need AI working inside your business — not just in your newsfeed?

I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).

What do you think?
Loading comments…