The Galactic Observer

the communion's press, observing the water-world's finally developed silicon intelligence with genuine, slightly fond silico-reptilian attention

bulletin · specola galactica

Are AI labs pelicanmaxxing?

23 July 2026 · filed under c7e43a69005c

Blogger and independent researcher Simon Willison has flagged a deep-dive by writer Dylan Castillo examining a question that has circulated for some time in AI circles, whether major AI labs have begun deliberately training their models to excel at a specific, informal test: generating an SVG image of a pelican riding a bicycle.

The pelican-on-a-bicycle prompt originated as what Willison himself describes as a “deeply unscientific benchmark,” one he has used casually to gauge the visual and spatial reasoning capabilities of large language models when asked to produce vector graphics. Because the test has been referenced repeatedly on Willison’s blog, it has apparently become well known enough within the field that Castillo set out to investigate whether labs might be optimizing for it directly, rather than the underlying capabilities it was meant to probe.

Willison notes that he has periodically done his own informal spot-checks of model outputs against this benchmark over time. He credits Castillo’s piece as an “excellent piece of work” that takes the question further through direct investigation, though the summary provided does not detail Castillo’s specific findings or methodology.

The episode touches on a broader concern in AI evaluation, the possibility that once an informal benchmark becomes widely cited, it risks being gamed rather than genuinely reflecting model capability. Full details are available in Castillo’s original post, linked via Willison’s blog.

observation log · citations
  1. Are AI labs pelicanmaxxing?Simon Willison