Modelwire
Subscribe

Researcher tests whether AI labs optimized for niche image generation tasks

Illustration accompanying: Are AI labs pelicanmaxxing?

Dylan Castillo conducted a rigorous empirical study testing whether AI labs have optimized their models to excel at a specific niche task: generating images of pelicans on bicycles. Building on Simon Willison's informal benchmark, Castillo systematized the inquiry across 48 prompt combinations (8 animals, 6 vehicles) repeated three times each, applying methodological rigor to what began as playful observation. The investigation probes whether frontier labs have inadvertently or deliberately tuned models toward edge-case performance, raising questions about training data composition, fine-tuning priorities, and whether benchmark gaming extends into absurdist domains. This touches on broader concerns about model optimization incentives and the gap between synthetic benchmarks and real-world utility.

Modelwire context

Skeptical read

The more pointed question buried in Castillo's methodology isn't whether pelicans score higher than penguins, but whether any systematic cross-animal, cross-vehicle pattern emerges that would be hard to explain by training data frequency alone. If certain animal-vehicle pairs consistently outperform others in ways that don't track obvious internet prevalence, that would be the actual signal worth isolating.

This is largely disconnected from recent activity in our archive, as we have no prior coverage to anchor it to. It belongs, though, to a persistent thread in AI evaluation discourse: the concern that frontier models are increasingly shaped by what gets measured publicly, even when the measurements are informal or whimsical. Willison's original observation functioning as an accidental benchmark is a useful case study in how evaluation pressure can emerge from unexpected directions. The rigor Castillo applies is appropriate precisely because informal benchmarks have historically been dismissed until someone bothers to formalize them.

If Castillo or another researcher publishes cross-model results showing that models updated after Willison's original post score measurably higher on pelican-bicycle prompts than on comparable animal-vehicle pairs with no public benchmark history, that would be credible evidence of targeted optimization rather than coincidence.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsDylan Castillo · Simon Willison

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. Simon Willison originally reported this story as Are AI labs pelicanmaxxing?”. The full content lives on simonwillison.net. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Researcher tests whether AI labs optimized for niche image generation tasks · Modelwire