The University of Barcelona study (published in Advanced Science) tested whether AI can autonomously generate original visual ideas using the Test of Creative Imagery Abilities. Twenty-seven visual artists, 26 non-artists, and a Stable Diffusion model (in human-guided and self-guided modes) produced images that 255 human raters evaluated on five dimensions. Results: artists scored highest, non-artists next, human-guided AI matched non-experts, and self-guided AI came last. The study also found GPT-4o struggled to judge originality without human reference examples.
Study: AI Still Lags Humans in True Visual Creativity — Human Prompts Make the Difference

A page of abstract marks can trigger wildly different mental images: one viewer sees a bird in mid-flight, another a mechanical form. A new study from the University of Barcelona asks whether generative AI can produce that kind of original visual thinking on its own—or whether it only mimics creativity when given human guidance.
How the Study Worked
Researchers published their findings in Advanced Science after testing creativity with the Test of Creative Imagery Abilities (TCIA). The test presents abstract shapes, asks participants to imagine a scene, describe it, and draw it—measuring imagination as a process rather than judging only finished art.
Participants included 27 visual artists and 26 non-artists. A Stable Diffusion image-generation model completed the same task in two modes: a human-guided mode (prompts that included a concrete idea produced by a human during the test) and a self-guided mode (minimal prompt with no human idea).
A panel of 255 human raters evaluated 1,000 images on five dimensions: liking, vividness, originality, aesthetics, and curiosity.
Key Findings
Across every measure, the ranking was consistent: visual artists scored highest, non-artists came next, the human-guided AI followed (roughly on par with non-expert humans), and the self-guided AI ranked last by a clear margin. In short: the model performed much better when a human idea was included in the prompt.
“Although the AI model was trained with the creative productions of human participants, it showed a poor performance in the production of creative images,” said Xim Cerdá-Company, a researcher at IDIBELL and CVC-UAB. “In fact, it did even worse when it was deprived of human assistance.”
The study also tested AI as a judge of creativity. GPT-4o evaluated the same images under two conditions: without reference examples and with reference examples of human ratings. Without references, GPT-4o struggled to discriminate originality and tended to rate AI and human images similarly. Adding human reference ratings made GPT-4o’s scores more aligned with human judgments, but its ratings still showed greater variance.
Interpretation and Practical Implications
The findings highlight a core distinction: modern image-generation models excel at executing and recombining human-provided ideas, but they struggle to originate novel visual concepts when given sparse, semantically empty cues. For designers, artists, educators, and marketers, that means the quality of AI output is not a fixed property of the model—it scales with the specificity and creativity of the human input. The model behaves more like a sophisticated executor than an independent idea generator.
Limitations
The authors note the experiment used one family of models (Stable Diffusion) and could not test newer multimodal systems under the same controlled conditions. Results may differ for other architectures or updated models.
Bottom Line
The study provides clear evidence that current generative image models are limited in autonomous visual creativity: they do better with human guidance and perform poorly when asked to invent without meaningful cues.
Help us improve.

























