Across 9,198 people and 215,542 model observations, humans held a small average lead and a larger advantage at the creative extreme. We examine why “think like a genius” prompts were not enough.
ACADEMIC WORK REVIEWED
A large-scale comparison of divergent creativity in humans and large language models
ResearchersDawei Wang, Difang Huang, Haipeng Shen and Brian Uzzi.
Nature Human Behaviour · 23 December 2025 · 10:531–540 (2026) · Peer-reviewed article · Source language: English
English review: 10 September 2026 · The source date appears in the citation above.
Published review
Source and editorial checks are complete. DijitalPi did not conduct a customer experiment for this paper.
LET'S READ THE RESEARCH TOGETHER
We first explain the researchers' question, method and findings. We then discuss how to interpret the results, clearly separating DijitalPi's commentary from the source.
01 / WHAT DID THE RESEARCHERS WANT TO UNDERSTAND?
How do humans and large language models differ when generating new ideas?
Wang, Huang, Shen and Uzzi compared humans and LLMs at large scale on the same divergent-creativity task.
Explaining the context · DijitalPi commentary
Creativity includes the spread of results and the extreme top end, not only the average idea.
The task does not represent every form of brand strategy, long-form writing or collaborative creative execution.
02 / HOW WAS THE RESEARCH CONDUCTED?
How did they test the question?
The study compared responses from 9,198 people with 215,542 observations across model conditions. It also tested temperature, strategic prompts and genius, celebrity or demographic personas.
Understanding the method · DijitalPi commentary
The large sample makes it possible to inspect the centre and tails of the distribution. Model observations and unique human participants are different units.
A persona does not represent a real human group; it is an intervention that changes model output.
03 / RESEARCH FINDINGS
A small human average advantage became larger at the creative extreme.
Humans were slightly more creative on average, showed greater variation and reached higher values in the most creative tail. Genius or demographic personas helped up to a point and then could reverse real-world human patterns; strategic prompting produced mixed results.
People
9,198Human participants in the comparison.
Model observations
215,542Observations across model and prompt conditions.
Raising average idea quality may still remove rare, unusual options. Measure diversity and the top tail separately.
More dramatic personas do not improve quality without limit. Prompt design should be treated as an experiment.
Starting a workshop with AI can anchor people on its first suggestions. An independent human round helps preserve different paths.
05 / CONCLUSION AND OPEN QUESTIONS
What did we learn, and what do we still not know?
Humans held a small average advantage and a clearer advantage at the most creative extreme. The result does not mean humans win every creative task or that LLMs are useless.
Which sequence and selection method produces the best combined human-AI campaign ideas? That requires a field experiment.
Dawei Wang, Difang Huang, Haipeng Shen and Brian Uzzi. University of Hong Kong; Chinese Academy of Sciences; University of Chinese Academy of Sciences; Northwestern University.
The academic work belongs to the researchers named above. This page contains DijitalPi's explanatory review and original business examples; it is not a full translation of the paper.