Skip to the review
DijitalPi
TREN
Contact

INTERNATIONAL RESEARCH / PEER-REVIEWED ARTICLE

Are humans or AI more creative?

Across 9,198 people and 215,542 model observations, humans held a small average lead and a larger advantage at the creative extreme. We examine why “think like a genius” prompts were not enough.

ACADEMIC WORK REVIEWED

A large-scale comparison of divergent creativity in humans and large language models

ResearchersDawei Wang, Difang Huang, Haipeng Shen and Brian Uzzi.

Nature Human Behaviour · 23 December 2025 · 10:531–540 (2026) · Peer-reviewed article · Source language: English

ENGLISH REVIEW AND COMMENTARY: DIJITALPI

We explain the paper in plain English, discuss what it may mean for organisations and identify our own examples separately.

English review: 10 September 2026 · The source date appears in the citation above.

Published review

Source and editorial checks are complete. DijitalPi did not conduct a customer experiment for this paper.

LET'S READ THE RESEARCH TOGETHER

We first explain the researchers' question, method and findings. We then discuss how to interpret the results, clearly separating DijitalPi's commentary from the source.

01 / WHAT DID THE RESEARCHERS WANT TO UNDERSTAND?

How do humans and large language models differ when generating new ideas?

Wang, Huang, Shen and Uzzi compared humans and LLMs at large scale on the same divergent-creativity task.

Explaining the context · DijitalPi commentary

Creativity includes the spread of results and the extreme top end, not only the average idea.

The task does not represent every form of brand strategy, long-form writing or collaborative creative execution.

02 / HOW WAS THE RESEARCH CONDUCTED?

How did they test the question?

The study compared responses from 9,198 people with 215,542 observations across model conditions. It also tested temperature, strategic prompts and genius, celebrity or demographic personas.

Understanding the method · DijitalPi commentary

The large sample makes it possible to inspect the centre and tails of the distribution. Model observations and unique human participants are different units.

A persona does not represent a real human group; it is an intervention that changes model output.

03 / RESEARCH FINDINGS

A small human average advantage became larger at the creative extreme.

Humans were slightly more creative on average, showed greater variation and reached higher values in the most creative tail. Genius or demographic personas helped up to a point and then could reverse real-world human patterns; strategic prompting produced mixed results.

People
9,198Human participants in the comparison.
Model observations
215,542Observations across model and prompt conditions.

Source: Nature Human Behaviour article and DOI ↗

04 / DIJITALPI'S EXPLANATION

How should we interpret these findings?

Raising average idea quality may still remove rare, unusual options. Measure diversity and the top tail separately.

More dramatic personas do not improve quality without limit. Prompt design should be treated as an experiment.

Starting a workshop with AI can anchor people on its first suggestions. An independent human round helps preserve different paths.

05 / CONCLUSION AND OPEN QUESTIONS

What did we learn, and what do we still not know?

Humans held a small average advantage and a clearer advantage at the most creative extreme. The result does not mean humans win every creative task or that LLMs are useless.

Which sequence and selection method produces the best combined human-AI campaign ideas? That requires a field experiment.

Return to the researchers' original publication

Nature Human Behaviour article and DOI

DIJITALPI'S APPLICATION COMMENTARY

What could this look like in your organisation?

We created these scenarios to make the topic concrete. They are not cases from the paper or measured client results.

EXAMPLE 01

The average rises, diversity falls

AI ideas score well individually but repeat the same message patterns.

Is the creative pool stronger?

Open the recommendation for example 01

Inspect the average and the distribution together.

Report similarity, number of themes and the most unusual high-scoring ideas separately. One average score can hide a loss of diversity.

EXAMPLE 02

The “think like a genius” prompt

The team assigns famous-inventor roles to the model to increase creativity.

Does a persona guarantee creative output?

Open the recommendation for example 02

Treat the role as a test condition, not evidence.

Compare it with a version that has no persona. Show outputs in mixed order to independent evaluators and check for repeated ideas.

EXAMPLE 03

AI speaks before people think

A workshop opens with model suggestions, and the team builds every later idea around that first list.

Was independent human input preserved?

Open the recommendation for example 03

Run a short independent round first.

Have participants record ideas before seeing AI output, then combine the pools. This reduces anchoring on the first suggestions.

TRY IT WITH YOUR TEAM

Measure the diversity of an idea pool.

Evaluate AI and human ideas with the source hidden. Compare average score, theme diversity and the ideas in the top ten percent separately.

Explore our AI content-production approach

Source and ownership

Dawei Wang, Difang Huang, Haipeng Shen and Brian Uzzi.
University of Hong Kong; Chinese Academy of Sciences; University of Chinese Academy of Sciences; Northwestern University.

Nature Human Behaviour · 23 December 2025 · 10:531–540 (2026) · Peer-reviewed article. Open the original publication · Find on Google Scholar

The academic work belongs to the researchers named above. This page contains DijitalPi's explanatory review and original business examples; it is not a full translation of the paper.

← Return to all research articles

ASK AI ABOUT DIJITALPI

Let an AI explain what DijitalPi does.

Opens your chosen assistant with a ready research prompt. It reads the site live and answers.