AI lifts the middle of human creativity, study finds

GPT-4 beat the average human on a 100,000-person word-association creativity test, but the top 10% of humans outscored every AI model. Models also repeated favorites like "microscope" in 70% of responses, while humans showed far more diversity.

Categorized in: AI News Creatives
Published on: Aug 26, 2026
AI lifts the middle of human creativity, study finds

A peer-reviewed study published in Scientific Reports on January 21, 2026, compared human and AI creativity using a simple word-association test. The headline result: GPT-4 beat the average human score. But the full picture is more nuanced-and more reassuring for people who make things for a living.

The exercise is called the Divergent Association Task, or DAT. Participants provide ten single English nouns as different from one another as possible in meaning. "Cat," "dog," and "hamster" stay in a tight semantic neighborhood. "Cat," "thimble," and "liberty" are farther apart. The scoring system converts each word into a numerical representation based on its contexts across a large body of language, then calculates the distances among the first seven valid words. Seven words produce 21 possible pairs; their mean semantic distance, multiplied by 100, becomes the DAT score.

The method was introduced by Jay Olson and colleagues in a 2021 paper in the Proceedings of the National Academy of Sciences. Across 8,914 participants, DAT performance showed moderate to strong relationships with established creativity measures, including tasks that ask people to invent unusual uses for common objects or bridge apparently unrelated ideas.

How the 2026 comparison was built

That validation is why the 2026 comparison is more informative than asking a chatbot to write a poem and voting on whether it feels inspired. Humans and models received the same instruction, and both were scored with the same automated rule. The DAT measures divergent verbal association, one recognized component of creative cognition. It does not measure the whole process by which an idea becomes a useful invention, an affecting story, or a workable design.

The researchers selected exactly 100,000 English-speaking participants from a larger dataset. The sample was evenly split between men and women, with 20% drawn from each of five age bands: 18-29, 30-39, 40-49, 50-59, and 60 or older. Most were in the United States, with smaller groups from the United Kingdom, Canada, Australia, and New Zealand. This makes the sample unusually large and deliberately balanced on age and sex. It does not make it a random sample of the world's population-the headline's "humanity" is shorthand for this English-speaking volunteer dataset.

The team tested OpenAI's GPT-3.5, GPT-4, and GPT-4-turbo; Anthropic's Claude 3; Google's GeminiPro; and several open models. Each model condition produced 500 responses. The researchers started a new conversation on every iteration so one answer would not influence the next, and applied the same scoring rule to both groups.

What the scores actually showed

In the main comparison, GPT-4 had the highest model average and exceeded the overall human mean by a statistically significant margin. GeminiPro's mean was statistically indistinguishable from the human mean. GPT-4-turbo performed worse than the older GPT-4 in this task-a useful warning against assuming that a newer or more efficient model automatically becomes more divergent.

So GPT-4 did not sit across from one representative person and win a creative duel. Five hundred GPT-4 generations produced a higher mean DAT score than the mean of 100,000 human responses.

The most creative half of the human sample, however, had a higher average than every model tested. The top 10% moved farther ahead. The researchers constructed benchmarks from the upper 50%, upper 25%, and upper 10% of human scores. The average of the upper human half remained above the mean of every model in the curated comparison. The top decile defined the clearest separation.

This detail is easy to overstate. It does not mean each of 50,000 people beat every one of the AI's 500 answers. The paper compared the centers of groups, not a sequence of one-to-one contests.

Models repeat themselves

A striking pattern emerged: models repeated themselves. Across GPT-4 sessions, "microscope" appeared in about 70% of response sets, and "elephant" in roughly 60%. GPT-4-turbo was more repetitive still: "ocean" appeared in more than 90% of its sets. The human sample behaved differently. Its most frequent words were "car" at 1.4%, "dog" at 1.2%, and "tree" at 1.0%.

There is no contradiction between a high DAT score and that repetition. The task rewards semantic distance among words inside a single answer; it does not directly reward originality across 500 separate answers. A model can discover that "microscope," "elephant," and several other favorites form a reliably distant set, then reuse those ingredients. Each individual list may travel far across semantic space even while the population of lists follows a familiar route.

Humans, taken together, generated a much broader collection of routes. Their average individual list was less divergent than GPT-4's, but their choices were less concentrated across the sample. Individual novelty and collective diversity are not the same thing.

The team also tested whether GPT-4's score could be changed without retraining it. Temperature is a model setting that changes how heavily generation favors the most probable next token. GPT-4's mean DAT score rose significantly with temperature; at the highest tested setting of 1.5, its mean reached 85.6, higher than 72% of the human scores. Prompting strategy mattered too. Asking GPT-4 to think through etymology-the origins and structures of words-produced the highest mean score among the tested strategies.

Limits of the study

These findings make any claim about "a model's creativity" conditional. Which model? Which version? Which prompt? Which temperature? How many samples? Who selects the final output?

The DAT measures divergent association, not creativity in full. The authors did not stop at lists of nouns. They used related automated measures to examine haiku, movie synopses, and flash fiction. GPT-4 scored above GPT-3.5 across those writing formats on their measure of semantic divergence, while human-written haiku and synopsis samples retained an advantage on key comparisons.

A brilliant work can be made from words that are semantically close. A technically distant combination can be useless, incoherent, or merely strange. Creativity normally asks for novelty and some form of fit: usefulness in engineering, insight in science, emotional force in fiction, or coherence within an aesthetic choice. The paper acknowledges that automated measures do not fully capture usefulness, convergent thinking, or expert judgment.

The DAT prompt has been public since 2021, and the training data of commercial models are opaque, so the researchers could not rule out prior exposure. A model may have learned examples of the task or discussions of how to score well. The 100,000 people also came with limited metadata-the researchers did not know their occupations or creative experience.

Why this matters for creatives

The study helps explain why language models can feel creative in daily use. They can produce remote associations quickly, consistently, and at negligible marginal cost. For a person staring at an empty page, that can be genuinely useful. AI for Creatives is increasingly about working with these tools as ideation partners rather than treating them as replacements.

But ideation is only one stage. Creative work also requires deciding what is appropriate, noticing when an apparently odd connection contains an insight, rejecting fluent nonsense, and revising a promising fragment until it belongs to a larger whole. The upper human tail matters because creative industries do not always hire for the average idea. Their value often sits in rare responses and in the judgment to recognize them. The same logic applies to Generative Art workflows: the tool produces options, but someone has to choose which one deserves to exist.

The model repetition result adds another reason to keep humans in the loop. If many people lean on the same system with similar prompts, each may receive something that feels fresh in isolation while the wider culture quietly becomes more alike.

The most defensible conclusion is therefore narrower-and more useful-than declaring a winner. On this brief test of divergent verbal association, GPT-4 beat the human average. The most creative half of the sample still had a higher average than every model tested, and the top 10% moved farther ahead. Machines have become very good at raising the floor of ideation. The ceiling still depends on uncommon human divergence, and on the human ability to know which unexpected association deserves to become something more.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)