Wednesday, August 26, 2026

In January 2026, the largest human–AI creativity comparison yet tested language models against 100,000 people. GPT‑4 beat the average person at generating unexpected associations — but even the most creative half of humanity outperformed every model tested, with the top 10% far ahead.

Must Read

Try the experiment before reading the result.

Write down ten English nouns that are as different from one another as possible. Not opposites. Not ten objects you can see. Not a tour through one category. Ten words whose meanings and uses take the largest possible leaps.

It sounds easy until “dog” makes “cat” available, “ocean” calls up “mountain,” and the mind discovers how strongly one thought pulls the next into its neighbourhood.

In a peer-reviewed study published in Scientific Reports on January 21, 2026, researchers used that small exercise to stage the largest human–AI creativity comparison yet. They tested several language models and set their responses against a dataset of 100,000 people.

GPT-4’s average score exceeded the average human score. That is the result built for a headline.

The rest of the distribution is the more interesting story. The mean score of the most creative half of the human sample exceeded every model mean in the researchers’ comparison. The upper quartile pulled farther ahead, and the top 10% opened a larger gap again.

This was not a clean victory for either side. It was a strong machine middle beneath a human upper tail.

A ten-word test behind a very large claim

The exercise is called the Divergent Association Task, or DAT. Participants are asked to provide ten single English nouns that are as different from one another as possible in every meaning and use.

A list containing cat, dog and hamster stays inside a tight semantic neighbourhood. Cat, thimble and liberty are more distant. The scoring system converts each valid word into a numerical representation based on the contexts in which it appears across a large body of language, then calculates the distances among the first seven valid words.

Seven words produce 21 possible pairs. Their mean semantic distance, multiplied by 100, becomes the DAT score.

The method was introduced by Jay Olson and colleagues in a 2021 paper in the Proceedings of the National Academy of Sciences. Across 8,914 participants, DAT performance showed moderate to strong relationships with established creativity measures, including tasks that ask people to invent unusual uses for common objects or bridge apparently unrelated ideas.

That validation is why the 2026 comparison is more informative than asking a chatbot to write a poem and voting on whether it feels inspired. Humans and models received the same instruction, and both were scored with the same automated rule.

It is also why the boundary needs to be stated early. The DAT measures divergent verbal association, one recognised component of creative cognition. It does not measure the whole process by which an idea becomes a useful invention, an affecting story, a workable design or a piece of art that survives revision.

What “beat the average person” means

The researchers selected exactly 100,000 English-speaking participants from a larger dataset. The sample was evenly split between men and women, with 20% drawn from each of five age bands: 18–29, 30–39, 40–49, 50–59, and 60 or older.

Most were in the United States, with smaller groups from the United Kingdom, Canada, Australia and New Zealand. All had reached the study through the public DAT website after hearing about it through news, social media or word of mouth.

That makes the sample unusually large and deliberately balanced on age and sex. It does not make it a random sample of the world’s population. The headline’s “humanity” is shorthand for this English-speaking volunteer dataset, not a census of human creative potential.

For the machines, the team tested OpenAI’s GPT-3.5, GPT-4 and GPT-4-turbo; Anthropic’s Claude 3; Google’s GeminiPro; and several open models, including Pythia, StableLM, RedPajama and Vicuna. An expanded comparison added systems released between January 2023 and June 2025.

Each model condition produced 500 responses. The researchers started a new conversation on every iteration so one answer would not influence the next. They applied the same scoring rule to both groups and excluded outputs with too few valid words.

In the main comparison, GPT-4 had the highest model average and exceeded the overall human mean by a statistically significant margin. GeminiPro’s mean was statistically indistinguishable from the human mean. GPT-4-turbo performed worse than the older GPT-4 in this task, a useful warning against assuming that a newer or more efficient model automatically becomes more divergent.

So GPT-4 did not sit across from one representative person and win a creative duel. Five hundred GPT-4 generations produced a higher mean DAT score than the mean of 100,000 human responses.

The most creative half changes the picture

An average compresses a distribution into one point. The researchers reopened it by sorting the human DAT scores and constructing benchmarks from the upper 50%, upper 25% and upper 10%.

For each human benchmark, a bar in the paper’s comparison represented the mean of a random 500-response subsample drawn from that segment. Every language model was likewise represented by the mean of 500 generated responses.

The average of the upper human half remained above the mean of every model in the curated comparison. The upper quartile scored higher again. The top decile defined the clearest separation.

This detail is easy to overstate. It does not mean each of 50,000 people beat every one of the AI’s 500 answers. Some model responses overlapped with, or exceeded, some high human scores. The paper compared the centres of groups, not a sequence of one-to-one contests.

It also does not establish a permanent human ceiling. The field moves quickly, and the study’s model set stops in June 2025. It tells us where those tested systems sat under those prompts and settings, not where every system available in August 2026 sits today.

Still, the result matters. As Silicon Canals noted in its earlier report on this comparison, “average” and “exceptional” are separate questions. The large human dataset makes the distinction visible rather than rhetorical.

The model found a high-scoring recipe

One result buried beneath the averages says something subtle about machine creativity.

Across separate GPT-4 sessions, “microscope” appeared in about 70% of response sets. “Elephant” appeared in roughly 60%. GPT-4-turbo was more repetitive still: “ocean” appeared in more than 90% of its sets.

The human sample behaved differently. Its most frequent words were “car” at 1.4%, “dog” at 1.2% and “tree” at 1.0%.

There is no contradiction between a high DAT score and that repetition. The task rewards semantic distance among words inside a single answer. It does not directly reward originality across 500 separate answers.

A model can discover that microscope, elephant and several other favourites form a reliably distant set, then reuse those ingredients. Each individual list may travel far across semantic space even while the population of lists follows a familiar route.

Humans, taken together, generated a much broader collection of routes. Their average individual list was less divergent than GPT-4’s, but their choices were less concentrated across the sample.

This is a useful distinction for creative work. A system can be strong at producing an unexpectedly varied answer on demand while still tending to produce similar answers when many people give it the same job. Individual novelty and collective diversity are not the same thing.

Temperature and prompts moved the score

The team then tested whether GPT-4’s score could be changed without retraining it.

Temperature is a model setting that changes how heavily generation favours the most probable next token. Lower settings make responses more deterministic. Higher settings permit less likely continuations and usually increase variation, although they can also reduce coherence.

GPT-4’s mean DAT score rose significantly with temperature. At the highest tested setting, 1.5, its mean reached 85.6, higher than 72% of the human scores. Repeated words also became less common as temperature increased.

Prompting strategy mattered too. The researchers tried instructions that encouraged different linguistic routes. For GPT-4, asking it to think through etymology, the origins and structures of words, produced the highest mean score among the tested strategies.

The finding makes any claim about “a model’s creativity” conditional. Which model? Which version? Which prompt? Which temperature? How many samples? Who selects the final output?

Those are not technical footnotes. They describe the creative system being evaluated. A human who changes the prompt, requests more candidates and recognises the one worth keeping is part of the process that produces the result.

This is divergent association, not creativity in full

The authors did not stop at lists of nouns. They used related automated measures to examine haiku, movie synopses and flash fiction. GPT-4 scored above GPT-3.5 across those writing formats on their measure of semantic divergence, while human-written haiku and synopsis samples retained an advantage on key comparisons.

The exercises offered supporting evidence that the DAT captures something relevant beyond a word list. They did not turn semantic distance into a complete theory of art.

A brilliant work can be made from words that are semantically close. A technically distant combination can be useless, incoherent or merely strange. Creativity normally asks for novelty and some form of fit: usefulness in engineering, insight in science, emotional force in fiction, or coherence within an aesthetic choice.

The paper acknowledges this. Automated measures do not fully capture usefulness, convergent thinking or expert judgment. A future benchmark would ideally combine computational scoring with human evaluation and test models on tasks that were kept private until the moment of assessment.

That last point matters because the DAT prompt has been public since 2021. The training data of commercial models are opaque, so the researchers could not rule out prior exposure. A model may have learned examples of the task or discussions of how to score well. The team’s public code and data repository improves transparency, but it cannot reveal what sits inside a proprietary training corpus.

The 100,000 people also came with limited metadata. The researchers did not know their occupations or creative experience. It is plausible that writers, musicians, editors and other practised creators were concentrated in the upper tail, but the study could not test that explanation.

AI lifts the middle; humans still matter at selection

The study helps explain why language models can feel creative in daily use. They can produce remote associations quickly, consistently and at negligible marginal cost. For a person staring at an empty page, that can be genuinely useful.

But ideation is only one stage. Creative work also requires deciding what is appropriate, noticing when an apparently odd connection contains an insight, rejecting fluent nonsense, developing a promising fragment and revising it until it belongs to a larger whole.

The upper human tail matters because creative industries do not always hire for the average idea. Their value often sits in rare responses and in the judgment to recognise them.

The model repetition result adds another reason to keep humans in the loop. If many people lean on the same system with similar prompts, each may receive something that feels fresh in isolation while the wider culture quietly becomes more alike.

The most defensible conclusion is therefore narrower, and more useful, than declaring a winner. On this brief test of divergent verbal association, GPT-4 beat the human average. The most creative half of the sample still had a higher average than every model tested, and the top 10% moved farther ahead.

Machines have become very good at raising the floor of ideation. The ceiling still depends on uncommon human divergence, and on the human ability to know which unexpected association deserves to become something more.

 

- Advertisement -spot_img
- Advertisement -spot_img
Latest News

Karen Arnold followed 81 high-school valedictorians for 14 years and found that nearly 90% entered professional careers—but most pursued conventional paths within established institutions....

A valedictorian has solved a difficult problem repeatedly. Across several years, subjects and teachers, that student has understood what...
- Advertisement -spot_img

More Articles Like This

- Advertisement -spot_img