Stories on a Diet

When talking about the risks of AI we usually warn about confabulations (hallucinations) or cognitive outsourcing, but there is a more insidious risk connected to LLMs that is often overlooked: homogenisation. AI-models are pulled toward a grey average in their generation. Researchers at the University of Maryland and Google Deepmind found that stories written by several AI-models tend the cluster around a sameness not shared by human writing. This goes beyond stylistic elements such as using AI-specific words and shows that narrative features of AI-stories also hover around a grey average. This research should warn policy makers and teachers (again) to make sure students keep getting exposed to rich and diverse human-written texts during their education, especially when an institution has embraced AI in their curriculum.

In order to research the storytelling capabilities of LLMs that goes beyond stylistic preferences, the researchers developed a method called StoryScope. With StoryScope stories are compared on narrative features based on Narrabench rather than using specific words or phrases. Narrabench is a classification system to be used as a benchmark for narratives generated by LLMs to test their narrative understanding. It uses twelve primary narrative features (Perspective, Style, Time, Revelation, Paratext, Motivation, Agent, Network, Event, Plot, Structure, and Setting) in four narrative dimensions (story, narration, discourse, and situatedness). The researchers behind Narrabench stress the importance of LLMs being able to generate good narratives as they pervade our everyday lives: “They can be used to entertain, inform, persuade, and maintain shared beliefs in communities across generations.” And they will. And they are.

Narrabench taxonomy graph
The twelve primary narrative features of Narrabench. The colours represent the ‘Big-4’ narrative dimensions (story [in red], narration [in green], discourse [in blue], and situatedness [in yellow]. The shades indicate how much the feature was already present in previous benchmarks. (“NarraBench: A Comprehensive Framework for Narrative Benchmarking”)

So, the StoryScope researchers used Narrabench to determine literary features in AI-generated stories: Would you be able to identify AI-generated stories when looking at these literary elements rather than stylistic ones? They had human writers and five AI-models (Gemini 3 Flash, Kimi K2.5, DeepSeek V3.2, Claude Sonnet 4.6, and GPT-5.4) to write a story by using prompts like the one below:

Representative benchmark story-generation prompt. (“StoryScope: Investigating idiosyncrasies in AI fiction”)

The researchers used 10 of the 12 Narrabench features, leaving the situatedness dimension out (this covers social context and audience expectations) for their research. Around 60,000 stories were created and they used another LLM to analyse the outcomes and bring the number down to 600 usable stories. The narrative features were turned into vectors (lists of numbers) to compare for overlap and uniqueness of a story. These vectors were put on a two-dimensional projection as given below.

Narrative feature vectors projection
Simplification of the 10 Narrabench narrative features found in the produced stories. Each dot is a vector representation and the closeness to another dot indicates similarity on these features. It does not tell how good the story was. All LLMs are quite apart from human writing. Claude is the most distinct of the LLMs. (“StoryScope: Investigating idiosyncrasies in AI fiction”)

They found that “the five models occupy a tight cluster that is well-separated from human stories” and concluded that “LLMs consistently reduce collective diversity.” Each LLM also has its own ‘fingerprints’: tendencies in storytelling not much shared by the other models. Claude is ‘cool and restraint’, GPT is ‘gossipy’, DeepSeek tends to give much information early, Gemini has the tidiest endings, and Kimi has the least fingerprints and is thereby the most generic.

There were a couple of core narrative choices that separate LLM stories from human stories. AI stories tend to be more explicit and there is less ‘between-the-lines.’ There are less subplots and time jumps such as flashbacks and flashforwards. Rather than naming specific feelings, AI focuses on the body and the senses extensively to show emotions. Human writing refers more to specific texts whilst LLMs vaguely allude to other writing. Humans break the fourth wall, which is addressing the reader directly, more often than AI does. Finally, human writing spans over more locations, they have more dialogue relative to the story, and present more morally ambivalent protagonists.

These results should not come as a surprise. Sophisticated as LLMs can be, they are still prediction machines going for tokens with a high probability. It is not a specific model that has this tendency, it is the generative AI architecture. You can’t prompt your way out of this problem, how rich that prompt might be. As LLMs predict the next token based on the previous one and use probability calculations, AI generated text is pulled toward an average centre.

Other researchers (also, again, partly from Google Deepmind) had already found that when LLMs edit essays, they tend to advise a grey average, not only in style but also in opinion. In “How LLMs distort Our Written Language” researchers showed that when AI was used to correct errors and improve clarity of a human text, the text became less creative, not in the voice of the author, and even altered its semantic meaning, including the opinion of the writer, moving towards a more neutral stance on an issue. The shift is not programmed in the LLM and the researchers think the changes are caused by LLM-preferred linguistic patterns. LLMs are trained in text patterns stored numerically in their vector space. These patterns determine the most probable word that should follow a generated text to create a coherent whole.

LLM-generated revisions display a larger and more consistent semantic shift than human-written revisions of the same essays. The first graph (in grey) shows human revisions on a text. The green graphs are LLM revisions. The importance lies in the direction of the arrows. The green graphs have arrows going mainly into one direction whereas the grey one has arrows going in multiple directions. (“How LLMs Distort Our Written Language”)

It has also been known for quite some time that when you train new AI-models on synthetic data, information produced by other LLMS, they will eventually implode during training. This is a phenomenon is called model collapse: the training data has lost too much variety and its rarer options for the trained model to make sense. This is also the reason why Anthropic has started to destroy books to scan and feed to their new models rather than have other LLMs write new material to train on. Synthetic data is simply less diverse, less nutritious, less useful than human generated texts.

StoryScope, opinion distortion, and modal collapse are all a warning. Too much exposure to LLM texts will expose users to poorer texts and move them towards a grey average in language, narrative features, and even opinion. Not because the prompts were poorly written but because it is in the LLM’s nature to cluster around a grey average. The researchers of “How LLMs Distort Our Written Language” caution the educational field: “students who enjoyed writing before using AI tended to augment AI outputs with their own thinking, whereas students who struggled with writing were more likely to adopt AI-generated text wholesale.” It is vital for the educational field to teach students to become confident and thereby independent writers to make sure they will use AI responsibly and effectively later in their careers. This requires independent reading of rich, diverse human texts and independent writing. Not only to keep our own writing varied and interesting, but also to keep our ideas and narratives nutritious, distinctive, and worthwhile.


Sil Hamilton, Matthew Wilkens, Andrew Piper, “NarraBench: A Comprehensive Framework for Narrative Benchmarking,” 10 October 2025, https://arxiv.org/abs/2510.09869

Jenna Russell, Rishanth Rajendhran, Chau Minh Pham, Mohit Iyyer, John Wieting, “StoryScope: Investigating idiosyncrasies in AI fiction,” 10 August 2026, https://arxiv.org/abs/2604.03136

Marwa Abdulhai, Isadora White, Yanming Wan, Ibrahim Qureshi, Joel Z. Leibo, Max Kleiman-Weiner, Natasha Jaques, “How LLMs Distort Our Written Language,” 19 March 2026, https://arxiv.org/abs/2603.18161

Photo by Rain Bennett on Unsplash

Categories: , ,