Readers Prefer ChatGPT Stories to Human Ones, Until They Find Out Who Wrote Them
A study published in Judgment and Decision Making by Sydney Sears and Deena Skolnick Weisberg presented over 2,500 readers with six short stories of approximately 1,000 words each: three from literary magazines, three generated by ChatGPT 4.0 using prompts based on the theme, style, and perspective of the original human stories. Participants failed to distinguish who wrote what; their performance remained at chance level, even in direct comparisons between human and AI stories.
ChatGPT stories received higher scores for both perceived quality (1.54 vs. 0.97 on a scale from -3 to +3) and immersion (1.42 vs. 1.00). The bias lies in the label: when readers believe the author is human, scores rise for both groups. When they think it’s AI, they drop, regardless of who actually wrote it.
Why this matters to you. If you use AI to write, the bottleneck isn’t output quality; it’s reader perception. AI-generated texts tend to be more fluid, easier to process, and more optimistic. People confuse smoothness with quality. Disclose AI involvement and expect a perception penalty that has nothing to do with the actual content.
Personal attitudes toward AI matter as much as the text itself. Those with a positive view of models give higher scores to everything, and even higher scores to AI texts when they know what they are. Skeptics do the opposite. Experience with narrative doesn’t help distinguish; experience with AI systems does, but only weakly.
In Detail
The study comprises three experiments. In the first, 1,682 participants read one story each and rated it for perceived quality and immersion. Half received correct information about the author, half received incorrect information. The design is clean: it separates the text effect from the label effect, and both carry weight.
In the next two experiments (905 participants total), researchers made the task harder: each person reads both a human and an AI story and must identify which is which. Even with direct comparison, performance remains at chance. This is the most convincing finding: distinguishing isn’t difficult; it’s virtually impossible for an average reader on short texts.
What makes AI texts more appealing.
The researchers propose a concrete hypothesis: AI-generated texts are smoother, easier to read, and more emotionally positive than human ones. Cognitive psychology tells us that people prefer material that’s easy to process (so-called processing fluency). Quality literary narrative, by contrast, is often intentionally difficult, asking readers to work to extract meaning. The short story format favors AI: maintaining coherence across 1,000 words is a very different task from sustaining a novel.
Attribution bias cuts both ways.
The label effect is symmetrical. Say “a human wrote it” and scores rise for both text types. Say “ChatGPT wrote it” and they fall. A reader’s preexisting attitude toward AI modulates the effect: those who trust models reward declared AI texts, while skeptics penalize them beyond reason. A similar result had emerged in an earlier study on AI-generated poetry.
What this means for AI writers.
For those using AI in professional writing, the study separates two questions usually conflated: “Is the output good?” and “How will it be received?” The answers diverge based on transparency. If you don’t disclose AI use, the text is evaluated for what it is. If you do, it’s evaluated for what the reader thinks it is. The difference is measurable: we’re talking significant shifts on a seven-point scale.
An earlier study by Stony Brook University and Columbia Law School (October 2025) showed that with simple prompts, expert readers preferred human texts. But when models were trained on individual author styles, experts preferred AI texts eight times out of ten for style imitation and two times out of two for quality. Model specialization shifts the threshold.
Limitations.
The short format favors AI: 1,000 words is territory where narrative coherence is manageable for a model. On longer formats, performance changes. Participants were recruited online and weren’t literature experts (literary experience didn’t help anyway). Experience with AI systems correlated positively with the ability to recognize AI texts, but the effect was modest. Data and materials are available on the Open Science Framework.