
Good writing has become suspicious.
Last year, I did a crash course on the politics and policy of AI. The final assessment included an essay question. The note accompanying it said: AI use isn’t prohibited, but you must document how you used it and submit that alongside your answer. I wondered: how would they actually assess that? The em dash? The “delve”? The contrast framing that AI loves almost as much as it loves bullet points?
This is the new anxiety for anyone who writes professionally.
The Surface Problem
People have caught on to the tells. Em dashes, transitional phrases, the particular way AI lands a paragraph: all of it has been catalogued, memed, and turned into a checklist. That checklist is now baked into the prompting process itself. “Avoid em dashes. Don’t use the word ‘delve.’ Don’t start sentences with ‘In conclusion.'”
But while you can swap the vocabulary, you cannot change the underlying skeleton that persists.
A research paper published this year, StoryScope: Investigating Idiosyncrasies in AI Fiction, tested exactly this. Researchers at the University of Maryland built a pipeline to analyse 61,608 stories written to the same prompts by five major AI models and human authors. They extracted 304 narrative-level features. Not word choice or sentence rhythm, but structural decisions: how plots are resolved, how time is handled, how themes are communicated.
Further, they took AI-generated stories, ran them through a span-level rewriting tool designed to strip out surface AI artifacts (the clichés, the purple prose, the redundant exposition) and tested whether the narrative classifier could still detect them.
It could. Accuracy dropped by just 1.6 percentage points.
What the Research Actually Found
The specific patterns are worth naming, because they map precisely onto what good editors have been feeling but couldn’t put a finger on.
- AI stories over-explain their themes. Narrators state the moral explicitly 77% of the time, versus 52% for human authors. Dialogue becomes philosophical debate.
- References to other works stay deliberately vague, allusions rather than named texts, as though AI is wary of committing to a specific intellectual lineage.
- AI plots are clean to a fault. 79% of AI stories have no subplots, versus 57% for humans.
- Resolutions skew toward internal acceptance, the character learns a lesson, understands something about themselves. Rather than the messier, more ambiguous endings human writers favour.
- Causal chains are continuous and tidy. Nothing is left dangling.
And this surprised me most: AI conveys emotion through bodies and environments, not direct labels. A character doesn’t feel afraid; their chest tightens, the lamplight dims, cold sweat surfaces. Humans, by contrast, use explicit emotion labels nearly four times more often than AI.
The assumption that AI writes flatly and humans write with sensory richness turns out to be exactly backwards.
Human stories, meanwhile, break the fourth wall, address the reader directly, reference specific authors and texts, and carry nonlinear structures: flashbacks, time jumps, revelations that force you to reread earlier scenes. They occupy what the researchers called rarer narrative space. The human story was the most structurally unusual of all six versions written to the same prompt 57.8% of the time. By chance, that figure would be 16.7%.
Human writing is — surprise, surprise — more original.
The Prescriptive Reflex
This connects to something I’ve noticed in AI review tools. If you submit a descriptive essay that examines a problem without prescribing solutions, they’ll flag it as incomplete. Because AI defaults to the prescriptive mode. It assumes every writing owes the reader a set of instructions.
The StoryScope findings explain why. AI resolves narratives through internal understanding and acceptance. It closes loops.
A piece that ends with the problem examined but not solved, that trusts the reader to sit with the ambiguity, registers as unfinished. This is a structural logic baked into how AI conceives of an argument’s purpose.
And I once experienced this first-hand: I fed one of my essays into an AI detector. It came back flagging high AI probability because it was too well-written. When I pushed back and asked whether it was suggesting only AI could write that well, the tool immediately reversed course. The confidence score, it turns out, is partly a UI decision. ‘We think, maybe, 60%‘ doesn’t sell. So they sell certainty.
What Persists After the Polish
The things that survive a surface rewrite are structural than stylistic. The neat three-part arc. The clean resolution that ties everything together. The thematic unity that leaves nothing unexplained.
The features StoryScope identifies as core AI markers are the things you cannot fix by rephrasing a sentence.
That meandering paragraph, the one that detours into an anecdote before returning to its point, is not a flaw in most non-technical writing. It’s what holds the reader. Whereas extreme AI optimisation processes efficiently in a human mind. But disappears just as quickly.
A growing volume of online content is now running to 2,000 or 3,000 words. The only reason to write at that length, unless you consider yourself a Shakespeare, is if you’re writing for AI indexing rather than human reading.
Whether it works is not the point. Content written for machines reads like it was written for machines. And this research shows how.
The Diminishing Returns of Iteration
There’s a belief that AI orchestration equals authorship. That given enough iterations, you can improve any draft into something exceptional.
I disagree.
The first AI review does real work: catches structural gaps, logical non-sequiturs, clunky transitions. The second yields less. By the third pass, you’re likely making it worse in the ways that matter most. Each iteration pushes the piece further toward AI’s structural defaults. In other words, as far from human writing as possible.
Sherlock Holmes, in “The Adventure of the Norwood Builder,” identifies exactly this failure mode in a criminal who had constructed an almost perfect trap. Almost. As Holmes puts it: “He had not that supreme gift of the artist, the knowledge of when to stop. He wished to improve that which was already perfect — to draw the rope tighter yet round the neck of his unfortunate victim — and so he ruined all.“
That judgment of knowing when a piece is done separates a writer from a prompt engineer.
As compute costs make headlines, as companies pour money into AI tools and find the returns murkier than the projections, the ability to use AI minimally and extract maximum value from it will matter more. Traditional writers, who can identify the one useful iteration and leave the rest alone, are better positioned here than orchestrators whose entire workflow depends on multiple iterations.
The North Star
There is no going back to pre-AI writing. The idea is to retain the judgment, the point of view, and the originality that made the writing worth reading in the first place.
The StoryScope researchers put a number on human narrative originality: a mean rarity percentile of 0.71, against 0.49 for AI. That gap is worth protecting.
Because we’re writing for a human, eventually. That’s the north star. Everything else, the iteration count, the AI mentions, the surface polish, is just technique deciding whether to serve the work or replace it.
Leave a comment