Chapter 1 of 7. In May 2026 the four published articles on this site carried em-dashes at 8 to 13 per thousand words. Edited human prose runs 1 to 3. This chapter is about the mechanism behind that gap, because a tell you can only feel is a tell you cannot fix.
The number that started it
During a routine devlog audit on 22 May 2026, my writing advisor flagged the em-dash density across the four articles then published on this site. I pulled the pieces and counted by hand. The count came back at 8 to 13 per thousand words, against a baseline of 1 to 3 for normal edited prose. From the write-up of that audit:
The four pieces weren't stylistically heavy on em-dashes. They were statistically anomalous in a way that had a specific cause: LLM-assisted drafting leaves a fingerprint, and I'd been publishing without checking for it.
Nothing in those articles was wrong. Every sentence parsed. The facts held. No reader had complained. The prose still carried a measurable signature, and I had shipped it four times without noticing. I sat exactly where I suspect many writers using AI assistance sit: I could smell generated text in other people's work and could not see it in my own.
Why does AI writing sound flat?
Because of how the text is made. A language model writes by choosing the next token: it scores every candidate continuation for plausibility given the text so far, then picks from the top of that distribution. The most plausible continuation is by construction the most typical one. Make that choice once and you get a reasonable word. Make it several thousand times per draft and the output regresses toward the centre of the training data: the most common cadence, joined by the likeliest connectives. Flatness is not a malfunction. It is the objective being met.
The model also holds no running totals. It does not know it has opened three sections the same way or spent its fourth parenthetical aside in two paragraphs. Each sentence is produced from local context, not from a budget for the whole piece. Averaged choices at the sentence level, no memory of density at the document level. Those two properties together produce the texture readers describe as flat.
There is now a measured version of the claim. A 2026 preprint from UC Berkeley's D-Lab (van Nuenen, arXiv:2604.22142) had three frontier LLMs rewrite 300 personal narratives under three prompt conditions and tracked 13 linguistic markers, from function words to contraction use. The finding, verbatim: "Across models and prompt conditions, LLM rewriting produces a consistent pattern of stylistic normalization." Prompting for voice helped less than you would hope: "Voice-preserving prompts reduce the magnitude of the changes but do not eliminate their direction." One author, not yet peer-reviewed, so weight it as a preprint. The direction result is still the important part. Every model tested pushed the narratives toward the same centre, whatever the instruction. The flattening is a property of the rewriting, not of any one product.
That frame is what this course means by a tell. The definition from the original audit write-up, borrowed for all seven chapters:
An AI-prose tell is a surface feature that trained language models produce at rates statistically higher than edited human prose. Not because the model is doing something wrong (the outputs are grammatically correct, often fluent), but because the model is optimising for plausibility and coherence, and certain constructions are plausible and coherent at a frequency that human writers don't naturally hit.
Read that twice, because the word doing the work is "rates". No single em-dash convicts a draft, and an isolated hedge is not an error. The signal lives in density, and density is invisible to a sentence-by-sentence read.
What the tells look like on the page
Ten numbered tells went into the catalogue that came out of that audit. The five families below are the ones the write-up names as examples, quoted here as specimens, not used:
em-dash density flag above 3 per 1,000 words
hedging qualifiers "it's worth noting", "importantly"
throat-clearing "in order to", "with that in mind"
performative warmth "great question", "absolutely"
false continuity "of course", "naturally"
Assemble them and you get a paragraph like this one, invented for illustration:
It's worth noting that the migration is, of course, fully reversible — in
order to roll back, you simply restore the snapshot. Importantly, the
process is seamless, and naturally the same steps apply in reverse.
Every construction in it is grammatical. A copy editor checking for errors finds none. That is the trap: each tell is, in the audit's words, "borderline in isolation. At the density a generative draft produces them, the cumulative signal is detectable."
Readers are already primed to detect it. In a January 2025 survey of 1,000 Americans run via Pollfish by the content agency Hookline&, 82.1% said they can tell at least some of the time when an article they are reading has been written by AI. Handle that number with care. It is self-reported belief rather than a tested detection task, and it counts everyone who answered above "Never", including people who said "Rarely". What it establishes is narrower and still useful: your audience believes it can spot you, so it is looking. Publish at 11 em-dashes per thousand words and some readers will feel the wrongness without being able to name it.
The quietest tells are the most dangerous, and the lived proof came from running the audit script retroactively over my own four published articles. From the write-up:
It also surfaced two patterns I hadn't noticed on re-read: a cluster of "of course" and "naturally" constructions in the second article, and a run of "in order to" phrases that could all be shortened to "to." Both had survived because they're quiet.
Three instances each across a 900-word piece. I had re-read those articles before publishing. The loud tells got caught on those re-reads; these quiet ones slid under my eye on every pass. Counting is the only way to see them.
Why this costs more than style points
If flat prose only cost you aesthetics, a busy writer could reasonably decide not to care. The cost is larger than that.
Google's helpful-content documentation defines the target plainly: "People-first content means content that's created primarily for people, and not to manipulate search engine rankings." The same page draws the line for automation: "If you use automation, including AI-generation, to produce content for the primary purpose of manipulating search rankings, that's a violation of our spam policies." And the spam policies themselves define scaled content abuse as "when many pages are generated for the primary purpose of manipulating search rankings and not helping users", listing as an example "Using generative AI tools or other similar tools to generate many pages without adding value for users".
Averaged prose walks toward that definition by default. The centre of the training distribution contains no value that was not already published a thousand times, so a draft regressed to that centre adds none.
The positive case is on the same helpful-content page, in the framing Google names "experience, expertise, authoritativeness, and trustworthiness, or what we call E-E-A-T". The documentation adds: "Of these aspects, trust is most important. The others contribute to trust, but content doesn't necessarily have to demonstrate all of them. For example, some content might be helpful based on the experience it demonstrates, while other content might be helpful because of the expertise it shares."
Connect that to the mechanism and the stakes become concrete. Averaging strips specifics first, because specifics are by definition atypical. A dated incident with a hand-counted density figure sits nowhere near the centre of any distribution. Experience is exactly the signal flattening removes, and experience is a named component of the framework Google says it rewards. This chapter opens with one for precisely that reason.
How to fix it: count first, rewrite second
Before 22 May my method was the one most writers use. The write-up describes it, then delivers the verdict:
The problem I had before 22 May was that I was catching these by feel, on re-read, inconsistently. That's not a process. It's a mood.
The fix that held on this site has two halves, and the rest of this course builds them out in order.
The first half is a written catalogue. The style guide that came out of the May audit lives at the repo root and gives every tell four fields:
- the pattern (what to look for)
- why it reads as generated (the mechanism, never the label alone)
- the threshold, where one exists
- the fix
The fix column is the differentiator. In the write-up's words, "A style guide that names problems without naming remedies is documentation for its own sake." The em-dash entry does not say to use fewer em-dashes. It says to check whether the parenthetical is load-bearing; if it is not, cut the aside; if it is, rewrite the sentence so the aside becomes the main clause. That is a repair you can execute, not a vibe you can fail at.
The second half is a counter. audit-ai-tells.sh takes a file path, counts pattern instances against word count and exits non-zero on any breached threshold. Real output from a real run:
[audit-ai-tells] em-dash density: 11.2 per 1k words (threshold: 3)
→ line 14: "The reflex I'd been using to review code — does this function…"
[audit-ai-tells] hedging: 4 instances
→ line 22: "it's worth noting that the original version"
[audit-ai-tells] FAIL (2 patterns exceeded threshold)
Note what the script refuses to do. It does not rewrite flagged sentences, because "Automated rewrites on flagged sentences produce their own tells, often subtler ones, so the rewrite step stays manual by design." It also does not catch semantic tells such as fabricated specifics or marketing assertions. Surface-pattern density is the whole scope. Later chapters take up what sits outside it.
One honest limit from the same incident. The four over-budget articles were never retroactively edited: "Retroactive edits to published pieces create a different kind of trust problem. The record stops being reliable." The audit is a gate on everything going forward, not a rewrite of the past.
- Name the tells in writing
A catalogue in your repo: pattern, mechanism, threshold, fix. Chapter 2 walks the vocabulary layer; Chapter 3 covers rhythm, the layer readers detect first.
- Count deterministically
A script, not a feeling and not another model asked to judge. Chapter 7 builds the full pre-publish audit, including what happened when models were asked to count their own tells.
- Rewrite by hand
The process that produced the tell cannot be trusted to remove it. Chapters 4 and 6 show where the draft actually starts and where prompting stops working.
- Put your experience back in
Averaging removed the specifics. Chapter 5 is about restoring the version numbers and documented failures no model could invent.
The transferable rule
Flat is not a defect the next model release will patch away. It is what next-token selection converges on when nothing external pulls against it, and the Berkeley result says prompting pulls weakly. So stop treating AI-sounding prose as a mystery of taste. Treat it as a measurable property with a structural cause: averaged choices at the sentence level and no memory of density at the document level.
The gap between the model's cadence and yours is not a problem to be embarrassed about. It is the raw material of this course. Chapter 2 opens the catalogue: the vocabulary tells one by one, each with the mechanism that produces it and the repair that removes it.
Count your own baseline
Run the count-first method from this chapter on your own archive. Pick your three most recent published pieces, count the em-dashes in each, and turn the arithmetic into a small script you can rerun on every future draft.
Expected behaviour
- A script that takes a file path and prints em-dash density per 1,000 words
- Density figures recorded for at least three of your own published pieces
- A written baseline note stating your normal range and the threshold you chose
- The script exits non-zero when a piece breaches your threshold
PROVE IT Run the script on your most recent piece and paste the output line showing its density against your threshold.
Why does next-token generation drift toward flat prose?
Above what em-dash density per 1,000 words does the catalogue flag a draft?
Why can a sentence-by-sentence re-read miss a tell that a counter catches?
Show answer
A tell is a surface feature produced at rates statistically higher than edited human prose, and every individual instance is grammatical and borderline. The signal lives in density across the whole piece, which no local read can see. The quiet tells on this site survived repeated re-reads and only surfaced when the script counted them.
↺ re-read: “Why does AI writing sound flat?”