On this page

Chapter 2 of 7. Prerequisite: Chapter 1, where an audit of four published articles found em-dash densities of 8 to 13 per thousand words against a normal edited range of 1 to 3. This chapter names the rest of the catalogue, then shows what happened when one banned punctuation mark came back as fourteen identical clause tails.

Search for an AI word list and you get a different one every month. The lists keep growing, and the tools that promise to scrub a draft against them hand back prose with a different accent of wrong. Neither fact makes the lists useless. It makes them an inventory of symptoms rather than a diagnosis.

This chapter gives you the working inventory from this site's audit, then shows why mechanical substitution fails. The inventory comes from the pre-publish style audit built after the em-dash discovery: real tells, caught in shipped technical articles, on the site you are reading.

The AI writing tells list, from a 2026 audit of shipped articles

On 28 May 2026 the patterns found in the four published pieces were codified into a repo-root STYLE-GUIDE.md as ten numbered items. The core vocabulary entries, with the phrases quarantined in a code block so this chapter's own prose stays clean:

TELL                           EXAMPLE PHRASES OR SHAPE                THRESHOLD (where one exists)
em-dash density                the sentence — interrupted — resumed    flag above 3 per 1,000 words
hedging qualifiers             "it's worth noting", "importantly"      none set
transitional throat-clearing   "in order to", "with that in mind"      none set
performative warmth            "great question", "absolutely"         flagged at section opens
false continuity markers       "of course", "naturally"                none set

Two weeks later, the article auto-pipeline on the same site shipped with its own scrubber, and its list already differed:

opening hedges                 "Indeed,", "Notably,"
closing summaries              "In conclusion,"
engagement filler              "Let's dive in."
generic positivity             "powerful", "seamless", "comprehensive"
rule-of-three rhythm           three parallel items, paragraph after paragraph
sentence-length uniformity     every sentence within a few words of the same length

Look at the last two rows of the second list. They are not vocabulary at all. They are structure, and no find-and-replace can touch them. Hold that thought for Chapter 3.

A label is not an entry: the four-field format

The reason STYLE-GUIDE.md works where a pasted word list would not is that each of its ten items carries four fields, not one.

  1. The pattern

    What to look for, stated concretely enough that a script can count it. A phrase, or a sentence shape.

  2. Why it reads as generated

    The mechanism behind the label. Hedging qualifiers pad an assertion toward plausibility. Performative warmth is chat register leaking into a document. False continuity markers assert an agreement no reader ever made. Knowing the mechanism is what lets you recognise the tell in a new costume later.

  3. The threshold, where one exists

    One instance of anything proves nothing. The em-dash entry flags above 3 per 1,000 words because the tell is density, never presence.

  4. The fix

    The remedy, named. From the audit write-up: "A style guide that names problems without naming remedies is documentation for its own sake."

The fix field is where the blacklist mindset dies. The em-dash entry does not say "use fewer em-dashes". It says: check whether the parenthetical is load-bearing. If it is not, cut the aside. If it is, rewrite so the aside becomes the main clause. Read that instruction closely and you see the real target. The tell was never the punctuation mark. The tell is the density of asides, and the dash is only how the model spells them.

There is a boring operational lesson in the file's history too. Before STYLE-GUIDE.md existed, the em-dash guidance lived as a one-line inline note in a drafting workflow doc. From the same write-up: "That note wasn't findable at publish time. Moving it into a canonical reference that the toolchain can point at is the difference between a note-to-self and a constraint."

The list grows because it is a snapshot

The two lists above were written a fortnight apart, on the same site, by the same author, and they disagree on membership. That is not sloppiness. Vocabulary lists are a snapshot of the model's current habits, not a closed set. Models get updated and drafting prompts change. The surface habits shift underneath any list you froze.

This is why the growing-list feeling is accurate and why it does not matter. The membership churns. The mechanisms underneath it barely move. Every model this site has drafted with, optimising for plausible, coherent continuation, has overproduced softened assertions and parenthetical asides in some spelling. Catalogue the mechanism and the next costume is recognisable on sight. List only the words and every model release resets you to zero.

Ban the punctuation and the tell migrates

Here is the incident that settles the argument, from this site's devlog in July 2026.

A chapter of the newsletter-cpanel-security course was drafted under a hard em-dash budget, and the deterministic recount came back perfect: 0.00 em-dashes per 1,000 words of prose. The same chapter carried 14 trailing appositive clauses, a density of 6.6 per thousand words. The shape, reconstructed:

The token lives in CI, which is why the build never sees it.
The loader reads frontmatter first, which is where the schema drifted.
The flag defaults to false, which is the safe direction.

Compare that against the em-dash entry's mechanism field. It is the same parenthetical pivot, at an aside density more than double the 3-per-1,000 threshold the original em-dash entry flags at. The cadence had been re-punctuated, not removed. The metric read clean. The tell survived.

That is the blacklist failure in a single artefact. The model's habit is the aside. The word list only knew one costume, so enforcement at the vocabulary layer pushed the habit into a spelling the counter could not see. Swapping tell-words by hand does the same thing more slowly.

External evidence does exist for the narrower claim that surface instructions fail to redirect model style. A 2026 single-author preprint from UC Berkeley's D-Lab (van Nuenen, arXiv:2604.22142, not yet peer-reviewed) had three frontier LLMs rewrite 300 personal narratives under three prompt conditions and tracked 13 linguistic markers. Two findings, verbatim:

Across models and prompt conditions, LLM rewriting produces a consistent pattern of stylistic normalization.

Voice-preserving prompts reduce the magnitude of the changes but do not eliminate their direction.

The voice-preserving condition cut the drift by about a third, and the residual effect (|d| = 0.76) still approached the conventional threshold for a large effect. Scope caveats apply: it is a single preprint, and the task was rewriting personal narratives rather than drafting technical prose. What it supports here is exactly what the appositive incident showed from the other side. Telling the model what to avoid changes how far the prose drifts. It does not change where the prose is drifting to.

Humanisers over-correct into a different fake

The commercial answer to the word list is the humaniser: paste your draft, get back a version with the flagged vocabulary swapped out and some engineered irregularity swapped in. The output tends to read like this:

Honestly? Setting up CI is a bit of a beast. But once you crack it? Total game changer.

Nobody's technical writing sounds like that either. The register has moved from one manufactured accent to another, and a reader who could smell the first can smell the second.

The style audit on this site refuses to auto-rewrite for precisely this reason. The script flags and stops, and the rewrite stays manual by design. The audit write-up gives the reason: "Automated rewrites on flagged sentences produce their own tells, often subtler ones." A substitution pass at scale is itself a pattern. A blacklist run backwards is still a blacklist.

The published policy test was never a word list

One more reason to stop treating vocabulary scrubbing as the goal: the policy everyone is nervous about does not mention vocabulary. Google's spam policies page defines the relevant abuse at the level of purpose:

Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users.

Its listed example of the practice is "Using generative AI tools or other similar tools to generate many pages without adding value for users", and the dedicated generative-AI guidance page repeats the same test in almost the same words: "using generative AI tools or other similar tools to generate many pages without adding value for users may violate Google's spam policy".

Purpose and value. A draft that swaps its flagged words for unflagged ones has changed its vocabulary and left its value exactly where it was. Against the published test, the scrub did zero work.

So aim the word list at the audience it can actually serve. Vocabulary tells cost you readers, the humans who close the tab at the second hollow qualifier because the prose smells generated. The list is a reader-trust instrument. Treat it as a policy-compliance instrument and you end up laundering vocabulary on pages that still add nothing.

What replaces the blacklist

Two things, both already visible in this chapter.

First, mechanism-aware entries instead of bare words: pattern, why it reads as generated, threshold where one exists, fix, all in one canonical file. An entry in that format survives the costume change, because you catalogued the habit rather than its current spelling.

Second, measurement at the right layer. The appositive incident proved a word-level metric can read a perfect 0.00 while the tell sits one layer down, in the rhythm of the sentences themselves. Counting words was never going to be enough, and the two structural rows in the scrubber's list already said so.

That layer is Chapter 3: rhythm and structure, the parenthetical pivot, sentence-length uniformity, and the register experiment where two reasonable-sounding voice changes both lost a blind test to the original.

// EXERCISE

Write your own four-field catalogue

Replace the pasted word list with a catalogue built from your own drafts. Create a canonical tells file at your repo root and give every entry the four fields from this chapter: the pattern, why it reads as generated, the threshold where one exists, and the fix.

Expected behaviour
  • A canonical file at the repo root, findable at publish time
  • Every entry carries all four fields, with the fix stated as a repair you can execute
  • At least one entry states its threshold as a density rather than a raw count
  • At least one entry describes a sentence shape instead of a phrase, so it survives a costume change
  • Each pattern is concrete enough for a script to count

PROVE IT Grep one entry's pattern across your most recent draft and paste the match count next to that entry's threshold.

// CHECKPOINT — BEYOND THE BLACKLIST
multiple choice · auto-checked

A chapter drafted under a hard em-dash ban came back at 0.00 em-dashes per 1,000 words. What did the audit find?

exact answer · auto-checked

At what density per 1,000 words did the trailing appositive clauses appear in that em-dash-free chapter?

open · self-checked

Why does the style audit refuse to auto-rewrite flagged sentences?

Show answer

Automated rewrites produce their own tells, often subtler ones, and a substitution pass at scale is itself a pattern. The only rewrite that removes a tell goes through the mechanism: find the aside, decide whether it is load-bearing, restructure by hand. That judgement does not automate.

↺ re-read: “Humanisers over-correct into a different fake

Lived experience

Sources

  • Google Search's guidance on using generative AI content on your website
    Google Search Central (Google for Developers)
    The dedicated generative-AI guidance page; evidence that Google's published test is value added for users, not the presence or absence of particular vocabulary
    developers.google.com
  • Spam policies for Google web search
    Google Search Central (Google for Developers)
    The verbatim scaled-content-abuse definition; the purpose-and-value test that a word-level find-and-replace cannot pass on its own
    developers.google.com
Back to guide overview