A banned-words list won't fix AI-sounding drafts

Who this is for

For the one person handling all the content on a small B2B marketing team who keeps rewriting AI drafts.

What you will take away

  • Banned-words lists work at the word level, while the AI writing tells readers actually notice work at the sentence and paragraph level, so the two never meet.
  • The corrective fragment, a one or two-word sentence that snaps back at the previous clause, showed up in 44% of the 79 pieces we measured and contains no banned words.
  • AI content detection tools score probability without naming which sentence gave you away, so you get a verdict and no evidence to act on.

We ran a 46-item banned-words list against 67 of our own pieces and found four hits across the entire set, which on paper looks like a list quietly doing the job it was hired for. Then we ran a control with the rules switched off, and that batch came back with zero banned terms in it as well, meaning the model wasn't reaching for those words in the first place and the list was scoring a result it had no hand in producing. However good the numbers looked, it was standing guard over a door nobody was trying to walk through.

Why banned-words lists miss the real AI writing tells

Banned-words lists work at the word level, while the AI writing tells your readers actually notice work at the sentence and paragraph level, which is why the two never meet. You can strip out "delve," "leverage," and "seamless" and still be left with prose that has a recognizable gait, because the vocabulary was never the thing carrying the machine feeling. Readers don't consciously flag vocabulary as they move down the page; what they register instead is rhythm, and rhythm survives every word you ban.

We measured our own output to check this, across 79 pieces and 64,995 words, and the single most common pattern was the corrective fragment, the one-word or two-word sentence that snaps back at the previous clause, as in "Not leads." It showed up in 44% of the pieces we measured, which makes it the closest thing we have to a signature, and not one of those fragments contains a banned word even though every one of them is a tell.

Three more patterns landed at 23% each in that same set, and they're all structural too. There's the not-just-X-but-Y antithesis, where every claim gets a mirrored upgrade; the announced triad, where a sentence promises three things and then dutifully delivers exactly three; and the withheld payoff, where a paragraph builds toward a point and then parks it on its own line for effect.

If you're the only person at your company who writes anything, you already know what this feels like from the inside, because the draft passes every rule you set and still reads like it came off a machine. You can't say why, so you rewrite the whole thing and lose the afternoon the tool was supposed to save.

What AI content detection actually catches, and what it leaves for you

AI content detection tools score probability rather than structure, so they'll tell you a piece looks machine-written without telling you which sentence gave it away, which is a scan result rather than a fix. You still have to find the line and change it yourself, with nothing from the tool to work from.

A content quality check that only reports a percentage puts you in the same position as the banned-words list, where something got flagged or nothing did, and either way your next move is guesswork. The four hits we found across 67 pieces told us a rule fired, but they told us nothing about whether the writing sounded human.

Naming the pattern is the part that changes a draft, because "this paragraph ends on a withheld payoff, and so did the last two" is a note a writer can act on in thirty seconds, while "82% likely AI-generated content" is a note that sends you back to the top of the page.

What we check instead of a banned-words list

We check three things: originality against the live web, reading level against a tone band, and named structural patterns.

Originality means a plagiarism check that runs against what's actually published right now rather than a static index, because if your draft has quietly reproduced a sentence from a competitor's post, that's a problem no vocabulary rule catches and no AI content detection score surfaces. Reading level means a readability score measured against a band you set rather than a generic grade target, since technical B2B writing that scores like a press release has a problem and so does a founder's blog that reads like a compliance memo. The number only means something next to the tone you're aiming for.

Structural patterns mean we count the four we named above, per piece, and report each one back to you by name rather than as a total, so corrective fragments, antitheses, announced triads and withheld payoffs each get their own line in the report. We track frequency rather than presence, because one corrective fragment is a stylistic choice while eleven in the same piece is a fingerprint.

Those three checks together make up the pre-publish checklist we run before anything goes out, and each one points at a specific line in the draft instead of producing a score you have to sit and interpret. If you tried an AI writer and everything came out generic, this is roughly where it went wrong: the tool controlled the words and left the shape of the sentences alone, so fixing the shape is what takes the generic feeling with it.

Common questions

Do AI content detection tools actually work in practice?

They return a probability score that doesn't identify which sentences triggered it, so you end up with a verdict and no evidence behind it. For a pre-publish checklist, a named structural pattern is more useful than a percentage, because you can act on it immediately instead of rewriting the whole draft.

Which AI writing tells turn up most often?

In our own measurement of 79 pieces and 64,995 words, the corrective fragment appeared in 44% of everything we looked at, with three more patterns hitting 23% each: the not-just-X-but-Y antithesis, the announced triad, and the withheld payoff. All four are structural rather than lexical, and none of them involve a word that any banned-words list would ever flag.

Does a banned-words list improve AI-generated content in any measurable way?

Our 46-item list produced four hits across 67 pieces, which is a thin return for a rule that size, and a rules-off control found zero banned terms in the same conditions, so the list may be policing a base rate the model never reaches anyway. It didn't hurt anything, and it also didn't measurably change how the AI-generated content read.

What should a content quality check include before you publish?

Three things, in our setup: a plagiarism check against the live web rather than a static index, a readability score measured against a tone band you define, and named structural patterns counted per piece, each of which points at a specific line to change rather than returning a score you have to interpret.

Why does my readability score look fine when the draft still reads wrong?

Sentence length and word complexity are the two things a readability score is built to measure, and repetition of shape isn't one of them, so a draft can score cleanly and still stack eleven corrective fragments in a row, which is exactly what readers notice when they can't say why the piece feels off. Structure and readability are separate checks, and you need both of them running.


All articles