Our humanizer never beat 1.9% on an AI detector in 18 tries

Who this is for

For solo marketers and owner-operators deciding whether an AI humanizer is worth paying for.

What you will take away

  • Across 18 scans on Originality.ai, the best score our humanizer ever produced was 1.9% human, and no output ever scored higher than the text it started from.
  • Cutting em dashes from 7.8 per piece to 0.00 across twenty runs moved the detector score not one point, so word-level edits do not reach what a detector reads.
  • We stopped scoring against detectors in August 2026 and now measure web duplication and reading level against tone, then rewrite the sentences that carry the repeated shapes.

Eighteen scans through a single AI detector, and the best score we ever recorded was 1.9% human.

We spent six weeks trying to beat an AI detector with our own humanizer, and we failed by a margin that wasn't close to narrow. We never once moved a piece past 1.9% on Originality.ai, no matter which prompt or text we ran. So if you bought an AI writing tool partly to humanize AI text and slip past detection, this is the measurement we wish someone had put in front of us before we started.

What the experiment actually set out to measure

We took two source texts and humanized each one three times over, using three different methods so that no single variable could explain a bad result. The first method chained passes through the production app, with each rewrite feeding into the next. The second ran independent passes on our current prompt, and the third ran independent passes on a previous prompt we'd already shipped to customers. That comes to 18 scans of humanizer output in total, all of them run through Originality.ai on aiModelVersion 4.

The setup rules out every obvious excuse in advance. If the problem were one bad prompt, the second prompt would have shown it. If the problem were one unlucky text, the repeats would have shown it. And if the problem were a single fluke run, chaining should have compounded any gain across passes, which it plainly didn't.

The maximum score across all 18 outputs was 1.9% human, and no output ever scored higher than the text it started from. On one piece, three chained passes ran 0.0, 0.0, 0.0, 0.0 from source through to final rewrite, which is four zeros in a row from a process meant to improve the number each time. We also re-scanned the source texts afterwards and they returned identical values, so the flat results weren't noise in the instrument.

Why word-level edits never reach an AI content detector

Word choice isn't what the detector measures, and we have a clean test for that. A separate change drove em dashes from 7.8 per piece down to 0.00, twenty times out of twenty. Perfect suppression of the single most cited AI tell in circulation, and the detector didn't move at all.

Which tells you something useful about AI content detector accuracy: the thing flagging your draft isn't counting dashes or swapping "utilize" for "use." It's reading something structural that survives a lexical rewrite intact. Every listicle about the words that supposedly give AI away is solving a problem the detector doesn't have.

The baseline is worse than the marketing suggests. Across 13 pieces from nine organisations between April and August 2026, unhumanized output averaged 0.4% human, with a median of 0.1% and a maximum of 2.3%. Tools promising to bypass AI detection imply you're sitting at 90% and need a nudge to 99%. The real distance you're being sold is 0 to 99, and we couldn't close it by a single point with full access to our own model, our own prompts, and six weeks behind it. Ask for the raw scans before you believe the pitch.

What we measure now instead of a detection score

We stopped scoring against detectors in August 2026. The product now measures what a reader actually experiences, because those are the only things that were ever under our control.

First, whether anything in the draft already exists on the web. You get the sentence, flagged, with the page it matches. If your draft says "a single source of truth for customer data" and forty published articles say it word for word, that's not plagiarism, but it is a line your reader has already skimmed past somewhere else. You rewrite it or you keep it knowing what it is.

Second, whether the reading level fits the tone you asked for. You told us the piece should sound like one engineer explaining something to another. The draft comes back at grade 16, with clauses stacked three deep and a 41-word sentence in the second paragraph. That mismatch is visible before you publish, not after someone tells you the blog reads like a white paper.

Third, the Humanizer goes after the repeated sentence shapes directly. It finds the specific lines carrying them and rewrites those, and you see every change side by side with the original to accept or reject one at a time. Those repeated shapes are what makes writing feel machine-made, and they are exactly what a lexical rewrite leaves intact. We counted seven of them across our own corpus in why every AI draft sounds the same.

All three of those measurements point at something you can actually change, which is the whole reason we bothered to build them. A detection score gives you nothing to change, as eighteen scans taught us at considerable expense.

The feature is still called the Humanizer, because we named it back when we thought the job was beating detectors and the name outlived the theory that produced it. It does a different job now, rewriting the specific sentences holding a piece back rather than chasing a number that never moved once in eighteen attempts. Better job, and a worse name than it deserves.

If your buying decision rests on a detection score, buy nothing and spend the money on an editor. If it rests on whether the draft reads like a person wrote it and says something true, that's a problem we can work on.

Common questions we get about detection

Can any tool reliably bypass AI detection in practice?

We found no evidence of it across 18 scans on Originality.ai, where our best humanized output reached 1.9% human and no output ever beat its own source text. Three chained rewrites on a single piece scored 0.0 every time, which is the outcome you'd expect if word-level rewriting simply doesn't touch what the detector reads. If a vendor claims otherwise, ask to see the raw scans and the detector version they ran.

Does removing em dashes help humanize AI text?

No, and we tested it properly before saying so. We cut em dashes from 7.8 per piece to 0.00 across twenty consecutive runs, and the detector score didn't move a single point. Word-level fixes don't reach whatever the detector is actually reading in your draft. Remove them because they're overused, not because a score will improve.

Can an AI content detector judge human writing correctly?

Our test measured machine-written text, not human text, so we can't speak to false positives. What we can say: on AI output, Originality.ai was consistent. Source texts re-scanned to identical values, which means the flat results reflect the instrument agreeing with itself rather than random variance.

What should I be looking for if not a detection score?

Look at what a reader experiences directly: whether the draft duplicates text that already sits on the web, whether the reading level matches the tone you asked for, and which sentences the Humanizer will rewrite, since the repeated shapes are what make writing feel generated in the first place. All three you can actually fix, which is the whole reason we measure them. A detection score you can't fix, as eighteen scans taught us.

Is paying for a tool to make AI writing sound human actually worth it?

That turns on what the money is buying, because the two products sold under the same label do very different jobs. A tool that flags your weak sentences and then rewrites those specific lines earns its price, since you can see the before and after and judge it yourself. A tool selling you a detection score it can't show it shifted by even one point out of a hundred, the way six weeks and eighteen scans never shifted ours past 1.9%, doesn't. So ask for the measurements first, then read the drafts and decide from those.

Not sure which kind of tool you are looking at? Five kinds of AI writing tool, compared.