How to choose an AI writing tool: five categories compared
Who this is for
For a solo marketer or small team about to buy an AI writing tool and unsure which of the five kinds they need.
What you will take away
- A 46-item banned-words list run against 67 of our published pieces returned four hits in total, so vocabulary is not where AI writing gives itself away.
- Across eighteen humanizer scans covering two source texts and three methods, the best result was 1.9 percent human and no output beat its own source.
- The five categories answer five different questions. Pick the one that matches your actual bottleneck, which is usually the rewrite rather than the first draft.
You're about to buy an AI writing tool, and the thing you're worried about is almost certainly not the thing the tool solves.
Five kinds of tool sit in this market. They look similar from the outside, they use a lot of the same words on their websites, and they solve five genuinely different problems. Most people buying one don't know which of the five they've picked until the invoice clears and the drafts start coming back wrong. So any honest AI writing tool comparison has to start with the categories, not the features.
The five categories, sorted by what they do
The word filter flags or bans vocabulary that "sounds like AI." Em dashes, "delve," "tapestry," "in today's landscape." It runs a list against your draft and tells you what to cut.
The rewriter takes a draft you already have and paraphrases it, usually with the promise that the output will read as human. This is the category most people mean when they say AI humanizer.
The generator takes a brief and produces a first draft from nothing. This is the biggest category by far and the one most buyers actually end up with.
The SEO scorer grades a draft on keyword coverage, heading structure and word count, measured against whatever's currently ranking for your term. It tells you what's missing relative to page one.
The voice system holds a profile of how you write, then scores every draft against that profile. Not against a generic standard of "good writing." Against yours.
Those five categories aren't competing products so much as answers to five separate questions, which is why comparing them feature by feature gets you nowhere. Pricing is genuinely all over the place across every category, and we haven't measured it, so this piece won't pretend to tell you what anything costs. What it can tell you is which question each category answers, and which of them answer a question nobody's actually asking.
A comparison of the five categories at a glance
| Word filter | Rewriter | Generator | SEO scorer | Voice systemThe Words Studio | |
|---|---|---|---|---|---|
| What it exists to do | Remove flagged vocabulary | Paraphrase a draft you have | Produce a first draft | Grade keyword and topic coverage | Hold your voice profile and score drafts against it |
| What it measures | Whether banned words are present | Nothing | Nothing | Overlap with the pages now ranking | Voice, brief, funnel stage, originality, reading level |
| Knows your voice | No | No | Loosely, via your prompt | No | Yes, from a stored profile |
| Catches structural repetition | No | No | No | No | Yes |
| Benchmarks against what's already winning in search | No | No | No | Yes, against page one | No |
| The signal you get back | A count of hits | None. Output, not information | None | Gaps against competitors | Per-draft score, named misses |
| It fails when | Your problem isn't vocabulary | It has to sound like one person | Volume isn't your bottleneck | It ranks and still reads assembled | You have no voice to profile yet |
One row goes to the SEO scorer outright, and it is the row that matters most if organic search is your main channel: it is the only category here that measures your draft against a real external benchmark rather than against itself.
Why the word filter doesn't work as advertised
The word filter fails because vocabulary is not where AI writing gives itself away, and we have the numbers on this from our own work.
We ran a 46-item banned-words list against 67 of our own published pieces and got four hits in total across all 67. If banned vocabulary really were the tell, a list that long should have lit up like a switchboard rather than turning up almost nothing.
Then we tested the most famous flag of all, and the result was worse for the category. We made a change that drove em dashes from 7.8 per piece down to 0.00, and it held at zero twenty times out of twenty, which is perfectly clean by the standard every word filter enforces. The detector score didn't move at all, not by a little and not in one direction over another.
Which means the whole category is organised around the wrong variable. You can strip every flagged word out of a draft and the draft still reads exactly as machine-made as it did before, because the thing giving it away was never sitting in the word choice.
Structure is where a draft gives itself away, and the counts make the gap hard to argue with. We tracked seven structural habits across 79 pieces and 64,995 words, and the em dash everyone bans turned up in only 4 percent of them. The most common habit in the set, the one almost nobody names or tests for, appeared in 44 percent, which puts an eleven-to-one gap between where the market is looking and where the problem actually lives. The full count is in why every AI draft sounds the same, and the banned-words test is in why a banned-words list won't fix AI-sounding drafts.
Best for: someone who wants a fast, cheap sanity check on obvious tics before hitting publish.
Key strength: it's instant and it requires no setup at all, since you run the text through and get a list back.
Key weakness: it measures the least predictive thing available. A draft can score perfectly and still be obviously synthetic to a reader.
Who should skip it: anyone treating a clean word-filter pass as evidence the writing is fine. It isn't evidence of anything except that you know which words are on the list.
Why the rewriter doesn't work either
The rewriter fails at the specific job it's sold for, which is making text read as human to a machine, and we tested that directly.
We ran eighteen scans of humanizer output, covering two source texts and three different methods. The best result across all eighteen runs was 1.9 percent human, and not one output scored higher than the text it started from, which means every pass left things the same or made them worse.
The baseline is what makes that finding damning rather than merely disappointing. Across 13 pieces from nine different organisations, the mean came in at 0.4 percent human, so the floor really is the floor. Everything sits down there whether it's been rewritten or not, and paraphrasing doesn't lift you off it.
To be clear about what that finding is and isn't: it says our rewriting didn't move detector scores. It doesn't say we beat detectors, and it doesn't say detectors are wrong. The honest, narrow claim is that pushing a draft through a paraphrase step changed the score by nothing worth having. The full run is written up in our humanizer never beat 1.9% on an AI detector in 18 tries.
Best for: someone who needs a different phrasing of a sentence they're stuck on, treated as a thesaurus for clauses rather than a compliance step.
Key strength: it genuinely does unstick you, because paraphrase is a real writing aid when you're staring at a line that won't come.
Key weakness: it strips out whatever made the original yours. Rewriters flatten toward an average, and your voice is by definition not the average.
Who should skip it: anyone buying an AI humanizer as insurance, because eighteen scans say the insurance doesn't pay out.
What the generator is actually good at, and where it stops
The generator is good at volume, and volume is a real problem for some teams, just not the problem most buyers have.
If you need forty product descriptions by Friday and nobody's going to read them closely, a generator earns its keep in an afternoon. That's a legitimate use case and there's no shame in it.
The trouble starts when the bottleneck was never blank-page speed in the first place. If you're one person writing for your own business, the draft isn't the hard part, because the rewrite is where the hours go. You get a draft back in seconds, then spend far longer making it sound like a human who works at your company, which nets you very little and burns the energy you needed for the thinking.
Best for: teams with a genuine volume requirement and a low bar for distinctiveness.
Key strength: it removes the blank cursor, which is a real psychological cost.
Key weakness: a prompt is a very lossy way to describe a voice. You can tell a generator to be "direct but warm" and it will produce something that a hundred other companies also asked for.
Who should skip it: anyone whose content has to sound like a specific person. The generator has no way of knowing whether it succeeded, because it has nothing to compare the draft to.
What the SEO scorer is actually good at
The SEO scorer is good at telling you what's missing relative to whatever's ranking, and it's the only category on this list that measures against a real external benchmark.
If you're publishing into competitive search and you don't know why page two won't budge, a scorer will show you the topics and terms the winners cover and you don't. That's useful information you can't easily get any other way.
Its limit is that it grades coverage, not character. A draft can hit every subtopic, nail the heading structure, land the word count, and still read like it was assembled rather than written. The scorer will give that draft a green light because the scorer isn't looking at the thing you're worried about. It's a content quality checker for topical completeness, and topical completeness is only one kind of quality.
Best for: teams where organic search is the primary channel and there's already a human editor handling voice.
Key strength: it gives you an external benchmark, so it isn't guessing what good looks like, it's reading what's already working on page one.
Key weakness: it can push you toward sameness. Optimising hard against the current page one tends to produce a draft shaped like the current page one.
Who should skip it: anyone whose distribution is email, social or sales conversations. You'd be paying to solve a problem you don't have.
The question worth asking instead of the detector one
Stop asking whether you can pass a detector. Start asking whether the draft reads like you, and whether you can prove it.
That's the pivot, and everything above points at it. The word filter can't answer it, because it only knows a word list. The rewriter can't, because it has no idea who you are. The generator can't, because it has no reference point. The scorer can't, because it's measuring against your competitors rather than against you.
None of those four tools holds a record of how you write, which is precisely why none of them can tell you whether a given draft matches it. They can tell you the draft is clean, or paraphrased, or finished, or topically complete, and every one of those verdicts sidesteps the question. What they can't tell you is whether the writing is yours.
What a voice system does differently
A voice system holds a profile of how you write and scores every draft against it, which makes it the only one of the five categories that can answer the question you actually care about. The Words Studio is a voice system, so everything below is a description of what we built and why.
Four checks run on every draft, and each one closes a gap the other categories leave open. First, voice match: the draft is scored against your stored profile, so you get a number and a list of named misses rather than a vague sense that something's off. Second, brief match: the draft is checked against the pillar and funnel stage you picked, because a consideration-stage piece that drifts into a sales pitch is a miss even if the sentences are lovely. Third, originality: the draft is checked against the live web to see whether anything in it already exists out there. Fourth, reading level against tone: if you asked for conversational and the prose comes back at a reading level nobody conversational has ever used, that gets flagged.
Brand voice software as a category tends to be described in fuzzy terms, so it's worth being concrete about what the profile actually holds. It's a written record of your habits: the traits and personality the writing carries, the words you reach for and the ones you refuse, how you open a piece, how you move between ideas, and how you close. The seven structural habits we counted across those 79 pieces and 64,995 words are the layer underneath that, and a word list cannot reach them.
Best for: one person or a small team where everything published has to sound like the same identifiable human, and where a full rewrite of every draft isn't affordable.
Key strength: it measures the structural layer, which is where our own counts say the tell actually lives. And it gives you a signal per draft rather than a vibe.
Key weakness: it needs a voice to profile. If you've published six things and they don't sound alike yet, there's nothing stable to measure against, and you'd be better off writing another ten pieces first.
Who should skip it: teams who genuinely don't care whether the output sounds distinctive. If your content is functional and nobody reads it for the writing, a generator covers you.
How to choose an AI writing tool, by use case
Pick the category that matches your actual bottleneck, not the one with the most convincing homepage. Choosing one without three months of trials comes down to a single question per category.
If you need a lot of low-stakes copy fast and distinctiveness genuinely doesn't matter, buy a generator and stop reading comparison articles.
If organic search is your main channel and a human editor already handles voice, buy an SEO scorer. Your gap is coverage, and that's what it measures.
If you're one person writing for your own business, and everything you publish has to sound like you because you're the brand, you need a voice system. An AI content tool for small business that doesn't hold your voice is just a faster way to produce content you'll have to rewrite. That is what we built The Words Studio to be.
If you're buying to pass a detector, don't buy anything in this list. Our own numbers say vocabulary changes don't move the score and paraphrasing doesn't either. You'd be spending money on a result nobody's demonstrated.
And if your drafts are technically fine but you can't say what's wrong with them, that's a voice problem wearing a quality costume. A content quality checker that only grades topical coverage will keep telling you the draft is good while you keep feeling it isn't.
See it on your own writing
Point The Words Studio at your website and it builds a voice profile from how you already write. Then every draft comes back scored against that profile, checked against the live web, and levelled to the tone you picked.
Common questions buyers ask before choosing
What's the difference between an AI humanizer and brand voice software?
An AI humanizer paraphrases text with the goal of changing how a machine classifies it, while brand voice software holds a stored profile of how you write and scores drafts against that profile. One transforms text blindly, with no reference point to work from. The other measures the draft against a known reference, which is the only way to tell whether it sounds like you.
Do banned-word lists actually make writing sound more human?
No, and that's based on our own measurement rather than a hunch. A 46-item banned-words list run against 67 of our published pieces returned four hits in total, and driving em dashes from 7.8 per piece down to 0.00, twenty runs out of twenty, moved the detector score not at all. Vocabulary simply isn't where machine writing reveals itself.
Do AI humanizers work the way they're sold?
Not for the job they're sold for, going by eighteen scans of humanizer output across two source texts and three methods. The best result was 1.9 percent human, and no output ever scored higher than the text it started from. Against a baseline of 0.4 percent human across 13 pieces from nine organisations, paraphrasing changed nothing worth paying for.
Which of these tools makes sense for a small business?
If you're the brand and everything you publish has to sound like you, choose a voice system that holds a profile of your writing. If you need volume and nobody reads your copy for the writing, a generator is enough. Match the tool to your real bottleneck, which is usually the rewrite, not the first draft.
What should I look for in a content quality checker?
Ask what the thing measures against, because that single answer tells you what it can and can't see. A checker grading keyword coverage compares you to ranking competitors, while a checker grading voice compares you to yourself. The second type can also check brief match, funnel stage, originality against the live web, and whether the reading level fits the tone you asked for.