AI for Writing, and Where It Quietly Fails
Using AI for writing works better in a specific order. Here's the order, the failure modes I actually hit, and the per-draft cost with real numbers.

Every page ranking for AI for writing right now answers the question "which tool". None of them answer "in what order", which is the question that actually determines whether the output is any good.
I checked. The July 2026 top ten was two roundups, four vendor tool pages, an academic writing product, a resource hub, and one Reddit thread. Roughly 18,100 searches a month for the phrase, difficulty around 28. Not one of those results describes a working sequence.
Here's the short version. Using AI for writing works when research comes first and drafting comes second, and it produces confident nonsense when you reverse that. The model is excellent at shaping material you hand it and unreliable at inventing material it doesn't have. Everything below is the same idea applied at different stages, plus the specific places I've watched it go wrong.
The Order That Works
Research, then structure, then draft, then cut. In that sequence, and the sequence is doing more work than the tool choice.
The reason is boring. A model generating from nothing generates from patterns, and patterns are averages, so you get the average article about your topic. Every claim in it will be plausible and roughly half of the specific ones will be wrong in a way that takes longer to check than it would have taken to look up. Hand the same model six sources you loaded yourself, plus the numbers you want it to use, and the failure mode disappears almost entirely, because now it's arranging rather than inventing.
I don't have a controlled experiment to point at here. What I have is the pattern across a few hundred drafts, which is that every genuinely bad piece I've produced started with drafting.
The Failure Mode Nobody Mentions
This one took me a while to spot and it's the most useful thing on this page.
I run keyword research through a paid data API, and one of the standard moves is to ask for related keywords around a seed term so you can map a topic properly. Sounds harmless. What actually happens is that the expansion drifts, steadily and quietly, toward whatever adjacent topic has the most search volume. Ask for keywords around a narrow subject and by the third hop you're looking at a list that belongs to a completely different, bigger niche, because volume is what the expansion optimises for and it doesn't know you've changed subject.
The same thing happens inside a draft. Ask a model to expand a section and it drifts toward the most heavily represented version of that topic in its training, which is usually the most generic version. You asked about a specific workflow, you got the general one, and it reads fine so you don't notice.
The fix in both cases is the same. Constrain hard at every step, check the output against what you actually asked for rather than against whether it reads well, and be suspicious specifically when the result is smooth. Smoothness is the tell.
What It's Reliable For, Stage By Stage
I keep this rough table in my head and it's saved me a lot of wasted rewriting.
| Stage | Reliable? | Where it breaks |
|---|---|---|
| Outlining from sources you supply | Yes | Will over-structure, wants five headings where two work |
| Drafting from a supplied brief plus facts | Mostly | Flattens specifics into generalities if the brief is thin |
| Drafting from a topic alone | No | Invents statistics, invents case studies, sounds certain |
| Tightening a paragraph you wrote | Yes | Sometimes removes the interesting part |
| Fact checking its own output | No | Confidently confirms its own errors |
| Rewriting for a different audience | Mostly | Changes register, keeps the original's structural habits |
| Coming up with the angle | No | The angle is the whole job and it defaults to the obvious one |
The bottom row is the one I'd underline. The angle is the thing that makes a piece worth reading and it's the thing a model is worst at, because "what has nobody said about this" is a question about absence, and absence isn't in the training data.
The Cost, Since Nobody Prints It
A quick set of real numbers, because the roundups talk about subscriptions and never about what a draft costs.
Anthropic's pricing page on 28 July 2026 lists Claude Sonnet 5 at an introductory $2 per million input tokens and $10 per million output, running through 31 August 2026, and Haiku 4.5 at $1 and $5. The same page puts one token at roughly 0.75 words in English.
Run the arithmetic on a 2,000-word draft, which is about 2,700 output tokens. If you've done the research first and you're feeding in something like 40,000 tokens of source material, Sonnet lands around eleven cents and Haiku around five. If you skip the research and prompt from a topic, the input cost basically vanishes and you're paying about three cents, but you're also getting the version with invented statistics in it. That's a bad trade at any price.
Which is a slightly funny result. The good workflow costs more per draft than the bad one, and the difference is eight cents.
Add a detector scan if you want one. Originality.ai's Pro plan, per their pricing page the same day, is $14.95 a month for 2,000 credits, with one credit covering 100 words of AI detection. Twenty credits for a 2,000-word piece, so about fifteen cents a pass. Not a meaningful line item unless you're rescanning constantly, which people do.
Does Any Of This Get Penalised
Worth answering directly since it comes up in every conversation and mostly gets answered with folklore.
Google's spam policies define scaled content abuse as "when many pages are generated for the primary purpose of manipulating search rankings and not helping users", and list "Using generative AI tools or other similar tools to generate many pages without adding value for users" as an example. The helpful content guidance adds that "If you use automation, including AI-generation, to produce content for the primary purpose of manipulating search rankings, that's a violation of our spam policies."
I loaded both of those on 28 July 2026. Note what they don't say. There's no clause about a model touching your prose. The offence described is volume plus thinness plus intent, which is a different thing entirely from writing one well-researched article with help.
Whether that distinction survives contact with an actual algorithm is a separate question and I genuinely don't know. The documented position is clear enough though.
Things I Do Differently Now
A short list, all of them learned by getting it wrong first.
I write the angle myself, before anything else, in one sentence. If I can't state what this piece says that the existing ones don't, I don't start. This kills maybe a third of my ideas at zero cost, which is the point.
I gather sources before drafting and I record what each one says with the URL and the date I loaded it. Not because anyone audits me, but because the alternative is discovering mid-edit that I can't remember where a number came from, and then either cutting it or spending twenty minutes re-finding it. If a source won't load, the number doesn't go in. I've left visible holes in published pieces over this and it's never once cost me anything.
I draft in sections rather than all at once. Long single-pass output degrades toward the middle in a way that's hard to see when you're reading it fresh.
I read the finished thing out loud, or at least mutter it. That catches the flat, evenly-paced rhythm that's the giveaway of unedited machine prose faster than any tool does. It's free and it takes four minutes.
What The Research Step Actually Looks Like
I've said "research first" three times now without describing it, which is the sort of thing that makes advice useless. So, concretely.
Before I write a word, I build a list. Each row has four fields, which are the claim, the value, the URL, and the date I loaded it. That's the whole system. It lives in a scratch file and gets thrown away afterwards.
The rule attached to it is the part that does the work. If a number isn't in that list, it doesn't go in the article. Not "I'm fairly sure it's around forty percent", not "I read somewhere that". If the page won't load, the claim gets cut or hedged into something I can actually stand behind.
That sounds pedantic and it takes maybe twenty minutes. What it buys is that the draft stage becomes arranging rather than inventing, which is exactly the mode where a model is reliable. It also means the finished piece contains specifics, and specifics are most of what separates writing worth reading from writing that merely reads well.
I've published articles with visible holes in them because a source returned a 403 and I wouldn't quote a number I hadn't personally loaded. Nobody has ever complained. I suspect admitting a gap reads as more trustworthy than filling it smoothly, though I'd struggle to prove that.
Students, Essays, And The Awkward Part
Two of the pages ranking for this phrase are academic writing products, and one is a university-facing resource hub, which tells you a lot about who's actually searching. So it'd be evasive to skip this.
The distinction that matters is between using a model to understand something and using it to produce something you'll claim you wrote. The first is genuinely excellent. Ask for the argument of a paper explained three different ways, ask what the counterargument is, ask what a term means in this specific field. That's tutoring and it works.
The second is a different activity with a different risk profile, and the risk isn't only getting caught. It's that you end up with a document you can't defend in conversation, which is a bad position to be in during a viva or a job interview about work on your CV.
There's also a practical trap specific to academic writing, which is that models are unreliable about citations in a way that's uniquely damaging in that context. A plausible-looking reference to a paper that doesn't exist is the single most common failure I've seen, and it's an instant credibility loss. If you take one thing from this section, it's to check every citation by loading it, not by recognising the author's name.
Where This Fits With Everything Else
Drafting is one of four jobs people lump together, and the other three have their own tools and their own failure modes. The full breakdown is in the AI writing tools overview, which also has the subscription-versus-tokens math laid out properly.
If your actual goal is search traffic rather than word count, the leverage is upstream of writing. Picking the right target matters more than the prose, and that's the SEO side of the problem rather than the writing side. And if you're working with existing text rather than starting fresh, what a rewriting pass changes is worth understanding before you run one, because it changes less than people assume.
For anyone comparing specific products rather than process, the tools writers actually reach for is the comparison, and it's a different question from this one.
The Honest Limit
I'll end on the thing I can't fix with a workflow.
A model will happily produce four thousand competent words about anything, and competent words are worth close to nothing now that everyone can produce them. What's scarce is having done something and being willing to report what happened, including the parts that didn't work. That's not a prompt. There's no setting for it.
The best use of AI for writing, in my experience, is that it removes the friction between having something to say and having it written down. It does not help with the having something to say. If you don't, faster drafting just gets you to the empty result sooner.


