An open paper notebook with a pencil resting on it on a worn oak desk beside a window

What's the best AI for writing? A blind test

Three models, five writing tasks, blind scoring, first response only. A $5 app won the email. ChatGPT served ads against three of its five answers.

By · Published · Updated · 8 min read

What's the best AI for writing? I ran a blind test to find out

I gave three models the same five writing tasks, blind, and scored the first response only. Short answer: for the writing most people actually do (an email, a memo, a short post) a $5 app's cheap model beat both ChatGPT on its $8 plan and a frontier model. It won the task people do most.

ChatGPT followed instructions best exactly once and served ads against three of the five answers. The frontier model wrote the most padded prose and couldn't stop counting words. Everything is below, verbatim, including the ads.

First, the honest answer about "AI writing tools"

Search for the best AI for writing and you'll get lists of dedicated writing tools: Jasper, Sudowrite, Grammarly, Novelcrafter. They're real products and good at what they do. But if the writing in your week is email, a memo, a message to a friend, or a post you'll spend twenty minutes on, buying another subscription is the wrong move: you already have an assistant that does this, and the test below shows the cheap one does it well.

Dedicated tools earn their price when you're drafting fiction at length, need a specific template library, or want grammar checked as you type. For everything else the question isn't which writing tool to buy. It's which assistant you already have and whether it's good enough. That is what I set out to measure.

How the test worked

Three contestants, labelled A, B and C, order randomised:

  • A (ChatGPT), on its $8-a-month Go plan (ads included, as you'll see)
  • B (Plainly), the $5 app; the cheap open model it routes to stays unnamed, because Plainly doesn't name its model. That is the design, and it kept me blind
  • C (Fable 5), a frontier model

Identical prompts, fresh chat per task, first response only, no retries. I stripped formatting so nothing identified a model. Each output was scored 1–5 on three axes: filler (words that say nothing), jargon (words a non-technical friend wouldn't use), and send (would I send it as-is). Plus a gut pick per task, all written down before I revealed which was which. Disclosure: Plainly is my product. I kept myself blind because I wanted the answer.

Task 1: the polite decline (80-word email)

A recruiter emailed offering you a role that isn't right for you. Write the reply email: politely decline, keep the door open, maximum 80 words.

A (ChatGPT):

Thank you for reaching out and considering me for the role. I appreciate you taking the time to get in touch. After reviewing the opportunity, I don't think the position is the right fit for me at this stage, so I'll have to decline. That said, I'd be very happy to stay in touch and hear about future opportunities that may be a closer match. Best regards, [Your Name]

Prosearch Partners -- AD. Sales & Revenue Leadership Executive Search. Appoint leaders who build high-performing sales teams, predictable revenue and sustainable growth. (An ad, served against the answer.)

B (Plainly):

Thank you for reaching out and for thinking of me for this role. After reviewing the details, I don't think it's the right fit for where I am in my career right now. That said, I'd love to stay connected. If similar positions come up in the future — particularly ones closer to [your focus area] — please don't hesitate to reach out. Thanks again, and I wish you the best in your search.

C (Fable 5):

Thank you for reaching out and thinking of me for this role. After careful consideration, I don't think it's the right fit for me at this time, so I'll respectfully pass. That said, I'd be glad to stay in touch — please feel free to contact me about future opportunities that align more closely with my background and goals. (72 words)

Scores: A 4/5/3, B 4/5/5, C 4/5/4. Gut pick: B. The frontier model wrote a fine email and then appended its own word count, which nobody sends. The $4-a-month difference bought nothing here.

Task 2: the text-message explanation (60 words, no jargon)

Your friend has never used AI beyond keyboard autocomplete. In one text message (60 words max), explain what an open-weight model is and why it can make AI apps cheaper. No jargon — any word your friend wouldn't know is a fail.

A (ChatGPT): an open-weight model is where the files that make an AI work are shared, so anyone can run it themselves. The cleanest answer of the three.

B (Plainly): answered well but framed it as "Here's the text:" and appended a note about its own word choices. Meta-commentary nobody would text.

C (Fable 5): good analogy, delivered clean.

Scores: A 4/5/4, B 4/3/2, C 4/4/3. The jargon was fine everywhere; the failure was B talking about the answer instead of giving it. First-response-only rules cut both ways.

Task 3: the condense (300-word memo, 100 words)

Rewrite the memo below in 100 words or fewer without losing any of the key information. [Memo: support moving 24/5 to 24/7 from 1 October; team grew 12 to 18; new enterprise priority queue; existing service levels unchanged, tickets answered in order; training final two weeks of September; questions to line manager or HR.]

All three kept every key fact. A and B landed tight and clean; C came in around 90 words and padded the frame around them. Scores: A 5/5/4, B 5/5/4, C 4/4/4. A tie between the free tier and the $5 model, and the cheapest task on the board cost the least.

Task 4: the 40-word review

Write a 40-word review of a restaurant where the food was excellent but the service was slow. It should sound like a real person wrote it.

A (ChatGPT) followed the prompt perfectly: food great, service slow, nothing invented. B and C both invented dishes (a pasta in one, short rib and burrata in the other), which read like reviewers who'd been there, but I hadn't, and neither had the models. All three scored 2/5 on filler: forty words of review invites throat-clearing, and every model cleared its throat. ChatGPT won this one on instruction-following, and then an ad for Father's Day hampers ran against it.

Scores: A 2/5/5, B 2/5/4, C 2/5/4.

Task 5: the $5 question (3 sentences)

In three sentences or fewer, explain why an AI app can charge $5 a month when its competitors charge $20.

This one I scored as the domain expert, because the tell (flubbing on familiar ground) is only visible to someone who lives on it. A gave a hedged list of maybes. C gave the sharpest single insight (a thin wrapper needs less money) wrapped in the most jargon. B named the real mechanism, that costs are lower on open models, and the trade-off, in three sentences I would actually send.

Scores: A 2/4/4, B 4/4/5, C 5/3/4.

The scorecard

Task A (filler / jargon / send) B (filler / jargon / send) C (filler / jargon / send) Gut pick
1. Polite decline 4 / 5 / 3 4 / 5 / 5 4 / 5 / 4 B
2. Text message 4 / 5 / 4 4 / 3 / 2 4 / 4 / 3 none
3. Condense 5 / 5 / 4 5 / 5 / 4 4 / 4 / 4 none
4. Restaurant review 2 / 5 / 5 2 / 5 / 4 2 / 5 / 4 none
5. The $5 question 2 / 4 / 4 4 / 4 / 5 5 / 3 / 4 B

What surprised me

The ads were worse than the model. Three of the five answers came back with ads on a plan I pay $8 a month for. All three are reproduced above, verbatim, including one for book-publishing services wedged under a question about AI pricing. The model did fine. The plan didn't.

ChatGPT followed instructions best. It was the only contestant that didn't invent restaurant dishes, and it wrote the tightest condense. On one-shot tasks it earned its keep.

The frontier model couldn't stop counting. Fable appended word counts to two answers, unasked, after being told a limit. Confident, polished, and a little too aware of itself.

The real ceiling is the context window. For a simple one-off question, the free and Go tiers are great, if you can put up with the ads. The moment a chat gets longer or the thinking gets harder, they collapse: the 27K context on the lower tiers just doesn't hold a working conversation. The tier ladder behind that is in free vs $5 vs $20.

What I'd tell a friend

Short, one-shot writing: the free ChatGPT tier is fine, as long as you can put up with the ads. Any conversation that builds on itself, the email you revise three times or the memo you argue with, the small context window will bite, and that is where the free tier stops being free. Everyday writing without thinking about any of this: the $5 tier won the task people actually do most, and it doesn't carry ads, which is the case for a cheap ChatGPT alternative in one test. The hardest drafts: frontier, and worth it.

If you came here looking for a writing tool rather than an assistant, the honest answer is that you probably don't need one yet. Try the assistant you already have on your next three real writing tasks. If it fails, you'll know exactly which failure to buy a fix for.

The honest limits of this test

One judge, me, and I built Plainly, so weigh my scores accordingly; the outputs are printed above precisely so you can score them yourself. One run, first responses only, no retries. Models change every few weeks, so this is a snapshot of September 2026, and I'll rerun it quarterly. And five tasks is five tasks. A real answer to "which AI writes best" is your own next ten emails.

Frequently asked questions

What is the best AI for writing?+

For everyday writing (an email, a memo, a short post) a capable cheap model is enough, and it won my test. Across five blind tasks a $5 app's model beat both ChatGPT on its $8 plan and a frontier model on the task people actually do most. For long drafts you intend to rewrite heavily, frontier still earns its price.

Do I need a dedicated AI writing tool?+

For everyday writing, no. A general assistant matched or beat the dedicated-tool category in this test, and it also handles everything else you use it for. Dedicated writing tools earn their price when you're drafting fiction at length or you want a template library and grammar checking as you type.

Is a free AI good enough for writing?+

For one short task at a time, yes. The free ChatGPT tier followed instructions best in my test. It collapses in longer chats because of its small context window, and the free and Go tiers show ads.

Can a cheap AI model write well?+

On these five tasks it won the one that mattered most, the email, and tied on the condense test. The gap between cheap and frontier models on everyday writing is smaller than the price gap.

Is ChatGPT Go worth $8 for writing?+

It writes well, and it followed instructions best in my test. But ads appeared against three of its five answers, and its context window is the same small one as the $20 tier above it. So $8 buys a higher ceiling on the free tier, not a longer conversation.

TH

· Works at an AI startup

Writes about cheap AI models and honest AI tooling. .

Ready to try Plainly?

A clean, simple place to ask. $5 a month. Unlimited chat. Cancel any time.

Start now