Machines have been writing text for more than seventy years. In 1952, Christopher Strachey programmed the Manchester Mark 1 to write love letters. In the mid-1960s, Joseph Weizenbaum at MIT built ELIZA, a program that played a therapist by turning patients’ words back into questions. His own secretary asked him to leave the room so she could talk to it in private. Story generators followed in the 1970s. From 1990 until 2019, the Loebner Prize held a yearly contest for the chatbot judges found most human-like. The prize for a program judges couldn’t tell from a person was never won. Now it’s a business to catch them.
Machines have been fixing text almost as long. Ralph Gorin wrote the first spelling checker at Stanford in 1971. Grammatik brought grammar and style checking to personal computers in 1981, and Microsoft Word built in a grammar checker in 1992. Long before ChatGPT, these programs were suggesting changes to make prose better, not just fixing typos. Grammarly, founded in 2009, turned that into a company valued at $13 billion. Nobody asked what percentage of a column the spell checker, grammar checker or prose checker wrote. The Washington Post now requires guest writers to confirm their work was not “created or manipulated with artificial intelligence or editing software.” Taken literally, that bans spell check.
With chatbots like ChatGPT, the question of who wrote this has moved to the forefront.
Enter the AI detector.
What an AI Detector Really Does
An AI detector looks for patterns in the prose that suggest a computer generated it. Pangram, a leading detector, was trained on millions of examples of human and AI writing until it learned to tell them apart. “We’re reverse-engineering their stylistic fingerprint,” its CEO, Max Spero, told The New York Times. It doesn’t read for meaning. It doesn’t check facts, judge arguments or ask whether anything in the piece is new.
For example, when I first started writing my op-eds, the chatbots kept giving me the same cliché: something “in the clothing of” something else. It showed up constantly, and it was very annoying. I don’t see it at all now. I’m guessing Anthropic noticed it too and stopped Claude from doing that. That would be a simple example of how a detector might know. A human writer would not use it so much, and certainly not constantly in every piece. The harm is just that it makes for boring and weak writing. A human reader will notice bad writing.
The things that irritate me most in AI-generated text, I also find in human writing: lots of padding for length and overexplaining. I constantly have to make the chatbots trim. They learned it from reading newspapers and books that sell by the pound, and apparently there is enough of that out there that they picked up the habit. So if you start looking for it, you will get the same “bad writing” from humans or AI.
What Problem Are We Trying to Solve Here?
One of my bosses always asked, whenever someone suggested some new software to write: What problem are we trying to solve here? Quite often the room got very quiet. Nobody had really thought about that.
In trying to detect what role AI had in a piece of writing, we should ask the same question.
Possible Problems This Can Solve
People who want AI detection have real reasons. Here are the main problems they hope it will solve:
Schools want students to do their own writing. Why do you care? They will have access to these tools in life. If they can write well with AI, they can write well.
Platforms want to catch meaningless spam: content farms, fake reviews, bot posts. A detector detects AI writing, not spam. It doesn’t look at the content. And humanizers can disguise the sentence patterns it does look at, as we explain later.
Readers want to know a real person is behind the byline. Why do you care whether a real person is behind the byline? A detector doesn’t tell you who you’re connecting with anyway. If you need to know, there’s DocuSign.
Publishers and prize committees want to reward human craft. This is not a chess match. Human craft with the help of AI is what will be out in the world to compete against.
Someone has to answer for what’s written, including a fabricated fact or citation. You are already responsible for what you publish.
Research needs integrity. Pangram estimated that 21% of peer reviews submitted to a major AI conference this year were fully AI-generated. A detector doesn’t tell you whether the review is correct, which is what matters. And AI is becoming, or may already be, smarter than humans. Why wouldn’t you want something smarter than a person reviewing things?
Writers want to protect their livelihoods from free machine text. This is protecting someone else’s business. On a related topic, a detector also spares teachers from grading student work that may be better than they could write themselves. Years of training in long division don’t protect you against a child with a calculator.
On the schools’ point: working out your own thinking with the help of AI is fine. Most people’s parents helped them long before AI. You still have to direct the AI; it doesn’t do it for you. People who haven’t written with it think it’s more automatic than it is. Programmers learn the same lesson: a chatbot can’t do a major software project for you if you don’t know how to do it yourself.
Carpenters use nail guns now. How fast you can put a roof on with a regular hammer is not something worth testing.
Some Purposes of Writing
Before deciding what a detector should protect, it helps to ask what writing is for. People write for many reasons, including these:
Create something the reader enjoys.
Teach the reader something.
Communicate information.
Persuade: change the reader’s mind or get them to act.
Record what happened, what was decided or what someone knew.
Work out your own thinking. Writing forces you to find out what you actually believe.
Connect with a particular person: letters, condolences, love notes.
Get something done: instructions, manuals, emails, proposals.
Regarding the Purpose of Writing, What Problems Does an AI Detector Really Solve?
Basically nothing.
You want the best possible writing, original ideas, things worth reading. You can ask a chatbot those questions without a special tool. I do this with all my own writing. I ask: “What is original about this piece?” or “Read, Analyze, Critique” or “Thoughts?” (works great).
I also can’t grow my own food. If the grocery stores closed, I’d be in trouble. Nobody thinks schools should ban supermarkets.
In my opinion, most of the proposed AI problems are just typical overreach by people with an agenda orthogonal to the purpose of the thing.
What Problems Does an AI Detector Create?
It creates an artificial grade, assigned by a robot, on the value of the writing and whether you cheated somehow.
It teaches that a poorer piece that shows no AI is considered better.
The Arms Race
Now there is an arms race between AI detectors and AI humanizers. An AI humanizer is designed to make an AI-assisted piece undetectable.
The humanizers are doing well. In a University of Notre Dame study, not yet peer reviewed, once AI-written text went through a humanizer called Undetectable AI, more than 96% of it passed as human. Undetectable AI offers a money-back guarantee if its output gets flagged, and it claims more than 22 million users.
The same study found the opposite problem for honest writers. Research abstracts that had only been lightly edited with AI were flagged 64 to 80% of the time. Pangram says its newest version fixed that; nobody outside the company has checked.
Honest AI editing carries more risk than paid evasion.
Pangram does better against amateur disguises. A peer-reviewed study from Vrije Universiteit Brussel found it caught 37 of 40 AI-written papers that GPT-4o had simply been told to make sound human. Against a real humanizer, the Notre Dame numbers are the ones that count.
Grammarly sells both sides. It offers an AI detector, and a feature that “humanizes” text to get past detectors.
For those playing at home, computer science already has a name for this: a generative adversarial network, or GAN. The generator is the one that’s supposed to win. This is how you make counterfeit money with machine learning.
Whose Help?
A ghostwritten book passes. Prince Harry’s memoir Spare was written with J.R. Moehringer. A human wrote those sentences, so a detector scores it human.
Presidents have speechwriters. Celebrities have ghostwriters. Academics have editors nobody credits. Under the Post’s rule, a human editor rewriting your paragraphs is fine. Software doing the same thing is not. Even Pangram’s CEO told the Times that using AI as an assistant is “completely OK.”
Slide Rules
We used slide rules. When calculators arrived, the objection was that we’d never learn the real skill. By 1980 the National Council of Teachers of Mathematics was calling for calculators at every grade level.
In my early twenties, preparing text for publication was a big job. You couldn’t just bang out an article. Then word processors arrived, and I watched the number and size of journals explode. Research later found word processing made student writing both longer and better.
Nobody built a detector for text typed on a computer instead of a Selectric. Journals that got too long set length limits. The fix was norms about the output, not forensics on the process.
Summary
At best, a detector score tells you one thing: AI touched the prose. It can’t tell you whose ideas they are, whether the facts are right, or whether the piece is worth reading. Honest writers get flagged, anyone who pays for a humanizer walks through, and a ghostwritten book passes.
When it comes to AI detectors, readers don’t hear “stylistic patterns.” They hear cheating or spam. Hachette canceled a novel after Pangram’s CEO posted that it scored 78% AI. Dartmouth’s provost, flagged when Semafor scanned more than 300 newspaper guest columns, now faces a campus review. To me this is just a witch hunt.
My old boss would ask: what problem are we trying to solve? For a reader, the only problem is whether a piece is worth their time. A chatbot can help answer that. A detector is not designed to.
Pangram will probably score this piece near 100% AI-written.
It can’t tell you whether it is worth reading.
It also can’t tell you whether Pangram itself has any useful place in the world of writing.

