← Back to blog
NewsAbout 6 min read

Someone Catalogued 13,000 Ways AI Writing Gives Itself Away

Published Oct 4, 2026
Someone Catalogued 13,000 Ways AI Writing Gives Itself Away

A study from the writing-analysis company Graphite has done something useful for anyone who reads a lot of machine-generated prose: it stopped guessing and started counting.

The team took 10,000 articles published before ChatGPT existed, asked AI models to rewrite them, and then compared the two piles. What fell out was a list of roughly 13,000 phrases that show up at least twice as often in the AI versions as in the human originals. Not whole sentences. Phrases. Small habitual turns of language that a model reaches for when it is filling space.

The tells moved, they did not disappear

The obvious part of that finding is the catalogue itself. The more interesting part is what happened as it was assembled.

Graphite's chief AI officer, Greg Druck, said the trend line runs in two directions at once. Older, louder markers such as em-dashes have been trained down, because developers noticed them and went after them. But when one set of tics gets smoothed away, another set grows into the gap. The catalogue did not get shorter over time. It got less familiar.

That is the uncomfortable core of it. People keep treating AI detection as a solved problem in one direction or the other. Either the detectors work and can spot a model instantly, or they are broken and useless. The reality is more like erosion. The surface gets scrubbed, the underlying grain stays.

Two labs, two directions

The most quoted line from the study is a split between vendors. Druck observed that Anthropic's Claude models have been moving closer to standard human word distributions across generations. OpenAI's GPT models, by the same measure, have been drifting further away.

Take that with the caution it deserves. This is one company's metric, on its own comparison set, measuring its own definition of "human-like." Word frequency distributions are not truth. A model can sit right on top of the human average and still produce sentences no person would write, and a model can wander off the distribution while writing something genuinely good.

Still, the direction is worth noting, because it lines up with a difference you can feel. Some models read like a careful writer who has been told to avoid cliches. Others read like a very confident assistant who has absorbed a house style and will not let it go.

Why phrase-level tells matter more than sentence-level ones

Most people met AI detection through whole-sentence giveaways. The stage-setting opener. The tidy summary at the end. The three-part list that arrives whether or not the content has three parts.

Models got better at not doing those things, at least when asked nicely. What is harder to sand off is the connective tissue. The way a paragraph is stitched to the next one. The reflexive move toward balance, where every claim gets a matching counterweight, whether or not the evidence supports one. The preference for the abstract noun over the specific one.

These survive rewriting instructions because they are not rules the model consciously follows. They are the shape of the distribution it learned from. You can tell a model to stop using a word. It is much harder to tell it to stop thinking in a particular rhythm.

An antique mechanical typewriter holding a blank sheet of paper under a green desk lamp

The phrase list also makes a subtler point about averages. Tells are not spread evenly. A model might use one phrase far more than people do and another less, in a pattern that only shows up across thousands of comparisons. No human reader notices the aggregate. A reader notices a sentence that feels off, and rarely knows why. That gap between what statistics can see and what a person can articulate is exactly why detection is hard to do well by eye.

The detection market this feeds

There is an obvious commercial consequence to a list like this. Any company selling AI detection now has a fresh, concrete set of features to train on, and any company selling AI writing now has a fresh list of habits to avoid.

That is the arms race in miniature. Detection companies mine patterns, writing companies train them out, and the catalogue shifts. The interesting question is not whether the next version of a detector can catch the current version of a model. It is what happens to trust when both sides get good enough that neither can be sure.

For a reader, the practical effect is a slow loss of a shortcut that never worked reliably anyway. Style was always a weak signal. Plagiarists wrote in their own voice for centuries, and honest writers have always borrowed phrases. The new catalogue just makes the weak signal measurable.

What this means if you publish

For a writer, the study reads less like a threat and more like a mirror. If 13,000 phrases tilt toward machine prose, then a draft that leans on those phrases will read as machine prose, no matter how good the ideas are. The fix is not a detector. The fix is specifics.

Machine writing defaults to the general because the general is always safe. Human writing is full of the particular: the name of the street, the number that does not round, the detail that only someone who was there would know. Those are the parts that no amount of phrase-scrubbing can fake, because they were never in the training data in the first place.

For an editor, the practical lesson is the opposite of automation. The tells are now subtle enough that a glance will not catch them. You have to read for rhythm and for the absence of specifics, not for the tell-tale em-dash that everybody already learned to look for. The most reliable edit is not deleting the obvious tics. It is adding the kind of concrete detail a model would not have invented, which serves the reader and happens to defeat the classifier at the same time.

The arms race nobody wins cleanly

There is a version of this story where detection gets solved. A classifier trained on enough of these phrase patterns flags AI text with high accuracy, and everyone goes back to trusting what they read.

That version has a problem. The moment a detector publishes, it becomes training signal. The next model is trained to avoid whatever the detector keys on, and the catalogue shifts again. Graphite's 13,000 phrases are a snapshot, not a settlement.

What is left is a slow redefinition of what "reads like a person" means. Each generation of models raises the floor on generic competence, and each generation makes the remaining human markers a little more distinctive by comparison. The tells do not vanish. They migrate into smaller and smaller gestures, until the only reliable signal left is the one that was always reliable: whether the writing knows something the reader did not, and says it plainly.

A phrase list cannot measure that. Neither, so far, can a classifier. It is the part of writing that resists being catalogued, and it may be the part that ends up mattering most.

Related articles