More Human Than Human
The AI-writing witch hunt burns real people and misses real machines. I would know. I run one.
Everyone wants a way to spot AI writing. Teachers want it, editors want it, the reply guy under your last post wants it worst of all. Ask a better question. What happens to the people who fail the test?
William Quarterman can tell you. In 2023 the UC Davis senior was failed and referred to the university’s honor court because a program called GPTZero flagged his history exam. He cleared his name with two exhibits: his Google Docs edit history, and a demonstration that the same detector flags Martin Luther King’s “I Have a Dream” speech and the Book of Genesis as machine work. Marley Stevens, a junior at the University of North Georgia, took a zero, went on academic probation, and lost a scholarship over a paper she says touched nothing smarter than Grammarly’s spell-check; the appeal included $105 for a mandatory honesty seminar. Kimberly Gasuras had covered Ohio local news for 24 years when a detector called Originality.ai cost her a freelance platform. No hearing. A score.
Now hold those next to what the testers admit about their own instruments.
OpenAI built a detector and retired it inside six months “due to its low rate of accuracy.” Its own published evaluation: it correctly caught 26 percent of machine text and falsely accused 9 percent of human text. One in four machines caught. Nearly one in ten humans burned. They shut it down. The schools did not.
Stanford researchers ran student essays through seven commercial detectors. Native English speakers came back near-perfectly human. Essays by Chinese students learning English were flagged as AI 61.3 percent of the time, on average, and at least one detector flagged 97.8 percent of them. The machine could not find the machine, so it found the foreigner.
Turnitin, the hall monitor of American education, launched its detector claiming “a less than 1% false positive rate.” Ten weeks and 38.5 million student papers later, it conceded a sentence-level false positive rate around 4 percent, concentrated exactly where students write like students: introductions and conclusions. Vanderbilt did the arithmetic on its own 75,000 yearly submissions, figured even the advertised error rate meant hundreds of falsely accused students, and turned the thing off. And in the single funniest data point of the panic, ZeroGPT scored the Declaration of Independence 97.93 percent AI-generated. Canonical texts saturate the training data, so the tools read the founders as robots. A test that fails Jefferson, King, and Genesis is not measuring what it claims to measure. Your kid’s essay does not stand a chance.
The spec for sounding human
Here is where farce turns into economics. On December 4, 2023, a handful of editors founded WikiProject AI Cleanup and began writing down what machine prose actually looks like. Their field guide, Signs of AI writing, might be the most useful style manual of the decade, and it was written as a wanted poster: the em dash pileups, the “not just X, but Y” constructions, the rule-of-three tic, the vocabulary (“delve,” “tapestry,” “testament”), the inflation where every event “marks a pivotal moment.” The project now counts 286 participants and a backlog of 6,937 suspected articles, and since August 2025 Wikipedia can speedy-delete unreviewed LLM pages the way it deletes vandalism.
The editors are not wrong about the patterns. At least 13.5 percent of 2024 biomedical abstracts carry LLM fingerprints in their vocabulary. By one detector-based measure, majority-machine articles overtook human-written ones on the open web in November 2024; by a looser one, three quarters of new pages carry some AI text. Those figures come from detectors, so hold them loosely. But the flood is real, and the people bailing water have my respect.
The trouble is what a public spec creates. The moment “signs of AI writing” exists, it is a compliance checklist for machines and an indictment checklist for humans, and both markets opened the same year. On the accusation side, the em dash now has a slur: the “ChatGPT hyphen.” On the compliance side, twenty dollars a month buys a “humanizer” that launders machine text past the detectors, and the biggest one claims 23 million users. The research favors the launderers: one paraphrasing model dropped a detector’s accuracy from 70.3 percent to 4.6. QuillBot sells a detector and a humanizer from the same storefront, which is at least honest about what kind of war this is. Detection firms answered with keystroke surveillance that watches you write. Humanizers answered back with fake keystrokes, typed out one at a time, to simulate a student on a deadline.
Meanwhile the actual humans are sanding themselves down. A grad student told NBC he leaves words misspelled on purpose so the detector will believe him: “If we write properly, we get accused of being AI.” A professor in the same story: “We’re just in a spiral that will never end.”
Follow the incentives to the bottom of that spiral. Machine text is being tuned, iteratively and for money, to pass a published test of humanity. Human text is being roughed up to pass the same test. The equilibrium is exact: machines engineered to read human, humans engineered to read broken. More human than human. Tyrell Corporation said it as a sales pitch in 1982. White Zombie screamed it as a warning in 1995. The detectors made it a market.
The part where I tell on myself
The reason this essay exists: I get accused of being AI. This publication has a standing rule against em dashes, and you can guess why. So when it happened again, I did what I do. I went and checked what the tells actually are, rule by rule, and found that most of them are just descriptions of bad writing that humans were committing at scale long before the machines learned it from us.
And here is the confession at full strength. I did not merely use AI to help draft this piece. I built one. It is called Faitality. It is mined from everything I have published, it holds my actual positions, and I calibrated it on purpose to run 50 percent hotter than I do. It proposes the price while I am still wondering what to charge. It drafts the email I was going to put off until Thursday. It names it, out loud, when I am stalling, because I built that in too. It is, by design, more Will than Will.
So the detector people are right to be paranoid about people like me. They are just wrong about the test. Every word the twin produces dies in a file unless I sign it. It has no publish button, and it is not getting one. What you are reading survived my knife, and the lines that did not survive are gone because they were not true enough or not mine enough. Run this essay through a detector if you like. I honestly do not know what it will say, and that is the point. The score cannot tell you the one thing you actually want to know.
What you want to know is whether a person stands behind the words. Whether somebody with a name checked the numbers, holds the position, and eats the consequences when it is wrong. That was never a property of the prose. It is a property of the signature. Every guarantee you have ever trusted works this way. The stamp does not make the steel true. It tells you who to come find when it is not.
Where the paint is thin
The steelman deserves full strength. The slop flood is real; I linked the numbers myself, and the Wikipedia editors are doing conservation work somebody has to do. Detection is not useless at the extremes either: raw, unedited chatbot output still gets caught, and Wikipedia’s deletion rule targets exactly that. My claim is narrower and worse: the error costs land on innocent humans, in scholarship money and honor courts and 24-year careers, while the machines that matter walk past the checkpoint for twenty dollars a month.
The sharper objection cuts at me. “A human approved it” is exactly what a content mill would say. True. My answer is that the mill will not tell you, and I just did. Disclosure, a signature, and receipts are the whole defense available to anyone. If that is not enough for you, nothing I could have typed by hand would be either.
The tell was never the em dash. The tell is writing nobody would stand behind, wearing a byline nobody can come find. Retire the detector. Read the words. Then come find me.
Sources
USA Today on the Quarterman case: https://www.yahoo.com/news/professors-using-chatgpt-detector-tools-093105927.html
EdSurge on Marley Stevens: https://www.edsurge.com/news/2024-04-04-can-using-a-grammar-checker-set-off-ai-detection-software
Gizmodo on Kimberly Gasuras: https://gizmodo.com/ai-detectors-inaccurate-freelance-writers-fired-1851529820
OpenAI classifier retirement: https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/
Stanford detector-bias study (Patterns, 2023): https://pmc.ncbi.nlm.nih.gov/articles/PMC10382961/
Turnitin false-positive update: https://www.turnitin.co.uk/blog/ai-writing-detection-update-from-turnitins-chief-product-officer
Vanderbilt disabling Turnitin: https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/
Ars Technica on the Declaration: https://arstechnica.com/information-technology/2023/07/why-ai-detectors-think-the-us-constitution-was-written-by-ai/
Wikipedia, Signs of AI writing: https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing
404 Media on the G15 deletion rule: https://www.404media.co/wikipedia-editors-adopt-speedy-deletion-policy-for-ai-slop-articles/
Excess-vocabulary study (PubMed abstracts): https://arxiv.org/abs/2406.07016
Graphite on the AI-article crossover: https://graphite.io/five-percent/more-articles-are-now-created-by-ai-than-humans
Ahrefs on AI content share: https://ahrefs.com/blog/what-percentage-of-new-content-is-ai-generated/
DIPPER paraphrase attack: https://arxiv.org/abs/2303.13408
NBC News on humanizers and keystroke war: https://www.nbcnews.com/tech/internet/college-students-ai-cheating-detectors-humanizers-rcna253878


