This guide publishes a list of words that are supposed to mark generated prose. Two billion words of human writing put nineteen of them between 0.76 and 1.94 times their pre-ChatGPT rate. 796 other words moved by two times or more and none of them is on the list.

The word list was written down before the counting

The list was not chosen here. It comes from this project's research log and the two sources that log cites: a copywriters' thread from August 2024 and Wikipedia's page on signs of AI writing. The list was published before the count, so naming words in advance is the correct procedure here.

None of the nineteen moved by more than a factor of two

Bar chart of nineteen words published as AI writing tells, each showing how its rate in Hacker News comments changed after ChatGPT. Unlocks leads at 1.94 times and delve follows at 1.88; underscore is lowest at 0.76; the rest cluster against the no-change line. The axis runs from 0.5 to 2 times.
The nineteen pre-registered words that clear the census floor, on an axis that stops at double. Same corpus, windows and method as the figure above; rendered from data/derived/copy-tells.json by scripts/metrics/render-copy-tell-charts.mjs.

Nineteen of the twenty-six surface forms clear the census floor of 2,000 occurrences. The largest riser is “unlocks” at 1.94 times and the largest faller is “underscore” at 0.76, against a census-wide median of 0.972. 796 words in this census moved by a factor of two or more and none of them is on the list.

Leverage was already thirty-one in a million

Before ChatGPT existed, “leverage” ran at 31.19 occurrences per million tokens on Hacker News. “Crucial” ran at 14.58 and “unlock” at 13.30. A word that common says little about who wrote the sentence it appears in.

The words that did move

The top of the ranking is “claude” at 601 times its pre-ChatGPT rate, then “slop” at 286 and “anthropic” at 282. Ten of the top twelve name an AI system or a failure mode and the other two stay in the figure because the ranking is raw. The words that rose on Hacker News are the words people use to discuss AI systems, not the words those systems are accused of producing.

Delve rose with the post about it

Line chart of the monthly frequency of the word delve in Hacker News comments, 2015 to 2026, per 100,000 items. The series holds near its baseline of 8.8 for eight years, crosses the threshold in April 2024, and peaks at 28.3 in April 2026.
“Delve” per 100,000 Hacker News items, monthly. Exact occurrence counts from the Hacker News search API with denominators from the site's sequential item ids, collected August 18, 2026 by scripts/metrics/hn-word-rates.mjs. Rendered from data/hn-word-rates.csv by scripts/metrics/render-charts.mjs. A different instrument from the figures above, which is why the multiple differs.

“Delve” reached 1.88 times and ranks 928th. A second instrument built on the Hacker News search API puts its inflection in April 2024, the month Paul Graham's post about the word was seen seventeen million times. The rise arrives with the public argument about the word and measures the conversation rather than the prose.

This is a corpus of human writing

Hacker News comments are written by people, so this cannot measure how often a model emits a word. It measures how common each word already was in human writing, and nothing here says models do not overuse them. A tech forum is also not marketing copy, where “elevate” and “unleash” presumably run higher than 2.37 and 1.04 per million.

Seven words this census cannot see

The census keeps a word only if it appears 2,000 times overall and 200 times in the baseline. Seven forms on the list fall under that floor and nothing here is evidence about them. The tokenizer reads single words, so stock phrases like “here's the kicker” and the rule-of-three list cannot be counted by this instrument at all. All figures above come from the collector's first stage only.

The list was never rare

A word is evidence of generated prose when it was rare before and is common now. The twelve words at the top of the ranking started at a median of 0.37 per million tokens and the nineteen tell words started at 3.22. The vocabulary row of our tell index is retired as a presence test, since finding “leverage” on a page says little about who wrote it. The copywriters were probably describing density and a word census cannot test that.