Twelve passages. Which ones are machine?
Most people are close to chance at this and quite confident they are not. Play first, read the explanation of each tell afterwards, and then find out why the whole exercise is less useful than it feels.
5 min · Beginner · Scored
Trust your instinct, then check it
Each round explains what gave the passage away, whichever way you answered.
“Reviewer 2 wanted the ablation table moved to the appendix. Reviewer 3 wanted it moved back. We did neither and nobody noticed.”
Every passage here was written by hand for this lab — the machine ones to exhibit documented tells, the human ones to carry documented human traits. That makes the tells learnable, which scraped samples would not.
In plain words
The strongest tell in current model writing is the “it’s not X, it’s Y” construction, usually followed by a sentence fragment for emphasis. Once you have noticed it you will see it everywhere, including in places where a human wrote it — which is exactly the problem with tells.
Why your score matters less than you think
This is the part the game is really for. Being good at spotting machine writing is a much weaker skill than it appears, and treating it as reliable causes real harm.
These tells are already going stale
Every trait above is a habit of a particular generation of models, and each one gets trained away as it becomes recognisable. Anything you learn here has a shelf life measured in months.
Detection tools do not work reliably
Automated AI-detectors produce false accusations at rates high enough to ruin the careers of the people they misclassify. Several universities have withdrawn them for exactly that reason.
Non-native writers are penalised worst
Careful, correct, slightly formal English — exactly what a second-language writer produces — is also what detectors flag as machine-written. The harm here falls unevenly and predictably.
Provenance beats detection
The durable answer is not better detection but signed provenance: cryptographic records of what was made where. Guessing from the text itself is a losing race.
In plain words
If you take one thing from this lab, make it this: never accuse someone of using AI on the basis of how their writing sounds. The false-positive rate is high, it falls hardest on people writing in a second language, and there is no way to prove innocence.
What this explains
Why model writing sounds like that
It was trained to produce the most probable continuation, then tuned by people rewarding answers that sound balanced and complete. Hedged, symmetrical, conclusive prose is the direct output of that objective.
Why editing matters more than prompting
The tells are surface habits. A writer who reads the draft and cuts every “it’s important to note” removes most of them in a minute — which is why AI-assisted writing that reads well is still human work.
Why detectors get deployed anyway
Institutions face a real problem and are sold a tool that claims to solve it. Understanding the failure mode is how you argue against a policy built on one.
Why provenance is the actual answer
Signed records of what was generated where survive model improvements. Guessing from output does not — the better models get, the worse detection performs, by construction.
Next
Knowing how it writes is knowing how it works
Every habit in this game traces back to the training objective. The next lab takes that objective apart, stage by stage.