A versioned history of a moving target
The Turing Test was passed.
The response was a rename.
In 1950 Alan Turing proposed a procedure specific enough to actually run. Three participants, five minutes of typed conversation, an interrogator asked to say which of two hidden correspondents was the human. He predicted that by the year 2000 a machine would survive that questioning about thirty percent of the time.
It took until 2025. A pre-registered study at UC San Diego ran the three-party protocol across 1,023 games. GPT-4.5 was judged to be the human 73 percent of the time, which is to say it was picked as the human more often than the actual humans were.
That figure carries a condition, and anyone repeating it without the condition is doing the thing this site is about. The 73 percent was achieved with a prompt instructing the model to adopt a humanlike persona. Given a minimal prompt instead, the same model scored 36 percent, which the authors could not separate from ELIZA at 23.
So the first honest question is whether coaching disqualifies the result. Turing answered that one himself, in the same paper, while dismissing the objection that a machine would give itself away by being too good at arithmetic:
The machine (programmed for playing the game) would not attempt to give the right answers to the arithmetic problems. It would deliberately introduce mistakes in a manner calculated to confuse the interrogator.
Turing 1950, section 6. In context.
A machine programmed for playing the game is not cheating at the game. That is the entry requirement, described in advance, by the person who wrote the test.
The result is real and it is not a rout. A 2026 replication in PNAS ran the full fifteen minutes with no early exit, substituted GPT-5 after GPT-4.5 was deprecated, and got 59.3 percent: above chance, but only marginally so once corrected for multiple comparisons. All of that is laid out with the numbers.
The ordinary response to a saturated benchmark is a new benchmark with a new name. ImageNet gave way to GLUE, GLUE to SuperGLUE, SuperGLUE to MMLU, MMLU to MMLU-Pro. Each retirement was announced. Nobody argues that MMLU was secretly ImageNet all along.
The Turing Test did not get a successor. It got edited.
It has no maintainer, no version number, and no changelog, so for seventy-five years the specification has been revised in public: usually within days of something clearing it, almost always by people who present the revision as a clarification of what the test had meant the whole time. The revisions are not tracked anywhere, which is precisely what makes them work.
This site tracks them.
Why the numbering is not a joke
Two months before that study, a different group ran the same game with no time limit at all, pre-registered and IRB-approved. GPT-4-Turbo was correctly identified as non-human by 36 of 37 interrogators. Not a narrow miss. A rout in the other direction.
So the two most careful three-party experiments anyone has run came back with opposite verdicts, and neither is sloppy. The tempting explanation is duration, and it is not available: the studies also differ in the model tested, the people recruited, and how the model was prompted. Prompting alone moved a result from 21 percent to 73 percent inside a single paper.
That is the case for version numbers, and the evidence makes it better than any argument could. While both experiments are called the Turing Test, they look like a contradiction that somebody must be wrong about. Called Turing 1.1.1 and Turing 1.3.2, they stop being a contradiction and start being two results with a stated relationship, which is the thing a person could actually act on.
Both are laid out with the numbers on Turing 1.0.
One
The Record
23 entries, 1637 to 2026, each with what was actually said. 11 of them moved the bar.
Two
The Instrument
Build a Turing Test out of nine choices. Get a version number. Find out who already proposed it.
Three
The Registry
15 published successors, each given a version against the 1950 original.
The shape of it
Plotting where the bar sat after each of the moments below produces a staircase. It is flat for the thirty years when nothing could clear it, and it moves three times in the eleven years since something could.
Higher means a harder or different test. Crossing a whole number is a breaking change: results below it stop applying. This line is the site's own reading of the discourse, not a measurement of it.
Read the steps as text, with the reasoning for each
-
1950 Turing publishes 1.0.0
The specification as written. Five minutes, text, a general-public interrogator, 30 percent.
-
1966 ELIZA 1.1.0
A few hundred lines of pattern matching convinced people it understood them. The lesson drawn was not that the test was passed but that a naive interrogator is not a real test, and the assumed judge quietly got better.
-
1980 The Chinese Room 2.1.0
Searle's argument does not raise the bar so much as relocate it. Behavior stops being sufficient and what happens inside starts to count, which is the exact move Turing designed the game to rule out. Under this site's scheme that is a breaking change, and it is the first one.
-
1991 The Loebner Prize 1.0.0
The only time in seventy-five years the bar moved down. To make an annual contest runnable, early Loebner rounds restricted conversation to a declared topic, which is materially easier than the 1950 protocol. Money was involved.
-
2014 Eugene Goostman 2.3.0
A chatbot posing as a Ukrainian teenager cleared 30 percent and was rejected within days. The rejection was correct, and it also permanently raised three things at once: no persona excuses, longer sessions, and a threshold above the one Turing named.
-
2018 Google Duplex 3.3.0
A system booked a haircut by telephone and the reaction was not that the test had been passed but that machines must announce themselves. The bar absorbed voice, and acquired a disclosure norm that has no equivalent in the 1950 paper.
-
2025 GPT-4.5 passes 5.9.2
The protocol was run as written, under pre-registration, and cleared decisively. The response now on offer requires expert adversarial interrogators, unbounded time, images, memory across sessions, and statistical indistinguishability rather than a deception rate. Every one of those is defensible on its own. Together they describe a test nobody has built, nobody has run, and nobody has proposed retiring the old name for.
The finding
Sorting the proposed successors onto a common set of axes turns up something that is hard to see one paper at a time. Of the 15 collected here, 9 do not measure conversational indistinguishability at all.
They measure creative output under constraints, or pronoun disambiguation from a fixed corpus, or how closely a model's internal activations match brain recordings, or a six-part rubric for autonomy. Several have no interrogator at all. Those are real questions and one or two are better questions. But a test of creativity is not a stricter Turing Test, and calling it one is not raising the bar. It is changing the sport while leaving the old scoreboard bolted to the wall.
So this site refuses to give those a version number. They are recorded as forks: descended from the original, not comparable to it, and honest about which.
What this is not
It is not a standards body. Nobody ratifies anything here, there is no submission process, and no result on this site has been blessed by anyone. A specification with one author and no adopters is a document, and pretending otherwise would be its own kind of goalpost-moving.
It is also not an argument that machines think. Turing thought that question was too meaningless to deserve discussion, which is why he replaced it with a procedure in the first place. The claim here is narrower and duller: the procedure was run, the procedure was cleared, and the response was to quietly rewrite the procedure. That is worth writing down.
Read Turing 1.0.0, the 1950 protocol restated in the author's own words, with every number sourced.