The base version
Turing 1.0.0
A. M. Turing, "Computing Machinery and Intelligence," Mind, Vol. LIX, No. 236, October 1950, pp. 433–460.
Everything on this site is numbered against this document, so it is worth being exact about what it does and does not say. It is quoted here rather than paraphrased, because nearly every dispute about the Turing Test turns out to be a dispute about a paraphrase.
Two things about it surprise people who have only encountered it secondhand. Turing does not argue that machines think; he argues that the question is not worth asking in that form and substitutes a game for it. And the famous numbers are part of a prediction about the year 2000, not a pass mark he handed down.
The specification, in his words
The framing: 'Can machines think?'
I propose to consider the question, "Can machines think?" This should begin with definitions of the meaning of the terms "machine" and "think."
433. Mind, Vol. LIX, No. 236, October 1950.
Opening sentence of the paper, Section 1 'The Imitation Game'. Matches the JSTOR-scanned original exactly (the original typesets 'I PROPOSE' in small caps, a typographic choice, not a wording difference).
What the meaning of the words would have to be determined by (the Gallup poll passage)
If the meaning of the words "machine" and "think" are to be found by examining how they are commonly used it is difficult to escape the conclusion that the meaning and the answer to the question, "Can machines think?" is to be sought in a statistical survey such as a Gallup poll. But this is absurd.
433. Mind, Vol. LIX, No. 236, October 1950.
This is the passage that directly answers 'what the meaning of the words would have to be determined by' — ordinary usage, surveyed Gallup-poll style — and Turing's own rejection of that approach ('But this is absurd'). It appears earlier in Section 1 than the 'too meaningless to deserve discussion' line (which is a separate passage in Section 6, recorded as its own entry below). The brief's item 2 conflates these two passages; they are quoted separately here for precision. Verified against JSTOR scan, matches exactly with no transcription artifacts.
The dismissal: 'too meaningless to deserve discussion'
The original question, "Can machines think?" I believe to be too meaningless to deserve discussion. Nevertheless I believe that at the end of the century the use of words and general educated opinion will have altered so much that one will be able to speak of machines thinking without expecting to be contradicted.
442. Mind, Vol. LIX, No. 236, October 1950.
Section 6 'Contrary Views on the Main Question', immediately following the fifty-years prediction (see 'fifty_years_prediction' entry — same paragraph). Note this is a separate claim from 'in about fifty years' time': the fifty-years line predicts a technical capability (imitation-game performance by ~2000); this line predicts a shift in linguistic/social convention ('at the end of the century') about applying the word 'thinking' to machines. Popular summaries often merge the two into a single 'Turing predicted machines would think by 2000' claim; he explicitly declined to make that claim, calling the original question meaningless. Matches JSTOR scan exactly.
The imitation game: three players (A, B, C), their labels, and the interrogator's task
It is played with three people, a man (A), a woman (B), and an interrogator (C) who may be of either sex. The interrogator stays in a room apart front the other two. The object of the game for the interrogator is to determine which of the other two is the man and which is the woman. He knows them by labels X and Y, and at the end of the game he says either "X is A and Y is B" or "X is B and Y is A."
433. Mind, Vol. LIX, No. 236, October 1950.
Section 1. 'front' is a transcription artifact in the source PDF's text layer — the JSTOR-scanned original reads 'apart from the other two.' Quoted here exactly as the source PDF extracts, per the verbatim-against-source.url rule; flagged here so it isn't mistaken for a transcription error introduced in this document. Everything else in the quote matches the original exactly.
The teleprinter: why the channel is text
In order that tones of voice may not help the interrogator the answers should be written, or better still, typewritten. The ideal arrangement is to have a teleprinter communicating between the two rooms. Alternatively the question and answers can be repeated by an intermediary.
434. Mind, Vol. LIX, No. 236, October 1950.
Section 1, immediately following the game setup. Describes the mechanism (teleprinter / typewritten text) that enforces the text-only channel. Matches JSTOR scan exactly, no artifacts.
The prediction: fifty years, 10^9 storage, 70 per cent
I believe that in about fifty years' time it will be possible, to programme computers, with a storage capacity of about 109, to make them play the imitation game so well that an average interrogator will not have more than 70 per cent chance of making the right identification after five minutes of questioning.
442. Mind, Vol. LIX, No. 236, October 1950.
Section 6, opening paragraph, immediately preceding the 'too meaningless' line (same paragraph, recorded as a separate entry above). THREE transcription artifacts in the source PDF, confirmed against the JSTOR-scanned original: (1) the comma is misplaced — 'it will be possible, to programme computers, with a storage capacity...' should read 'it will be possible to programme computers, with a storage capacity of about 10⁹,...' (comma belongs after 'computers', not after 'possible'); (2) '109' is the original's superscript '10⁹' (ten to the ninth = about one billion) with the superscript formatting stripped by PDF extraction — read literally it looks like the integer one-hundred-and-nine, which it is not; (3) 'per cent chance' drops the original's period: '70 per cent. chance' (the period marks 'cent' as an abbreviation in 1950 British typesetting). CRITICAL READING: the 70 per cent figure is the INTERROGATOR's ceiling on making the RIGHT identification, not the machine's deception/pass rate. The commonly cited '30 per cent' is the arithmetic complement (100 minus 70) — i.e., the interrogator being wrong, and therefore the machine passing as human, at least 30% of the time — and is correct as a derived figure, but it is not a number Turing wrote down, and it was never posed as a fixed pass/fail bar for machines to clear. It was a prediction about the state of technology circa the year 2000, not a definition of what 'passing the Turing test' means. 'Five minutes of questioning' is likewise part of this specific 1950 prediction about future computers, not a rule Turing laid down elsewhere for how long the imitation game itself must run (no duration appears in the Section 1 game description).
Permission to lie / deliberate mistakes in arithmetic
It is claimed that the interrogator could distinguish the machine from the man simply by setting them a number of problems in arithmetic. The machine would be unmasked because of its deadly accuracy. The reply to this is simple. The machine (programmed for playing the game) would not attempt to give the right answers to the arithmetic problems. It would deliberately introduce mistakes in a manner calculated to confuse the interrogator.
448. Mind, Vol. LIX, No. 236, October 1950.
Section 6, '(4) The Argument from Various Disabilities'. This is the passage on deliberate arithmetic mistakes; Turing does not use the word 'lie' anywhere in the paper (verified: no occurrence of 'lie'/'lying'/'deceiv*' in the full text). The nearest thing to an explicit license to deceive is structural, in the Section 1 game description itself: 'It is A's object in the game to try and cause C to make the wrong identification' (p. 433) — deception is built into the game's premise from the start, not stated as a separate permission. Deliberate wrongness specifically about arithmetic is this passage. Matches JSTOR scan exactly, no transcription artifacts found in this passage.
Numbers that get reported wrong
These are the figures that circulate in a mangled form, including in places that should know better. The mangling is not random: it consistently makes the test sound either easier or more official than it was.
seventy_vs_thirty
- Commonly said
- Popular accounts often state '30% is the threshold for passing the Turing test' or invert the framing to '70% deception rate,' treating the number as either a formal pass/fail bar Turing prescribed, or as a machine win-rate figure.
- Actually
five_minutes_scope
- Commonly said
- Cited as if Turing specified a fixed test duration as part of the imitation game's rules, or as a general methodological requirement for any valid Turing test.
- Actually
storage_capacity_units
- Commonly said
- PDF/OCR transcriptions of the paper (including the machine-fetchable mirror used as source.url in this file) commonly lose the superscript formatting, rendering '10⁹' as the flat digits '109' — which reads as the integer one-hundred-and-nine rather than ten-to-the-ninth-power.
- Actually
two_party_vs_three_party
- Commonly said
- The 2024 Jones & Bergen GPT-4 result (54%) is frequently cited in popular press and summaries as evidence an AI 'passed the Turing test' in the classical sense, without noting the format difference from the 2025 result.
- Actually
persona_prompt_dependency
- Commonly said
- Headline figures are often repeated without the PERSONA-prompt qualifier, implying these are baseline/default capabilities of the models.
- Actually
no_explicit_fifty_percent_threshold
- Commonly said
- Treated as though Turing himself defined 50% win rate as the formal passing threshold for the imitation game.
- Actually
two_predictions_conflated
- Commonly said
- Summarized as a single claim: 'Turing predicted machines would think by the year 2000.'
- Actually
gpt5_not_gpt45_in_study3
- Commonly said
- Coverage of the PNAS paper's 15-minute replication sometimes reports the 59%/56% figures as if they simply extend or reproduce the original 73%/56% (5-minute) result with more time.
- Actually
pnas_abstract_body_order
- Commonly said
- N/A — this is an internal inconsistency in the primary source itself, flagged here as a trap for anyone quoting the PNAS abstract in isolation.
- Actually
What has actually been measured
The empirical record against this protocol specifically. It is short, which is itself worth noticing: the test everyone has an opinion about has been run under pre-registration a handful of times.
GPT-4.5 (OpenAI API model string "gpt-4.5-preview-2025-02-27"), PERSONA prompt
2025-03-31 verifiedLarge Language Models Pass the Turing Test (Jones & Bergen, 2025) — preprint
When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant.
- Participants
- 1023
- Duration
- 5-minute hard time limit; median completed game was 8 messages across 4.2 minutes
- Interrogators
- Two independent populations: (1) 126 UC San Diego psychology undergraduates (subject pool, course credit) and (2) 158 Prolific workers (paid $13.75). 1023 total games analyzed after exclusions (445 undergraduate + 576 Prolific, plus additional games noted in the aggregate total).
- Baseline
- ELIZA 23% and GPT-4o (gpt-4o-2024-08-06) 21%, both significantly below chance (50%); human witnesses' implied win rate is the complement of interrogator accuracy
Three-party, live, simultaneous 5-minute text conversations: one interrogator vs. one human witness + one AI witness at once, interrogator then judges which was human. Pre-registered (two separate pre-registrations, one per population). Overall win rate 73% blends 69% (undergraduate study) and 76% (Prolific study).
Source · doi 10.48550/arXiv.2503.23674
LLaMa-3.1-405B (TogetherAI API model string "meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo"), PERSONA prompt
2025-03-31 verifiedLarge Language Models Pass the Turing Test (Jones & Bergen, 2025) — preprint
LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time—not significantly more or less often than the humans they were being compared to
- Participants
- 1023
- Duration
- 5-minute hard time limit; median 8 messages / 4.2 minutes
- Interrogators
- Same two populations as GPT-4.5-PERSONA: 126 UCSD undergraduates and 158 Prolific workers
- Baseline
- ELIZA 23% and GPT-4o 21%, both significantly below chance
Three-party, live, pre-registered. 56% overall blends 45% (undergraduate study) and 65% (Prolific study) — LLaMa-PERSONA's win rate was significantly above chance in the Prolific study but not in the undergraduate study.
Source · doi 10.48550/arXiv.2503.23674
GPT-4.5 ("gpt-4.5-preview-2025-02-27"), NO-PERSONA (minimal baseline prompt)
2025-03-31 verifiedLarge Language Models Pass the Turing Test (Jones & Bergen, 2025) — preprint
GPT-4.5-NO-PERSONA and LLAMA-NO-PERSONA had overall win rates of 36% and 38% respectively
- Participants
- 1023
- Duration
- 5-minute hard time limit
- Interrogators
- Same two populations: 126 UCSD undergraduates, 158 Prolific workers
- Baseline
- ELIZA 23%, GPT-4o-NO-PERSONA 21%; GPT-4.5-NO-PERSONA did not perform significantly better than ELIZA in either population (p > 0.06)
Three-party, live, pre-registered. Same underlying model as the 73% PERSONA result, with only the system prompt changed to a minimal baseline instruction (no persona guidance). This is the number that shows the 73% headline figure is prompt-dependent, not an intrinsic model capability figure.
Source · doi 10.48550/arXiv.2503.23674
LLaMa-3.1-405B ("meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo"), NO-PERSONA
2025-03-31 verifiedLarge Language Models Pass the Turing Test (Jones & Bergen, 2025) — preprint
GPT-4.5-NO-PERSONA and LLAMA-NO-PERSONA had overall win rates of 36% and 38% respectively
- Participants
- 1023
- Duration
- 5-minute hard time limit
- Interrogators
- Same two populations: 126 UCSD undergraduates, 158 Prolific workers
- Baseline
- ELIZA 23%, GPT-4o-NO-PERSONA 21%; LLAMA-NO-PERSONA outperformed ELIZA in the Prolific study (z=2.42, p=0.016) but not the undergraduate study
Three-party, live, pre-registered. Same underlying model as the 56% PERSONA result, minimal baseline prompt instead of the persona prompt.
Source · doi 10.48550/arXiv.2503.23674
GPT-4.5 and LLaMa-3.1-405B, PERSONA prompt (studies 1-2, unchanged from preprint)
2026-05-19 verifiedLarge language models pass a standard three-party Turing test (Jones & Bergen, PNAS)
Overall, across both studies, GPT-4.5-PERSONA had a win rate of 73% (69% with UCSD undergraduates, 76% with Prolific participants). LLAMA-PERSONA achieved a win rate of 56% (Undergraduates: 45%, Prolific: 65%).
- Participants
- 1023
- Duration
- 5-minute hard time limit (studies 1-2)
- Interrogators
- 126 UCSD undergraduates + 158 Prolific workers (identical populations/design to the arXiv preprint for studies 1-2)
- Baseline
- ELIZA 23%, GPT-4o 21%
Peer-reviewed version of record. Title changed from the preprint's 'Large Language Models Pass the Turing Test' to 'Large language models pass a standard three-party Turing test'. Published online 2026-05-19, print 2026-05-26, PNAS vol. 123 issue 21, article e2524472123, CC BY 4.0. The headline studies 1-2 figures (73%/56% persona, 36%/38% no-persona, 23%/21% baselines) are UNCHANGED from the preprint — verified by direct text match in the full-text XML. The journal version ADDS a third preregistered study not present in the March 2025 preprint (see jb_pnas_study3 entries below).
Source · doi 10.1073/pnas.2524472123
GPT-5 (OpenAI API model string "gpt-5-chat-latest"), PERSONA prompt — substituted for GPT-4.5, which OpenAI had deprecated by the time of this replication
2026-05-19 verifiedLarge language models pass a standard three-party Turing test (Jones & Bergen, PNAS) — Study 3, 15-minute replication
GPT-5-PERSONA 59.3% 89 61 150
- Participants
- 205
- Duration
- 15-minute fixed duration; interrogators were prevented from ending games early, so all games ran the full 15 minutes (unlike studies 1-2's 5-minute cap)
- Interrogators
- 205 Prolific participants after exclusions (mean age 41.1); 396 total games (150 GPT-5-PERSONA games, 151 LLAMA-PERSONA games, 95 ELIZA games)
- Baseline
- ELIZA 16.8% (16 wins / 95 games); GPT-5-PERSONA and LLaMa-PERSONA both significantly outperformed ELIZA (p<0.001)
Three-party, live, pre-registered replication testing whether results hold with an extended, uncapped interaction window. Only 3 AI witness types tested (LLaMa-3.1-405B-PERSONA, GPT-5-PERSONA, ELIZA — no NO-PERSONA arm in this study). GPT-5-PERSONA's win rate was only marginally above chance after correction for multiple comparisons (z=-2.27, p=0.023, p_corrected=0.05), a notably weaker result than the 73% from the original 5-minute GPT-4.5 study, though not a like-for-like comparison since the model changed from GPT-4.5 to GPT-5.
Source · doi 10.1073/pnas.2524472123
LLaMa-3.1-405B, PERSONA prompt (same model/prompt as studies 1-2)
2026-05-19 verifiedLarge language models pass a standard three-party Turing test (Jones & Bergen, PNAS) — Study 3, 15-minute replication
LLAMA-PERSONA 55.6% 84 67 151
- Participants
- 205
- Duration
- 15-minute fixed duration, no early exit
- Interrogators
- 205 Prolific participants after exclusions; 151 LLAMA-PERSONA games
- Baseline
- ELIZA 16.8% (16/95)
Three-party, live, pre-registered. LLaMa-PERSONA's win rate (55.6%, reported rounded as 56% in the abstract) was NOT significantly above chance in this study (p=0.17), unlike GPT-5-PERSONA which was marginally significant.
Source · doi 10.1073/pnas.2524472123
GPT-4 (OpenAI API model string "gpt-4-0613")
2024-05-09 verifiedPeople cannot distinguish GPT-4 from a human in a Turing test (Jones & Bergen, 2024) — preprint; later peer-reviewed at ACM FAccT 2025 (DOI below) with two additional co-authors (Ishika Rathi, Sydney Taylor)
We used a two-player formulation of the game, where a single human interrogator conversed with a single witness who was either a human or a machine. While this differs from Turing’s original three-player formulation, it has become a standard operationalisation of the test because it eliminates the confound of the third player’s humanlikeness and is easier to implement
- Participants
- 402
- Duration
- 5-minute time limit
- Interrogators
- Recruited via Prolific; 402 interrogators (168 female, 212 male, 9 nonbinary, 13 prefer not to say; mean age 35). 500 total participants recruited (including witnesses) to yield 101 Human, 101 GPT-3.5, 100 GPT-4, and 100 ELIZA games after exclusions and slight oversampling.
- Baseline
- ELIZA 22%, GPT-3.5 ("gpt-3.5-turbo-0125") 50%, actual human witnesses 67%
TWO-PARTY, not three-party: a single interrogator conversed with a single witness (human OR AI, not both at once) and judged whether that witness was human. Pre-registered. The paper explicitly flags this as a deviation from Turing's original three-party design, adopted because it is the standard operationalisation in the AI-detection literature. Peer-reviewed version at ACM FAccT 2025: DOI 10.1145/3715275.3732108.
Source · doi 10.48550/arXiv.2405.08007
Two rigorous studies, opposite verdicts
The best-known result is not the only careful one, and the other careful one disagrees with it. Temtsin, Proudfoot, Kaber and Bartneck ran a pre-registered, IRB-approved three-party game with no time limit at all, which is closer to the 1950 text than any five-minute protocol. GPT-4-Turbo was correctly identified as non-human by 36 of 37 interrogators.
Their paper predates the Jones and Bergen preprint by two months, so it is not a rebuttal to it. Two groups set out independently to run Turing's game properly and came back with opposite answers.
It is tempting to say the difference is duration, and that would be tidy. It would also be wrong. The studies differ in the model tested, the population recruited, and the prompting strategy, and Jones and Bergen's own data shows prompting alone moving a result from 21 percent to 73 percent at a fixed five minutes. Duration is one unmatched variable among several, and nobody has held the others still.
Which is the argument for numbering these protocols, made better by the evidence than by anything on the front page. As long as both studies are called the Turing Test, they read as a contradiction. Called 1.1.1 and 1.3.2, they read as two measurements that were never the same measurement.
Correction
An earlier pass of this site's research concluded that no independent pre-registered three-party study existed, at secondary confidence. That was wrong, and it is recorded here rather than quietly overwritten. The study had circulated under a different title, and a search for one missed the other.
Why 1.0.0 and not 0.1.0
Under semantic versioning, 1.0.0 marks the point at which a public interface is declared stable and further changes have to be classified. That is exactly what the 1950 paper did: it fixed a procedure precise enough that departures from it are visible as departures.
Everything since has been an unversioned change to a stable interface, shipped without a changelog. The instrument exists to give those changes numbers.