Skip to content
Goalposts

The registry

Published successors

Every proposal here was published as an improvement on, or replacement for, the Turing Test. Each has been mapped onto the same nine axes and given a version number, so that for the first time they can be compared to each other rather than only to the original.

Two columns do most of the work. Version says whether a proposal is a stricter version of the imitation game or a different measurement. Ever run says whether anybody has actually tested a system against it.

15 proposals. 9 do not measure conversational indistinguishability and are recorded as forks. 7 have ever been run against a system.

Versions of the imitation game

These stay recognizably Turing's test. A 1.x proposal only makes it harder; a 2.x or above changes what is being measured, and results do not carry across.

Integrative Turing-like tests for Language and Vision

2022 verified
4.1.0

Mengmi Zhang, Elisa Pavarino, Xiao Liu, Giorgia Dellaferrera, Ankur Sikarwar, Caishun Chen, Marcelo Armendariz, Noga Mudrik, Prachi Agrawal, Spandan Madan, Mranmay Shetty, Andrei Barbu, Haochen Yang, Tanishq Kumar, Shui'Er Han, Aman Raj Singh, Meghna Sadwani, Stella Dellaferrera, Michele Pizzochero, Brandon Tang, Yew Soon Ong, Hanspeter Pfister, Gabriel Kreiman · arXiv preprint

We systematically benchmark AI's ability to imitate humans in three language tasks (image captioning, word association, conversation) and three vision tasks (color estimation, object detection, attention prediction), collecting data from 636 humans and 37 AI agents. Next, we conducted 72,191 Turing-like tests with 1,916 human judges and 10 AI judges.
Format: Displaced Modality: Multimodal Threshold: Chance (50%)

Ever run

Extensive — the strongest empirical entry in the corpus alongside TuringBench, Rahimov et al., and the Generalized Turing Test. For the Conversation task: judges correctly called genuine human exchanges 'human' 66% of the time, and correctly called AI-generated exchanges 'AI' only 47% of the time; overall imitation-detectability was 0.57 for conversation (0.57 image captioning, 0.53 word association) — close to chance. In AI-AI conversation pairs, Blenderbot exchanges were judged human 67% of the time, more often than genuine human-human exchanges (64%). A simple SVM judge trained on single sentences matched or beat human judges at catching machine text despite far less context. Separately, the authors ran a live, interactive, 3-party version (one judge, one human agent, one AI agent, judge asking questions in real time) and found human judges reached up to 100% accuracy in an early pilot, degrading toward chance only as the number of allowed exchanges shrank — a striking contrast (offline/displaced judging near chance vs. live/interrogative judging near-perfect) directly relevant to this corpus's thesis that format changes what is measured.

How this was mapped onto the axes, and where the paper is silent

A genuine boundary case, and the call is close. The paper's headline contribution is a large-scale, offline, forced-choice test: crowd-collected human and AI responses to 3 language tasks and 3 vision tasks are shown to judges who pick which of two responses is human. Only one of the six tasks (conversation) is conversational in Turing's sense; color estimation and object detection are psychophysical judgments with no dialogue. Mapped as non-fork because the underlying evidentiary method is identical across all six tasks and identical to Turing's own move: behavioral output, forced-choice human-vs-machine judgment, nothing mechanistic — the paper generalizes indistinguishability judgment across modalities rather than replacing it with a different construct (contrast with the NeuroAI Turing Test in this corpus, which swaps behavior for internal representations). format='displaced' reflects the primary, large-N test: judges read/view pre-collected transcripts and vision stimuli rather than interrogating live. The paper also ran a much smaller supplementary live 3-party experiment (one judge simultaneously conversing with one human and one AI agent, GPT-3.5-Turbo) that gave very different results (see empiricalResults) — not the paper's primary config, but a striking finding about how format changes what is measured. duration is defaulted: the paper caps conversations by exchange count (up to 24, extended to 48), not minutes. threshold='chance' is an editorial call: the authors never declare a formal numeric pass bar but repeatedly interpret their 'imitation detectability' scores (0.53-0.57) against an implicit chance/0.5 reference. rigor is defaulted to informal despite substantial disclosure (public code/data, cross-platform replication, corrected permutation tests), since no pre-registration is mentioned. Version note: v1 (2022-11-23, the date recorded above) already contained the core methodology but with smaller numbers (769 human agents, 24 AI agents, 21,570 trials); the abstract quoted above reflects v3 (Sep 2025) figures. The live-interrogation sub-experiment using GPT-3.5-Turbo postdates v1, since that model did not exist until March 2023.

Can Machines Imitate Humans? Integrative Turing-like tests for Language and Vision Demonstrate a Narrowing Gap · doi 10.48550/arXiv.2211.13087

The Turing Test Is More Relevant Than Ever

2025 verified
1.2.0

Avraham Rahimov, Orel Zamler, Amos Azaria · arXiv preprint

This study argues against discarding the Turing Test, proposing instead using more refined versions of it, for example, by interacting simultaneously with both an AI and human candidate to determine who is who, allowing a longer interaction duration, access to the Internet and other AIs, using experienced people as evaluators, etc.
Interrogator: Motivated amateur Threshold: Chance (50%)

Ever run

Real and central to the paper. Section 4, Tables 1-3: four between-subjects MTurk conditions crossing {Simple, Enhanced} x {no prompt engineering, with prompt engineering}, using Llama 3.2 1B as the AI throughout. Results: Simple/No-Prompt, n=41, 68.29% correct identification; Enhanced/No-Prompt, n=29, 93.10% correct (chi-sq(1)=4.97, p=.026); Simple/With-Prompt, n=41, only 43.90% correct — the prompt-engineered LLM was misidentified as human more than half the time, read by the authors as the AI 'passing' this weaker version; Enhanced/With-Prompt, n=31, 70.97% correct (chi-sq(1)=6.21, p=.013) — the same LLM clearly failed the harder, three-party version. This directly supports the abstract's central claim with real numbers: an LLM can pass a weak version of the test but fails a more robust one. No independent replication of this specific Simple-vs-Enhanced protocol found.

How this was mapped onto the axes, and where the paper is silent

Important correction to my own working assumptions: the abstract's laundry list (longer duration, Internet access, expert evaluators) describes the paper's ASPIRATIONAL 'Ultimate Turing Test / Turing Test 2.0' discussed in Section 6, which is NOT what was actually run. What the paper actually implements and tests (Section 3) is narrower: a 'Simple Turing Test' (two-party, single chat window, 2-minute limit) versus an 'Enhanced Turing Test' (three-party: a tester converses simultaneously with a human responder and an AI in two separate windows, exactly Turing's 1950 structure, 5-minute limit). Internet access and domain-expert evaluators are proposed as future work, never implemented; extended duration also never appears in the tested design — the Enhanced Test's 5 minutes matches, rather than exceeds, the 1950 baseline. Config maps to the Enhanced Test, the paper's actually-validated, recommended design and central empirical claim. format='three-party' because the Enhanced Test's dual-chat, simultaneous-comparison structure matches dimensions.json's three-party definition directly. interrogator='motivated' rather than 'naive': testers are Mechanical Turk workers screened for quality and paid a bonus specifically for correctly identifying the human/AI — financially incentivized to win, though not domain experts; a reader could push this toward 'naive' instead. threshold='chance': the paper interprets its results relative to a 50% cutoff throughout (e.g. treating 43.9% correct-identification as the AI having 'passed'). rigor='informal': real statistics (chi-squared tests) and full model/prompt disclosure, but no pre-registration.

The Turing Test Is More Relevant Than Ever · doi 10.48550/arXiv.2505.02558

Dual Turing Test

2025 verified
2.2.0

Alberto Messina · arXiv

In this short note, we propose a unified framework that bridges three areas: (1) a flipped perspective on the Turing Test, the "dual Turing test", in which a human judge's goal is to identify an AI rather than reward a machine for deception; (2) a formal adversarial classification game with explicit quality constraints and worst-case guarantees; and (3) a reinforcement learning (RL) alignment pipeline that uses an undetectability detector and a set of quality related components in its reward model.
Format: Displaced Interrogator: Motivated amateur Threshold: Chance (50%)

Ever run

Never run. None. The paper self-describes as a 'short note' and is purely a formalization. No dataset, no participants, no trained detector, no reported accuracy numbers; Section 8 ('Proposed Immediate Actions') explicitly lists building a pilot benchmark as future work — it does not yet exist. The RL alignment pipeline (Sections 4-5) is likewise presented as a proposed training loop with no runs reported. No independent implementation found.

How this was mapped onto the axes, and where the paper is silent

The judge stays human throughout the paper's own contribution (Section 3): a fixed judge is presented, per round, with an unlabeled pair of replies — one human, one AI — to a sampled prompt, and must say which is which. The 'Inverted Turing Test' the abstract mentions is explicitly cited as prior art (Watt's classic variant where a machine judges), not this paper's own mechanism — so 'inverted' (machine judges which interlocutor is human) does not apply; this is still a human-judge game with the payoff flipped from being deceived to catching the deception. format='displaced' is a judgment call: prompts are drawn from a fixed prompt space rather than authored/adapted live by the judge, and the judge only classifies a pre-generated reply pair rather than conducting an interrogation — closer to reading a transcript than steering one; a reader could argue for 'three-party' instead since both outputs are presented side by side for one comparative verdict. threshold is the hardest fit: the paper's own pass criterion is a minimax DETECTION accuracy for the judge (>= 0.70 in one worked example), the mirror image of this schema's deception-rate framing; its one explicit statistical anchor is a binomial test against a 50% chance baseline, which is what 'chance' is mapped from here, but no option in this schema was built for a flipped pass criterion. interrogator='motivated' reflects that the entire premise is a judge deliberately trying to catch the AI, not Turing's naive general-public interrogator, though no domain-expertise or tool-use is specified. Sections 4-5 (an RL alignment pipeline using the detector as a reward-model critic) are a separate training-procedure proposal layered on top of the dual test, not part of the test itself — noted as a design element the schema doesn't capture, but not enough to push this into fork territory since Section 3's core mechanism is still a human-vs-AI discrimination game.

Dual Turing Test: A Framework for Detecting and Mitigating Undetectable AI · doi 10.48550/arXiv.2507.15907

Energy Efficient Imitation Game

2025 verified
2.0.0

Adam Winchell · arXiv

This work expands upon the original imitation game by accounting for an additional factor: the energy spent answering the questions. By adding the constraint of energy, the new test forces us to evaluate intelligence through the lens of efficiency, connecting the abstract problem of thinking to the concrete reality of finite resources.
Resource constraint: Energy budget

Ever run

Never run. None — a philosophical essay with no experimental section, run against no real system. The central instrument, the 'psychoergometer,' is explicitly described as nonexistent: 'psychoergometers do not exist; and if they did, it would, for obvious reasons, be impractical to use them at scale.' There is no Pareto frontier, no measured joules-per-answer data, and no deception-rate-vs-energy tradeoff plotted anywhere in the paper — verified by full-text search: 'Pareto' and 'joule' each occur zero times in the document. The paper is an argumentative essay (motivating example, formal game description in prose, a 'Contrary Views' section, closing thoughts on a hierarchy of intelligence and the Halting Problem), not an empirical study. No independent implementation found.

How this was mapped onto the axes, and where the paper is silent

A genuine same-game extension, not a fork: the paper explicitly keeps and re-derives Turing's three-party structure ('It is played with three players: the liar, the truthteller, and the interrogator... What will happen when a machine takes the part of the liar in this game?') and states, 'We can actually embed the original imitation game within this new game by asking "Can machines think?"' — so format='three-party' and constraint='energy' are both grounded directly in the text, the latter being the paper's whole contribution. Everything else is left unspecified by the paper and defaulted: no channel restriction, no session-persistence discussion, no time bound for the interrogation itself, no interrogator-sophistication spec, and no numeric pass criterion is ever given. evidence='behavioral' rather than 'mechanistic': the energy measurement (via a hypothetical 'psychoergometer') is an external property of the answering process, not an inspection of internal representations.

Can machines think efficiently? · doi 10.48550/arXiv.2510.26954

A Rigorous Turing Test

2025 verified
1.3.2

Sharon Temtsin, Diane Proudfoot, David Kaber, Christoph Bartneck · arXiv

Several studies claim that large language models have passed the Turing Test and hence can "think", yet none follow Turing's original instructions precisely. Passing the test holds significance as evidence that a machine demonstrates human-like intelligence, and as a marker for artificial-general intelligence in commercial and legal domains.
Duration: Unbounded Methodological rigor: Pre-registered and disclosed

Ever run

Yes — the strongest, most rigorously reported empirical run in this corpus. Pre-registered, IRB-approved. Recruited 210 participants across 42 CIHG and 42 man-imitates-woman benchmark game (MIWG) trials; after exclusions, 185 participants remained across 37 trials of each game type. Machine witness: GPT-4-Turbo via the OpenAI API (temperature 0.75, max tokens 1000, frequency penalty 0.0). Headline finding: in the CIHG, only 1 of 37 interrogators misidentified GPT-4-Turbo as human — 36/37 (97%) correctly identified it as non-human, significantly above the 50% chance level (p<0.05, 95% CI 88-100%). Mean CIHG duration ~14 minutes. The MIWG benchmark produced near-chance identification accuracy (~43%, 95% CI 28-59%, not significantly different from 50%) — interrogators could not reliably tell a man imitating a woman from an actual woman via text either. The paper's conclusion: 'GPT-4-Turbo did not pass the test,' directly contesting prior 'LLM passed the Turing test' claims on methodological grounds (those studies used time limits Turing never specified). Note: the published Results section contains one internally inconsistent sentence that appears to transpose the 43%/97% figures between the two games, contradicting both the abstract and the adjacent sentence; this was resolved by checking the paper's own reported odds ratio (42.86), which only reconciles with CIHG=97%/MIWG=43% (the abstract's reading), confirming that one inline sentence carries a labeling error.

How this was mapped onto the axes, and where the paper is silent

Not a fork, and arguably the opposite of one: a strict, literal replication of Turing's 1950 three-party game, run more carefully than prior claimed 'passes', not a redefinition of what is measured. format=three-party and evidence=behavioral are explicit and central ('We conducted Turing's three-player imitation game... following the guidelines identified by Turing'). modality=text and memory=single are explicit in the described setup (single-sitting, message-based game via the OpenAI API; each participant took part in only one trial in a single role). duration=unbounded is explicit and is the paper's central methodological point: the computer-imitates-human game (CIHG) was run 'without duration constraints,' and the authors argue this is precisely why they get a different (negative) result than time-boxed prior studies. interrogator=naive is explicit, not defaulted: the paper quotes Turing's own spec that the interrogator 'should not be expert' and confirms recruited participants were not required to have AI expertise. constraint=none and threshold=30pct are both defaulted — the paper never imposes a compute/energy budget, and never explicitly adopts Turing's 30% figure as its criterion; its own statistical tests are framed against a 50%-chance null, which arguably reads closer to 'chance' than '30pct' — a genuine judgment call, defaulted to the schema base rather than overridden since the paper is silent on a specific number rather than affirmatively adopting a different one. rigor=full is explicit and strongly supported: preregistered on AsPredicted (#148990), IRB-approved (Univ. of Canterbury HREC 2023/98/LR-PS), with exact model (GPT-4-Turbo), API parameters, and the full system prompt disclosed. Title note: the current arXiv title is 'A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence' (lowercase 'a'), which differs from the working title 'The Imitation Game According To Turing' found in earlier secondary coverage — this is the live, current title, recorded verbatim.

A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence · doi 10.48550/arXiv.2501.17629

The Generalized Turing Test (GTT)

2026 verified
2.1.0

Daniel Mitropolsky, Susan S. Hong, Riccardo Neumarker, Emanuele Rimoldi, Tomaso Poggio · arXiv

We introduce the Generalized Turing Test (GTT), a formal framework for comparing the capabilities of arbitrary agents via indistinguishability.
Format: Two-party Threshold: Chance (50%)

Ever run

Yes — a real, substantial empirical section, self-described by the authors as exploratory rather than a rigorous measurement campaign. Nine LLMs tested pairwise: Claude Opus 4.6, Claude Sonnet 4.6, GPT-5.4, Gemini 3.1 Pro Preview, DeepSeek-V3.2, Mistral Large 2512, Ministral 8B 2512, Qwen3 32B, and Grok 4.20. For each of two full-matrix protocols (plain GTT, and GTTQ which adds a querying phase), every one of the 81 ordered actor-distinguisher pairs was run 10 times, giving 810 trial records per protocol — at least 1,620 trials across the two main protocols alone, plus additional fixed-distinguisher experiments, consistent with the abstract's claim of 'thousands of trials.' Headline result (Table 1): ranked by mean aggregate Turing Score T, Gemini 3.1 Pro tops the table (T=0.784, F=0.750, D=0.819), followed by Opus 4.6 (T=0.734), GPT-5.4 (T=0.722, but lopsided: F=0.912, the best fooling/actor score in the table, against a comparatively weak D=0.531 as distinguisher), Sonnet 4.6 (T=0.678), DeepSeek V3.2 (T=0.603), Grok 4.20 (T=0.569), Mistral Large (T=0.478), Qwen3 32B (T=0.450), Ministral 8B lowest (T=0.428). The authors report this ordering is 'broadly consistent with familiar model leaderboards.' A controlled-turn ablation found a strong dose-response effect for a strong distinguisher: 'against Gemini imitating Claude Opus, the distinguisher's success rises from 40% at one turn to 90% for >=3 turns.' No human baseline or significance testing is reported for the main pairwise matrix; treat per-model numbers as descriptive, as the authors themselves caveat.

How this was mapped onto the axes, and where the paper is silent

This is the proposal in the corpus that strains the nine-axis schema hardest, because it removes both fixed roles Turing assumed. Formally: A >= B iff B, acting as 'distinguisher', cannot reliably tell an interaction with a fresh instance of B from an interaction with A instructed to imitate B. Neither A nor B need be human; the paper's own worked example is Gemini 3.1 Pro imitating, and being distinguished by, Claude Opus 4.6. The word 'interrogator' does not appear anywhere in the paper (checked by full-text search) — 'distinguisher' replaces it deliberately. FORMAT: each trial is one distinguisher conversing with one unknown interlocutor and rendering a binary verdict, with no simultaneous side-by-side pair and no separate human judge — structurally closest to 'two-party' (a witness judged in isolation), just with a non-human witness and non-human judge. MODALITY: interactions are explicit multi-turn LLM conversations, so 'text' is not a default. MEMORY: each trial is fresh and bounded, matching 'single'. EVIDENCE: purely behavioral — the distinguisher only sees transcripts, never weights or activations. DURATION genuinely does not map: trials are paced in turns (main GTT capped at 40 distinguisher turns, GTTQ query phase at 20), not minutes; left at the '5min' default for lack of a better option, but a 40-turn LLM exchange is almost certainly far more content than a human 5-minute teleprinter session — the schema needs a turn-count axis to represent this proposal honestly. INTERROGATOR is the single biggest mismatch: the distinguisher in every reported experiment is a model instance, never a human. None of naive/motivated/literate/expert describes 'the judge is itself an AI system'; defaulted to 'naive' only because there is no non-human option. THRESHOLD mapped to 'chance': the pass criterion is Pr[B succeeds] <= 1/2 + epsilon (epsilon=0.005 in the main empirical figure) — the distinguisher must do no better than a coin flip within a small tolerance. A stronger 'statistical indistinguishability' definition is also given but explicitly described as impractical to test and not what the empirical section operationalizes. RIGOR='informal', matching the authors' own characterization: 'this study is intended as a first empirical instantiation of the framework rather than a high-precision measurement campaign,' no preregistration, no significance testing on the main pairwise matrix.

The Generalized Turing Test: A Foundation for Comparing Intelligence · doi 10.48550/arXiv.2605.10851

Forks

These are published as successors to the Turing Test and do not measure conversational indistinguishability. They get no version number, because there is no shared baseline to number them against.

This is not a criticism of the work. Several are better tests than the one they replace. It is a criticism of the naming, which lets a change of subject read as a raised standard.

The Winograd Schema Challenge

2012 verified
fork

Hector J. Levesque, Ernest Davis, Leora Morgenstern · KR-2012 (Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning)

In this paper, we present an alternative to the Turing Test that has some conceptual and practical advantages. A Winograd schema is a pair of sentences that differ only in one or two words and that contain a referential ambiguity that is resolved in opposite directions in the two sentences.

why a fork
Replaces live interrogation entirely with a fixed corpus of binary pronoun-disambiguation questions, administered and graded automatically, with no interrogator, no dialogue, and no deception. The paper's own stated design goal (Sec. 2, 'The trouble with Turing') is to eliminate the conversational, deception-based format, not to make it harder.

The Winograd Schema Challenge

The Lovelace 2.0 Test of Artificial Creativity and Intelligence

2014 verified
fork

Mark O. Riedl · arXiv

Observing that the creation of certain types of artistic artifacts necessitate intelligence, we present the Lovelace 2.0 Test of creativity as an alternative to the Turing Test as a means of determining whether an agent is intelligent. The Lovelace 2.0 Test builds off prior tests of creativity and additionally provides a means of directly comparing the relative intelligence of different agents.

why a fork
Proposes creative-artifact generation under evaluator-specified constraints as a replacement for the Turing Test: a human judge sets constraints for a creative artifact (poem, image, story) and checks whether the output meets them without being a mere fluke. Nothing in the procedure involves a witness being mistaken for human, so it does not measure conversational indistinguishability.

The Lovelace 2.0 Test of Artificial Creativity and Intelligence

Beyond the Turing Test

2016 verified
fork

Gary Marcus, Francesca Rossi, Manuela Veloso · AI Magazine, Vol. 37 No. 1 (Spring 2016), pp. 3-4

The articles in this special issue of AI Magazine include those that propose specific tests, and those that look at the challenges inherent in building robust, valid, and reliable tests for advancing the state of the art in AI.

why a fork
The editorial proposes no single measurable test of its own: it introduces a special issue of separate proposals (visual question answering, embodied construction tasks, standardized-test benchmarks, the Winograd Schema Challenge, etc.) and calls for a plural, undefined 'suite of tests' the full-text PDF dubs the 'Turing Championships,' explicitly framing every contained proposal as a starting point rather than endorsing one procedure.

Beyond the Turing Test

Turing Test Revisited: A Framework for an Alternative

2019 verified
fork

Aladdin Ayesh · arXiv preprint

The paper presents a plausible generic framework based on categories of factors implied by subjective perception of intelligence. An evaluative discussion concludes the paper highlighting some of the unaddressed issues within this generic framework.

why a fork
Proposes replacing indistinguishability testing with a taxonomic '4 E's' framework (Experience, Emergence, Expression, Explanation) for classifying the subjective factors behind perceiving a system as intelligent. The paper itself states the framework 'lacks the instantiation rules and mechanisms to ensure the correct instantiation and application' -- there is no interrogator, conversation, format, or pass/fail criterion of any kind, operational or otherwise.

Turing Test Revisited: A Framework for an Alternative

Twenty Years Beyond the Turing Test: Moving Beyond the Human Judges Too

2020 verified
fork

Jose Hernandez-Orallo · Minds and Machines (Springer)

After recognising the abyss that appears beyond superhuman performance, we build on Turing learning to identify two different evaluation schemas: Turing testing and adversarial testing.

why a fork
A conceptual/position paper proposing two distinct evaluation schemas rather than one operational test: 'Turing testing' keeps imitation/indistinguishability but replaces the human judge with a learned discriminative model (the GAN-style 'Turing learning' paradigm); 'adversarial testing' drops imitation and indistinguishability entirely in favor of a 'challenge-solve-and-replace' benchmark-outpacing dynamic. Since it does not commit to one instantiated format/duration/threshold and half the proposal explicitly abandons indistinguishability, there is no single protocol to place on the nine axes.

Twenty Years Beyond the Turing Test: Moving Beyond the Human Judges Too

TuringBench

2021 verified
fork

Adaku Uchendu, Zeyu Ma, Thai Le, Rui Zhang, Dongwon Lee · Findings of the Association for Computational Linguistics: EMNLP 2021

In this work, we present the TuringBench benchmark environment, which is comprised of (1) a dataset with 200K human- or machine-generated samples across 20 labels {Human, GPT-1, GPT-2_small, GPT-2_medium, GPT-2_large, GPT-2_xl, GPT-2_PyTorch, GPT-3, GROVER_base, GROVER_large, GROVER_mega, CTRL, XLM, XLNET_base, XLNET_large, FAIR_wmt19, FAIR_wmt20, TRANSFORMER_XL, PPLM_distil, PPLM_gpt2}, (2) two benchmark tasks -- i.e., Turing Test (TT) and Authorship Attribution (AA), and (3) a website with leaderboards.

why a fork
The reusable, leaderboarded 'Turing Test (TT)' task the paper actually contributes is a binary classification problem benchmarked with five automatic text classifiers reading static pre-generated news articles -- machine-vs-machine forensic detection, not a human judging live or transcript-based conversation. It keeps 'Turing Test' branding but measures detector F1 on a security/forensics task, not conversational indistinguishability as judged by a human interrogator.

TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation

NeuroAI Turing Test

2025 verified
fork

Jenelle Feather, Meenakshi Khosla, N. Apurva Ratan Murty, Aran Nayebi · arXiv preprint

While behavioral similarity provides a strong starting point, two systems with very different internal representations can produce the same outputs. Thus, in modeling biological intelligence, the field of NeuroAI often aims to go beyond behavioral similarity and achieve representational convergence between a model's activations and the measured activity of a biological system.

why a fork
Replaces conversational interrogation entirely with a statistical comparison of a model's internal neural activations against brain recordings, requiring representational -- not behavioral -- indistinguishability from a brain within the range of natural inter-individual variability. There is no interrogator, no dialogue, and the paper argues behavioral matching alone (Turing's actual test) is 'incomplete' for this purpose.

Brain-Model Evaluations Need the NeuroAI Turing Test

Turing Test 2.0

2025 verified
fork

Georgios Mappouras · arXiv

Second, we present a new framework on how to construct tests that can detect if a system has achieved G.I. in a simple, comprehensive, and clear-cut fail/pass way. We call this novel framework the Turing test 2.0.

why a fork
The actual test procedure has no interrogator distinguishing a human from a machine at all. It defines an information-theoretic threshold and scores a single AI system pass/fail on whether it can autonomously generate genuinely novel information not already recoverable from its own knowledge (e.g. 'generate an image of an analogue clock that shows half past six'). There is no judge comparing a human and machine interlocutor; it is a general/generative-intelligence benchmark administered to one system at a time, despite keeping the 'Turing test' branding.

Turing Test 2.0: The General Intelligence Threshold

GROW-AI (Growth and Realization of Autonomous Wisdom) test

2025 verified
fork

Alexandru Tugui · arXiv

This study aims to extend the framework for assessing artificial intelligence, called GROW-AI (Growth and Realization of Autonomous Wisdom), designed to answer the question "Can machines grow up?" -- a natural successor to the Turing Test. The methodology applied is based on a system of six primary criteria (C1-C6), each assessed through a specific "game", divided into four arenas that explore both the human dimension and its transposition into AI.

why a fork
Replaces conversational indistinguishability entirely with a six-criteria maturity/autonomy rubric (autonomous growth, entropy/gravity understanding, algorithm efficiency, sensory/affective logic, self-evaluation, advanced autonomous wisdom), each scored via a separate 'game' by human expert evaluators and combined into a composite 'Grow Up Index.' The paper frames its goal as going 'beyond simple imitation, by focusing on autonomy, responsibility, and ethical maturation' -- the textbook fork case the brief describes.

The next question after Turing's question: Introducing the Grow-AI test