The Winograd Schema Challenge
2012 verified fork Hector J. Levesque, Ernest Davis, Leora Morgenstern · KR-2012 (Proceedings of the Thirteenth International Conference on Principles of Knowledge Representation and Reasoning)
In this paper, we present an alternative to the Turing Test that has some conceptual and practical advantages. A Winograd schema is a pair of sentences that differ only in one or two words and that contain a referential ambiguity that is resolved in opposite directions in the two sentences.
why a fork
Replaces live interrogation entirely with a fixed corpus of binary pronoun-disambiguation questions, administered and graded automatically, with no interrogator, no dialogue, and no deception. The paper's own stated design goal (Sec. 2, 'The trouble with Turing') is to eliminate the conversational, deception-based format, not to make it harder.
The Winograd Schema Challenge
The Lovelace 2.0 Test of Artificial Creativity and Intelligence
2014 verified fork Mark O. Riedl · arXiv
Observing that the creation of certain types of artistic artifacts necessitate intelligence, we present the Lovelace 2.0 Test of creativity as an alternative to the Turing Test as a means of determining whether an agent is intelligent. The Lovelace 2.0 Test builds off prior tests of creativity and additionally provides a means of directly comparing the relative intelligence of different agents.
why a fork
Proposes creative-artifact generation under evaluator-specified constraints as a replacement for the Turing Test: a human judge sets constraints for a creative artifact (poem, image, story) and checks whether the output meets them without being a mere fluke. Nothing in the procedure involves a witness being mistaken for human, so it does not measure conversational indistinguishability.
The Lovelace 2.0 Test of Artificial Creativity and Intelligence
Beyond the Turing Test
2016 verified fork Gary Marcus, Francesca Rossi, Manuela Veloso · AI Magazine, Vol. 37 No. 1 (Spring 2016), pp. 3-4
The articles in this special issue of AI Magazine include those that propose specific tests, and those that look at the challenges inherent in building robust, valid, and reliable tests for advancing the state of the art in AI.
why a fork
The editorial proposes no single measurable test of its own: it introduces a special issue of separate proposals (visual question answering, embodied construction tasks, standardized-test benchmarks, the Winograd Schema Challenge, etc.) and calls for a plural, undefined 'suite of tests' the full-text PDF dubs the 'Turing Championships,' explicitly framing every contained proposal as a starting point rather than endorsing one procedure.
Beyond the Turing Test
Turing Test Revisited: A Framework for an Alternative
2019 verified fork Aladdin Ayesh · arXiv preprint
The paper presents a plausible generic framework based on categories of factors implied by subjective perception of intelligence. An evaluative discussion concludes the paper highlighting some of the unaddressed issues within this generic framework.
why a fork
Proposes replacing indistinguishability testing with a taxonomic '4 E's' framework (Experience, Emergence, Expression, Explanation) for classifying the subjective factors behind perceiving a system as intelligent. The paper itself states the framework 'lacks the instantiation rules and mechanisms to ensure the correct instantiation and application' -- there is no interrogator, conversation, format, or pass/fail criterion of any kind, operational or otherwise.
Turing Test Revisited: A Framework for an Alternative
Twenty Years Beyond the Turing Test: Moving Beyond the Human Judges Too
2020 verified fork Jose Hernandez-Orallo · Minds and Machines (Springer)
After recognising the abyss that appears beyond superhuman performance, we build on Turing learning to identify two different evaluation schemas: Turing testing and adversarial testing.
why a fork
A conceptual/position paper proposing two distinct evaluation schemas rather than one operational test: 'Turing testing' keeps imitation/indistinguishability but replaces the human judge with a learned discriminative model (the GAN-style 'Turing learning' paradigm); 'adversarial testing' drops imitation and indistinguishability entirely in favor of a 'challenge-solve-and-replace' benchmark-outpacing dynamic. Since it does not commit to one instantiated format/duration/threshold and half the proposal explicitly abandons indistinguishability, there is no single protocol to place on the nine axes.
Twenty Years Beyond the Turing Test: Moving Beyond the Human Judges Too
TuringBench
2021 verified fork Adaku Uchendu, Zeyu Ma, Thai Le, Rui Zhang, Dongwon Lee · Findings of the Association for Computational Linguistics: EMNLP 2021
In this work, we present the TuringBench benchmark environment, which is comprised of (1) a dataset with 200K human- or machine-generated samples across 20 labels {Human, GPT-1, GPT-2_small, GPT-2_medium, GPT-2_large, GPT-2_xl, GPT-2_PyTorch, GPT-3, GROVER_base, GROVER_large, GROVER_mega, CTRL, XLM, XLNET_base, XLNET_large, FAIR_wmt19, FAIR_wmt20, TRANSFORMER_XL, PPLM_distil, PPLM_gpt2}, (2) two benchmark tasks -- i.e., Turing Test (TT) and Authorship Attribution (AA), and (3) a website with leaderboards.
why a fork
The reusable, leaderboarded 'Turing Test (TT)' task the paper actually contributes is a binary classification problem benchmarked with five automatic text classifiers reading static pre-generated news articles -- machine-vs-machine forensic detection, not a human judging live or transcript-based conversation. It keeps 'Turing Test' branding but measures detector F1 on a security/forensics task, not conversational indistinguishability as judged by a human interrogator.
TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation
NeuroAI Turing Test
2025 verified fork Jenelle Feather, Meenakshi Khosla, N. Apurva Ratan Murty, Aran Nayebi · arXiv preprint
While behavioral similarity provides a strong starting point, two systems with very different internal representations can produce the same outputs. Thus, in modeling biological intelligence, the field of NeuroAI often aims to go beyond behavioral similarity and achieve representational convergence between a model's activations and the measured activity of a biological system.
why a fork
Replaces conversational interrogation entirely with a statistical comparison of a model's internal neural activations against brain recordings, requiring representational -- not behavioral -- indistinguishability from a brain within the range of natural inter-individual variability. There is no interrogator, no dialogue, and the paper argues behavioral matching alone (Turing's actual test) is 'incomplete' for this purpose.
Brain-Model Evaluations Need the NeuroAI Turing Test
Turing Test 2.0
2025 verified fork Georgios Mappouras · arXiv
Second, we present a new framework on how to construct tests that can detect if a system has achieved G.I. in a simple, comprehensive, and clear-cut fail/pass way. We call this novel framework the Turing test 2.0.
why a fork
The actual test procedure has no interrogator distinguishing a human from a machine at all. It defines an information-theoretic threshold and scores a single AI system pass/fail on whether it can autonomously generate genuinely novel information not already recoverable from its own knowledge (e.g. 'generate an image of an analogue clock that shows half past six'). There is no judge comparing a human and machine interlocutor; it is a general/generative-intelligence benchmark administered to one system at a time, despite keeping the 'Turing test' branding.
Turing Test 2.0: The General Intelligence Threshold
GROW-AI (Growth and Realization of Autonomous Wisdom) test
2025 verified fork Alexandru Tugui · arXiv
This study aims to extend the framework for assessing artificial intelligence, called GROW-AI (Growth and Realization of Autonomous Wisdom), designed to answer the question "Can machines grow up?" -- a natural successor to the Turing Test. The methodology applied is based on a system of six primary criteria (C1-C6), each assessed through a specific "game", divided into four arenas that explore both the human dimension and its transposition into AI.
why a fork
Replaces conversational indistinguishability entirely with a six-criteria maturity/autonomy rubric (autonomous growth, entropy/gravity understanding, algorithm efficiency, sensory/affective logic, self-evaluation, advanced autonomous wisdom), each scored via a separate 'game' by human expert evaluators and combined into a composite 'Grow Up Index.' The paper frames its goal as going 'beyond simple imitation, by focusing on autonomy, responsibility, and ethical maturation' -- the textbook fork case the brief describes.
The next question after Turing's question: Introducing the Grow-AI test