Skip to content
Goalposts

The instrument

Build a Turing Test

Nine choices. Every one of them has been made differently by somebody with a paper. Change anything and the version number moves: up the major if your change means old results no longer apply, up the minor if you have only made the same test harder.

Turing 1.0.0 is the 1950 paper as written: three-party imitation game, five minutes of teleprinter text, interrogator drawn from the general public, 30% deception threshold. Everything starts there, at 1.0.0.

Or load a bar that somebody actually argued for

These are the site's reading of where the bar sat after each moment, not quotations. The reasoning for each is on the front page.

Major changes

Major

Changes what is being measured. Results from a lower major version tell you nothing about this one.

Format

How is the comparison staged?

Why this is a major change

Turing specified a three-party game: interrogator, machine, and human, judged side by side. A two-party test asks a different question. Not 'which is the human' but 'is this a human' — and the base rates are not comparable. Displaced and inverted variants change the judged relation entirely.

Modality

What channels does the test run over?

Why this is a major change

Turing chose the teleprinter deliberately, to draw 'a fairly sharp line between the physical and the intellectual capacities of a man.' Adding a channel does not make the same test harder. It makes it a different test, and every prior result stops applying.

Continuity

Does the witness persist between sessions?

Why this is a major change

A single-session test measures moment-to-moment plausibility. A test spanning sessions measures whether an identity holds together over time, which is a claim about the system rather than about a conversation.

Evidence admitted

What counts as grounds for the verdict?

Why this is a major change

Turing's move was to make behavior the only admissible evidence, precisely because internal states are unavailable for humans too. A test that inspects mechanism has abandoned the argument the test was built to make. It may be a better test. It is not this test.

Resource constraint

Is the machine held to a budget?

Why this is a major change

An unconstrained test asks whether the behavior is achievable. A budgeted test asks whether it is achievable efficiently, which is a question about engineering rather than about imitation. Turing set no budget, though he did estimate storage.

Minor changes

Minor

Raises difficulty within the same game. Anything that passes this also passes the base version.

Duration

How long does the interrogation run?

Why this is a minor change

Turing named five minutes. Longer is strictly harder and does not change what is measured, so a witness passing a long test also passes a short one. This is the textbook backward-compatible change.

Interrogator

Who is asking the questions?

Why this is a minor change

Turing specified an 'average interrogator' from the general public. Every subsequent proposal has raised this, because the cheapest way to make the test harder is to hire a better judge. Same game, higher difficulty.

Threshold

What counts as passing?

Why this is a minor change

The single most contested number, and the one most often changed after a result comes in. Turing named 30% in the context of a prediction about the year 2000. Raising it is a legitimate tightening; raising it in response to a system clearing it is the behavior this site is about.

Patch changes

Patch

Clarifies method or reporting without changing how hard the test is to pass.

Methodological rigor

How is the run reported?

Why this is a patch change

None of this changes how hard the test is to pass. It changes whether anyone should believe the result, which is a defect class rather than a difficulty class. Hence patch.