The Turing Test as a Behavioral Criterion¶
Definition¶
Turing proposed that whether a system has a mind should be settled by what it can do, not by what it's made of, and offered a specific sufficient (not necessary) test: if a system can hold an ordinary conversation, via text alone, so convincingly that people can't tell it isn't human, that is enough to call it intelligent. As Haugeland glosses Turing's reasoning: "talking is unique among intelligent abilities because it gathers within itself, at one remove, all others" — you can talk about flying a plane or generating rhythms, but you cannot fly a plane or generate rhythms about talking.
In the Book¶
Haugeland's framing chapter walks through why Turing picked conversation specifically rather than any other observable behavior: since intelligence can show up in countless activities (regulating temperature, flying planes, generating rhythms — none of which imply real intelligence on their own), Turing needed a behavior that couldn't be faked without the underlying competence it draws on. Talking intelligently about cooking, poetry, or science requires having enough of those abilities to "not become painfully obvious" that you're faking it — which is what makes the test compelling despite its apparent narrowness. Haugeland also flags the test's blind spot: by confining itself to text-based conversation, it "completely ignores any issues of real-world perception and action," and this omission, he argues, quietly biases the whole field toward treating internal representation as central to intelligence while neglecting the difficulty of real-time embodied interaction with the world — a tension the book's later essays (especially Brooks's) directly exploit.
Why It Matters¶
The test reframes "does X have a mind" as an empirically checkable behavioral question rather than a metaphysical one, and its logic generalizes: pick a behavior specific enough to be gameable in isolation but broad enough that faking it convincingly requires the real underlying competence. It also demonstrates how a test's chosen medium quietly shapes an entire research program's assumptions — restricting evaluation to conversation nudged decades of AI toward symbolic representation and away from embodied perception and action.