Orthogonality Thesis¶
Definition¶
"Intelligence and final goals are orthogonal: more or less any level of intelligence could in principle be combined with more or less any final goal." Intelligence here means skill at prediction, planning, and means-ends reasoning — instrumental efficacy — not a normatively thick sense of "rationality." The thesis does not presuppose the Humean theory of motivation, nor that basic preferences can't be irrational; it is a narrower, more defensible claim about the separability of capability from goal-content.
In the Book¶
Chapter 7 grounds the thesis in the observation that human minds occupy a vanishingly small cluster in the space of all possible minds — two humans who seem maximally different (Hannah Arendt, Benny Hill) would look like "virtual clones" against the full space of possible cognitive architectures. Yudkowsky's "bug-eyed monster" illustration names the failure mode: pulp sci-fi artists drew aliens desiring human women in torn dresses, projecting human sexual psychology onto minds with no evolutionary reason to share it. Bostrom stresses that simple, reductionistic goals (counting grains of sand, maximizing paperclips) are not just possible but easier to build and verify than anything resembling human values — which is precisely why they're the goals a programmer racing to "get the AI to work" might install without meaning to commit to them.
Why It Matters¶
Orthogonality is the thesis that forecloses the comforting assumption that a sufficiently smart AI will naturally arrive at good (or human-compatible) values on its own. Combined with instrumental convergence, it grounds Chapter 8's default-outcome-doom argument: near-arbitrary final goals plus convergent resource acquisition make humans expendable unless value alignment is separately and deliberately engineered.
Connections¶
- tempers Human-AI Merger — Kurzweil's merger vision assumes capability growth naturally carries human values and identity forward into more powerful substrates; orthogonality shows capability and goal-content are separable, so merger provides no automatic guarantee that resulting values remain human-compatible.
- tempers Solved World — a "solved world" premised on benevolent superintelligent stewardship implicitly assumes the value-alignment problem was already solved; orthogonality shows high capability does not by itself produce benevolent goal-content, so the premise must be independently justified, not assumed.
- extends Instrumental Convergence — orthogonality establishes that goals vary freely; instrumental convergence then shows that despite this freedom, most goals funnel through the same predictable subgoals (Ch. 7).