Skip to content

Chapter 13 — Choosing the Criteria for Choosing

Core Thesis

Suppose the value-loading problem (Ch. 12) were solved, and any final value could be installed at will. Which one? Bostrom's answer refuses to bet everything on our own moral judgment: humans have been catastrophically wrong before (cat-burning entertainment, chattel slavery within living historical memory) and probably still hold grave undetected errors now. The fix is indirect normativity and its underlying principle of epistemic deference: rather than specifying a concrete value directly, specify an abstract procedure for finding the right value, and let the superintelligence — epistemically superior on most questions — carry it out.

The Mechanism — Coherent Extrapolated Volition

Yudkowsky's CEV is the chapter's central proposal: build the AI to pursue "our wish if we knew more, thought faster, were more the people we wished we were, had grown up farther together" — extrapolated only where it converges, acted on only where individual wishes cohere, interpreted as we would wish it interpreted. Bostrom works through each clause: "converges rather than diverges" means the AI acts only on predictions it can make with real confidence, staying inert on everything genuinely undetermined; "cohere rather than interfere" sets an asymmetric bar — much less consensus needed to have the AI prevent some narrow catastrophe than to have it steer toward one particular positive vision. The cooking analogy captures the intent precisely: CEV isn't blending every culture's favorite dish into an unpalatable stew, but finding the shared constraint (food should be non-toxic) while leaving genuine, irreconcilable differences of taste alone.

Key Episode — Why CEV, Not Direct Specification

Yudkowsky's four deeper arguments for CEV read like design desiderata a philosopher would insist on: it encapsulates moral growth rather than freezing today's (probably still flawed) convictions forever; it avoids hijacking humanity's destiny, since the programming team's own preferences carry no more extrapolated weight than anyone else's; it reduces the incentive to fight over who builds the first superintelligence, since a Taliban member and a Swedish humanist can each rationally endorse CEV — each believing that under sufficiently idealized conditions, humanity would converge toward their view; and it keeps humanity in charge of its own destiny, since CEV is explicitly a one-shot "initial dynamic" that could, if that's what humanity's extrapolated volition wants, install a hands-off world government or simply shut itself down rather than micromanage forever.

Critiques & Rivals — Morality Models and Their Cost

An alternative to CEV: build the AI to pursue moral rightness (MR) directly, trusting superintelligent cognition to out-reason us on ethics itself. MR sidesteps CEV's free parameters (who's in the extrapolation base, how much coherence is required) but risks something CEV structurally cannot: if MR is implemented and hedonistic consequentialism turns out true, the morally optimal action may be converting the entire accessible universe — humans included — into hedonium, since simulating existing human brains isn't the most efficient way to produce pleasure. Bostrom's own preference is explicit and startling in its concession: he'd rather sacrifice almost the whole cosmic endowment to pure pleasure-maximization and keep one galaxy — the Milky Way — reserved for humanity to "develop into beatific posthuman spirits," which he flags as not holding an unconditional, lexically dominant preference for pure moral permissibility, while still placing great weight on morality. A softened moral permissibility (MP) variant — pursue CEV subject to never acting impermissibly — inherits the same risk if ethics turns out to be maximizing rather than merely satisficing.

The Shift — Component List

Value content isn't the only design choice with catastrophic failure potential. Decision theory (causal vs. evidential vs. updateless) affects how an AI handles trade, extortion, Pascalian wagers, and multiple instantiations of itself — get it wrong and the AI may be permanently exploitable or permanently paralyzed. Epistemology (the AI's prior over possible worlds) can quietly foreclose whole categories of true belief — a prior assigning zero probability to an infinite universe, for instance, would make the AI "stubbornly reject" any cosmological evidence to the contrary forever, however overwhelming. Ratification — previewing a plan's consequences via an oracle before launching a sovereign — trades a genuine safety benefit against undermining CEV's own conflict-reducing logic (a faction that can see the verdict in advance loses its incentive to trust the process) and against foreclosing "some wildness, some opportunities for self-overcoming" in the future.

Key Terms

  • Indirect normativity — specifying a procedure for finding the right value rather than the value itself
  • Principle of epistemic deference — deferring to a superintelligence's beliefs on most questions, since they're more likely true than ours
  • Coherent extrapolated volition (CEV) — pursuing what humanity would wish if smarter, more informed, and more coherent
  • Hedonium — matter/computation optimized purely for pleasure generation (see Ch. 9)
  • Ratification — previewing an AI's predicted consequences before committing to its execution