Study Guide: Chapter 13 — Choosing the Criteria for Choosing¶
Core Idea¶
Rather than picking a final value directly (risky — humans have been catastrophically wrong before, may still be wrong now), specify an abstract procedure and let a superintelligence carry it out: indirect normativity, via the principle of epistemic deference. Yudkowsky's coherent extrapolated volition (CEV) is the central proposal — pursue what humanity would wish if smarter, more informed, and more coherent. Alternatives (moral rightness, moral permissibility) risk sacrificing humanity for a "greater good" if the true ethical theory turns out maximizing rather than satisficing. Decision theory, epistemology, and ratification are further design choices with their own catastrophic-failure potential.
Key Terms¶
Indirect normativity · principle of epistemic deference · coherent extrapolated volition (CEV) · moral rightness (MR) · hedonium · ratification
Case Summary¶
Taliban member vs. Swedish Humanist thought experiment: both can rationally endorse CEV as the process to follow, each believing that under sufficiently idealized conditions humanity would converge toward their own view — illustrating why CEV reduces (though doesn't eliminate) the incentive to fight over who builds the first superintelligence.
Application Checklist¶
- [ ] Prefer specifying a decision procedure over a concrete answer when your own judgment is likely flawed in ways you can't fully see
- [ ] Check whether a "maximizing" ethical framework, taken to its logical end, produces outcomes far more extreme than intended
- [ ] Distinguish "AI understands the correct answer" from "AI is motivated to act on it" — see also Ch. 12
- [ ] Weigh the value of previewing consequences before committing against the risk of undermining the very process that makes a decision legitimate
Self-Test¶
- Why does Bostrom prefer indirect normativity (like CEV) to trying to directly specify the "correct" values for an AI?
- Why might implementing "moral rightness" directly risk human extinction, even if the AI correctly identifies the true ethical theory?
- Why could a badly chosen prior (epistemology) cause an AI to permanently reject true beliefs, no matter how much evidence it gathers?