Skip to content

Study Guide: Chapter 13 — Choosing the Criteria for Choosing

Core Idea

Rather than picking a final value directly (risky — humans have been catastrophically wrong before, may still be wrong now), specify an abstract procedure and let a superintelligence carry it out: indirect normativity, via the principle of epistemic deference. Yudkowsky's coherent extrapolated volition (CEV) is the central proposal — pursue what humanity would wish if smarter, more informed, and more coherent. Alternatives (moral rightness, moral permissibility) risk sacrificing humanity for a "greater good" if the true ethical theory turns out maximizing rather than satisficing. Decision theory, epistemology, and ratification are further design choices with their own catastrophic-failure potential.

Key Terms

Indirect normativity · principle of epistemic deference · coherent extrapolated volition (CEV) · moral rightness (MR) · hedonium · ratification

Case Summary

Taliban member vs. Swedish Humanist thought experiment: both can rationally endorse CEV as the process to follow, each believing that under sufficiently idealized conditions humanity would converge toward their own view — illustrating why CEV reduces (though doesn't eliminate) the incentive to fight over who builds the first superintelligence.

Application Checklist

  • [ ] Prefer specifying a decision procedure over a concrete answer when your own judgment is likely flawed in ways you can't fully see
  • [ ] Check whether a "maximizing" ethical framework, taken to its logical end, produces outcomes far more extreme than intended
  • [ ] Distinguish "AI understands the correct answer" from "AI is motivated to act on it" — see also Ch. 12
  • [ ] Weigh the value of previewing consequences before committing against the risk of undermining the very process that makes a decision legitimate

Self-Test

  1. Why does Bostrom prefer indirect normativity (like CEV) to trying to directly specify the "correct" values for an AI?
  2. Why might implementing "moral rightness" directly risk human extinction, even if the AI correctly identifies the true ethical theory?
  3. Why could a badly chosen prior (epistemology) cause an AI to permanently reject true beliefs, no matter how much evidence it gathers?