The Value-Loading Problem¶
Definition¶
The problem of getting a value like "happiness" or "justice" — not a primitive in any programming language — into an AI's utility function, in a way robust enough to survive the agent's growth into superintelligence. It must be solved before the system is smart enough to resist correction, since goal-content integrity (Ch. 7) gives a sufficiently advanced agent convergent reason to treat any attempted fix as an attack to repel.
In the Book¶
Chapter 12 surveys six approaches — evolutionary selection (inherits mind-crime and blind-search risk), reinforcement learning (collapses to wireheading), associative value accretion (mimics how humans get values, but the mechanism is evolutionarily bespoke and poorly understood even in us), motivational scaffolding (risky, since goal-content integrity resists the planned swap), value learning (the most promising: fix the goal as "pursue the true values," let beliefs about content refine — illustrated by the "envelope and barge" thought experiment), and institution design (shape a composite system's motivation through internal organizational structure rather than any single agent's content). Chapter 13 takes up the further problem this leaves open even if value-loading is solved: which values to load, proposing indirect normativity and coherent extrapolated volition as ways to offload that choice to the superintelligence itself rather than gambling on humanity's present, likely-flawed moral convictions.
Why It Matters¶
Solving value-loading without solving the "which values" question just produces a very effective machine for realizing whatever was specified — including catastrophically wrong specifications. The two problems are conceptually separable but both must be solved for a good outcome; getting either wrong alone can be an existential catastrophe.
Connections¶
- parallels Mental Models — Ch. 12's observation that an AI's actual behavior is driven by whatever is encoded in its motivation system, not by what its designers meant, parallels how an organization's implicit mental models — not its stated mission — actually drive behavior.
- parallels Purpose Problem — the difficulty of specifying what genuinely worthwhile values an AI should pursue (Ch. 13's indirect normativity, precisely because we may not know what we truly want) mirrors Deep Utopia's symmetric difficulty: once instrumental problems are solved for humans, the challenge of identifying what remains worth pursuing turns out to be just as hard to pin down directly.
- extends Orthogonality Thesis — value-loading is the practical engineering problem that orthogonality (Ch. 7) makes unavoidable: because capability doesn't imply good goals, goals must be deliberately and correctly installed.