Skip to content

Exploration-Exploitation Tradeoff

Definition

In a multi-armed bandit problem, a decision-maker repeatedly chooses among alternatives whose payoffs are drawn from unknown probability distributions rather than fixed known values. Every choice forces a tradeoff between exploration — trying a less-tested alternative to learn more about its true payoff — and exploitation — choosing whichever alternative has performed best so far. Because information is only revealed by trying an option, over-exploiting locks in a possibly-inferior choice while over-exploring wastes trials on options already known to be worse.

In the Book

Page grounds the model (Chapter 27) in Bernoulli bandit problems, each alternative equivalent to an urn with an unknown proportion of successes, illustrated with a chimney-sweeping company testing three sales pitches by phone and having to decide, after ten calls with mixed results, whether to keep exploiting the best-performing pitch or return to the others for more data. He compares two heuristics: "sample-then-greedy," which tests every alternative a fixed number of times before locking onto the best, against an adaptive heuristic that shifts an increasing share of trials toward whichever alternative is currently winning, in proportion to its measured success rate. The chapter's sharpest case is medical and ethical rather than commercial: when Robert Bartlett tested an artificial lung (ECMO) against existing alternatives and its success rate far surpassed the others, continuing to randomize patients to the inferior alternatives for the sake of a cleaner experiment would have caused unnecessary deaths — Bartlett stopped the trial and gave everyone the artificial lung, which Page notes is provably optimal: if an alternative has always succeeded, no exploration can improve on continuing to choose it.

Why It Matters

This tradeoff shows up wherever payoffs are learned rather than known in advance — clinical trials, ad placement, hiring, career choice, R&D portfolios — and it reframes "stick with what works" and "keep experimenting" not as a personality difference but as a genuine mathematical tension whose right balance shifts with how much uncertainty remains and how much of the payoff difference has already been revealed. It also supplies an ethical argument, via the ECMO case, for when continued experimentation stops being scientifically justified and starts being harmful.