Bandit algorithms instead of rigid A/B tests
Summary
Steve Hanov shows how a Bayesian bandit continuously directs traffic to promising variants. The example algorithm needs only a few lines of code. Beta distributions model binary success probabilities.
Ideas
- Thompson sampling balances exploration and exploitation probabilistically.
- Every observation updates the success distribution of a variant.
- Weaker variants receive less and less real traffic.
- The method maximises yield while the experiment is still running.
Insights
- Optimisation and unbiased insight pursue different goals.
- Adaptive experiments react faster but make classic evaluation harder.
- Uncertainty should visibly influence decisions instead of remaining hidden.
Facts
- Every variant is still tried out occasionally.
Recommendations
- Define the target metric and stopping criterion before the experiment.
- Simulate the algorithm with known probabilities.
References
Links to the original source and the Web Archive open in a new tab.