bk99.de entertain the web since 1997

Bandit algorithms instead of rigid A/B tests

Summary

Steve Hanov shows how a Bayesian bandit continuously directs traffic to promising variants. The example algorithm needs only a few lines of code. Beta distributions model binary success probabilities.

Ideas

  • Thompson sampling balances exploration and exploitation probabilistically.
  • Every observation updates the success distribution of a variant.
  • Weaker variants receive less and less real traffic.
  • The method maximises yield while the experiment is still running.

Insights

  • Optimisation and unbiased insight pursue different goals.
  • Adaptive experiments react faster but make classic evaluation harder.
  • Uncertainty should visibly influence decisions instead of remaining hidden.

Facts

  • Every variant is still tried out occasionally.

Recommendations

  • Define the target metric and stopping criterion before the experiment.
  • Simulate the algorithm with known probabilities.

References

Read the original article

Search the Web Archive