[MINI] Multi-armed Bandit Problems episode artwork

EPISODE · Oct 2, 2015 · 12 MIN

[MINI] Multi-armed Bandit Problems

from Data Skeptic

The multi-armed bandit problem is named with reference to slot machines (one armed bandits). Given the chance to play from a pool of slot machines, all with unknown payout frequencies, how can you maximize your reward? If you knew in advance which machine was best, you would play exclusively that machine. Any strategy less than this will, on average, earn less payout, and the difference can be called the "regret". You can try each slot machine to learn about it, which we refer to as exploration. When you've spent enough time to be convinced you've identified the best machine, you can then double down and exploit that knowledge. But how do you best balance exploration and exploitation to minimize the regret of your play? This mini-episode explores a few examples including restaurant selection and A/B testing to discuss the nature of this problem. In the end we touch briefly on Thompson sampling as a solution.

Episode metadata supplied by the publisher feed · Published Oct 2, 2015

Embed this episode

Ready to play

[MINI] Multi-armed Bandit Problems

0:00 12:47

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Data Skeptic?

This episode is 12 minutes long.

When was this Data Skeptic episode published?

This episode was published on October 2, 2015.

Can I download this Data Skeptic episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!