Suppose the arms of a two-armed bandit generate i.i.d. Bernoulli random variables with success probabilities ρ and λ respectively. It is desired to maximize the expected sum of N trials where N is fixed. If the prior distribution of (ρ, λ) is concentrated at two points $(a, b)$ and $(c, d)$ in the unit square, a characterization of the optimal policy is given. In terms of $a, b, c$, and d, necessary and sufficient conditions are given for the optimality of the myopic policy.
No takes yet. Share an insight, caveat, or question.
Thomas A. Kelley (1974) studied this question.