Optimal control of a linear process with unknown parameters is considered when the horizon is infinite and rewards are discounted. Active learning strategies are considered, i.e., agents consider the information value of possible actions, as well as current reward. Distributional assumptions are minimal in that no restriction to conjugate families is made. Convergence of beliefs and actions is established.
No takes yet. Share an insight, caveat, or question.
Kiefer et al. (1989) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: