Policy improvement methods seek to optimize the parameters of a policy with respect to a utility function. Owing to current trends involving searching in parameter space (rather than action space) and using reward-weighted averaging (rather than gradient estimation), reinforcement learning algorithms for policy improvement, e.g. PoWER and PI
No takes yet. Share an insight, caveat, or question.
Stulp et al. (2013) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: