This paper deals with reinforcement learning for process modeling and control using a model-free, action- dependent adaptive critic (ADAC). A new modified recursive Levenberg Marquardt (RLM) training algorithm, called temporal difference RLM, is developed to improve the ADAC performance. Novel application results for a simulated continuously-stirred-tank-reactor process are included to show the superiority of the new algorithm to conventional temporal-difference stochastic backpropagation.
No takes yet. Share an insight, caveat, or question.
Govindhasamy et al. (2005) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: