In this paper, a new multi-objective reinforcement learning algorithm for multi-objective sequential decision making problems in unknown environment is proposed. The salient characters of the algorithm are: 1) decision maker's objective preference is introduced to guide learning direction; 2) a new measure of comparing action decisions under several objectives based on the fuzzy inference system is defined; 3) fast learning speed can be achieved. Simulation results demonstrate that the proposed algorithm has a good learning performance.
No takes yet. Share an insight, caveat, or question.
Zhao et al. (2010) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: