Reinforcement learning controllers for robot manipulators depend strongly on reward tuning, and fixed weights may yield poor trade-offs under uncertainty and disturbances. This paper proposes a disturbance observer-based actor–critic RL (DOB–ACRL) with adaptive multi-objective reward shaping for a torque-saturated 2-DOF manipulator, where the reward weights are updated online using normalized indicators of tracking error, control energy, and effort. A Lyapunov analysis guarantees the uniform ultimate boundedness of closed-loop signals. The simulations show improved learning and performance over a static reward actor–critic baseline, reducing the RMS tracking error by up to 22.8%, the control energy by ~4.6%, the control effort by 1.9%, and the settling time by up to 29.2%.
Tam et al. (2026) studied this question.