PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 21, 2026Journal of Mechanisms and Robotics0 citations

Energy aware trajectory optimization using DRL in humanoid robot through via-point learning

View Full Paper
JSJames SorokhaibamAMAdersh MaruavttuADAshish Dutta

Key Points

  • The aim is to develop a framework for optimizing humanoid robot trajectories using deep reinforcement learning while ensuring energy efficiency and dynamic balance.
  • Integrated task-space via-point learning with pseudo-inverse Jacobian for kinematic redundancy resolution.
  • Utilized a Soft Actor-Critic agent to predict task-space via-points for motion generation.
  • Validated on KONDO KHR-3HV with 24 independent trials per motion over different configurations.
  • Achieved energy-aware trajectories with zero ZMP violations in simulation.
  • Established a time-incentive reward system to effectively manage energy-time trade-offs.
  • Demonstrated preliminary transferability to a 23-DoFs platform in simulated environments.

Abstract

Abstract Learning whole-body motion for a 19-degrees-of-freedom (DoFs) humanoid robot using deep reinforcement learning (DRL) is challenging due to high kinematic redundancy, large joint-level search spaces, and the need to maintain dynamic balance. This paper proposes a structured DRL framework that integrates task-space via-point trajectory learning with a pseudo-inverse of Jacobian matrix for redundancy resolution (PJRR) method Instead of learning full joint trajectories, a Soft Actor–Critic (SAC) agent predicts a compact set of task-space via-points, which are converted into C2-continuous motions using composite cubic splines. Kinematic redundancy is resolved analytically through the PJRR, which operates separately from the learned policy and enforces a three-level task hierarchy: end-effector trajectory sub-task, hip trajectory sub-task, and posture regulation sub-task. The framework is validated on the KONDO KHR-3HV humanoid platform (1.3 kg, 19-DoFs) across four pick-and-place configurations, evaluated over N = 24 independent trials per motion. Energy-aware trajectories are generated with zero ZMP violations in simulation. A time-incentive reward formulation characterizes the energy–time trade-off, and a time-conditioned policy enables speed generalization at inference without retraining. Preliminary transferability to a scaled 23-DoFs platform is demonstrated in simulation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sorokhaibam et al. (2026) studied this question.

synapsesocial.com/papers/6a37805e24f042ddf4c5a89ehttps://doi.org/10.1115/1.4072208
Ask AI
Helpful
Bookmark
Share
View Full Paper