PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 9, 2026Battery energy0 citationsOpen Access

Optimal Control of Mobile Energy Storage via Knowledge‐Guided Deep Reinforcement Learning

View Full Paper
XCXinlei CaiZMZijie MengQGQian Guo

Key Points

  • The aim is to develop a control strategy for mobile battery energy storage systems that maximizes profit under uncertain conditions.
  • Developed a deep reinforcement learning framework for mobile battery energy storage systems.
  • Introduced the Knowledge-Assisted Deep Deterministic Policy Gradient (KA-DDPG) algorithm to enhance learning efficiency.
  • Implemented a two-phase guidance strategy transitioning from offline to real-time actions.
  • KA-DDPG achieved a 3%–7% improvement in average profits over the Soft Actor-Critic baseline.
  • Demonstrated over 60% variance reduction in policy stability compared to standard DRL baselines.
  • Validated quicker learning phases under high uncertainty.

Abstract

ABSTRACT While mobile battery energy storage systems (MBESSs) are typically used to improve the stability of power systems, their ability to move also creates good opportunities for businesses to earn money through energy arbitrage. This profit depends heavily on decisions about timing and location, and is affected by uncertain conditions like fluctuating electricity prices and traffic. However, finding the best real‐time control strategy that considers long‐term profit and these uncertainties requires significant computing power. To tackle this issue, this paper presents a deep reinforcement learning framework for MBESSs designed to get the most profit from market arbitrage. Within this framework, we introduce the Knowledge‐Assisted Deep Deterministic Policy Gradient (KA‐DDPG) algorithm to learn the best policy more efficiently. The core novelty of KA‐DDPG lies in its probabilistic hybrid action selection mechanism that unifies the agent's learned policy, offline expert criteria, and random exploration to manage the complex hybrid action space. Additionally, a two‐phase guidance strategy is implemented to transition from offline‐based to real‐time‐based criteria actions, ensuring both learning acceleration and policy robustness under computational constraints. Our rigorous statistical evaluations demonstrate that the proposed KA‐DDPG approach leads to a 3%–7% improvement in average profits over the state‐of‐the‐art Soft Actor‐Critic baseline. Furthermore, it achieves exceptional policy stability, exhibiting a variance reduction of over 60% compared to standard DRL baselines and over 92% compared to the deterministic closed‐loop MPC. The KA‐DDPG algorithm also substantially speeds up the learning phase, validating its efficacy for real‐time MBESS control under high uncertainty.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cai et al. (2026) studied this question.

synapsesocial.com/papers/6a27ae3fa963992e16268415https://doi.org/10.1002/bte2.70132
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Energy Arbitrage Analysis for Market-Selection of a Battery Energy Storage System-Based Venture2025 · 5 citations
  2. 2Exploration in deep reinforcement learning: A survey2022 · 516 citations
  3. 3Critical Bus Voltage Support in Distribution Systems With Electric Springs and Responsibility Sharing2016 · 63 citations
  4. 4A model based balancing system for battery energy storage systems2022 · 22 citations
  5. 5Ride-sharing with travel time uncertainty2018 · 94 citations