PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 26, 2026Machine Learning with Applications0 citationsOpen Access

AI-powered predictive control for hybrid renewable microgrids: Integrating photovoltaic generation and battery storage in smart homes and industrial applications

View Full Paper
SFSebastián López FlórezGHGuillermo HernándezTCTpc Co-Chair

Key Points

  • The main aim is to establish a benchmark for deep reinforcement learning in microgrid energy management.
  • Developed a standardized evaluation protocol for energy management under different regimes.
  • Defined a common state/action/reward specification focusing on battery health, pricing, and constraints.
  • Adapted various reinforcement learning algorithms including value-based and policy-gradient methods.
  • Benchmarked agents using real-world datasets to assess performance.
  • Adapted RL agents reduced daily operating costs and improved battery health under realistic conditions.
  • DQN-based variants showed consistent performance improvements over standard methods.
  • PPO and actor–critic configurations maintained stability under price volatility and renewable energy fluctuations.
  • Performance varied depending on the dataset and operating conditions.

Abstract

The integration of renewable energy sources into microgrids remains challenging due to generation intermittency, storage inefficiencies, and progressive battery degradation. In this work, our main contribution is a standardized empirical benchmark of deep reinforcement learning for microgrid energy management across multiple operating regimes under a consistent evaluation protocol. To make these evaluations rigorous and comparable, we define a common state/action/reward specification that explicitly encodes battery-health proxies, dynamic pricing, and operational constraints, and we adapt a representative set of RL families accordingly: value-based methods (DoubleDQN, NoisyNet-DQN, PER-DQN, C51), policy-gradient approaches (PG), and actor–critic algorithms (PPO, A2C, SAC, DDPG, A3C-Energy). We benchmark the resulting agents on real-world driven datasets and environments (GEFCom2014, StoreNet, and a CityLearn-inspired setting), assessing performance under a consistent evaluation pipeline. Empirically, the reported results indicate that the adapted agents tend to reduce daily operating cost, smooth charge–discharge cycling, and improve battery-health proxies under realistic operating conditions, with gains that vary by dataset and operating regime. In particular, DQN-based variants exhibit consistent gains over standard exploration schemes, while methods such as PPO and actor–critic configurations maintain competitive and more stable behavior under highly volatile price signals and renewable fluctuations. Overall, the study provides scenario-conditional evidence on which RL families tend to perform well or poorly under the tested conditions, rather than a universal ranking across datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Flórez et al. (2026) studied this question.

synapsesocial.com/papers/699fe28895ddcd3a253e6468https://doi.org/10.1016/j.mlwa.2026.100864
Ask AI
Helpful
Bookmark
Share
View Full Paper