Simulation study demonstrates that entropy-adapted proximal policy optimization improves retailer profits and grid peak load reduction, indicating effective AI-driven electricity pricing.
The dynamic pricing mechanism in the electricity retail market plays a pivotal role in achieving optimal resource allocation on both supply and demand sides. However, the complex heterogeneity of consumer behavior at the retail level coupled with frequent price fluctuations in the wholesale market poses significant challenges to traditional optimization approaches. This paper models the dynamic pricing problem of power retailers as a Markov decision process within a continuous action space and proposes a solution framework based on proximal policy optimization algorithms. Centered on maximizing cumulative profits for power retailers, the framework integrates user price response characteristics, procurement cost constraints in wholesale markets, and regulatory boundaries for retail pricing into a unified modeling system, enabling adaptive generation of retail electricity prices across time periods through continuous interaction between agents and their environment. To address common issues with standard PPO algorithms–including premature policy degradation and insufficient exploration efficiency–the study introduces an adaptive exploration coefficient adjustment scheme based on policy entropy, along with a generalized advantage estimation method to reduce gradient estimation variance. Simulation experiments using actual operational data from a provincial electricity market in China during 2023 demonstrated that the proposed algorithm outperforms benchmark algorithms such as the Deep Deterministic Policy Gradient and Dual Delayed Deep Deterministic Policy Gradient in key metrics including average daily retailer profits, peak load reduction rates, and price series stability, validating the feasibility and practical effectiveness of reinforcement learning approaches for dynamic pricing in electricity retail markets.
No takes yet. Share an insight, caveat, or question.
Xu et al. (2026) studied this question.