This paper proposes a dynamic portfolio allocation framework that integrates deep reinforcement learning (DRL) with classical portfolio optimization to enhance rebalancing strategies and risk–return management. Within a unified reinforcement-learning environment for portfolio reallocation, we train actor–critic agents (Proximal Policy Optimization (PPO) and Advantage Actor–Critic (A2C)). These agents learn to select both the risk-aversion level—positioning the portfolio along the efficient frontier defined by expected return and a chosen risk measure (variance, Semivariance, or CVaR)—and the rebalancing horizon. An ensemble procedure, which selects the most effective agent–utility combination based on the Sharpe ratio, provides additional robustness. Unlike approaches that directly estimate portfolio weights, our framework retains the optimization structure while delegating the choice of risk level and rebalancing interval to the AI agent, thereby improving stability and incorporating a market-timing component. Empirical analysis on daily data for 12 U.S. sector ETFs (2003–2023) and 28 Dow Jones Industrial Average components (2005–2023) demonstrates that DRL-guided strategies consistently outperform static tangency portfolios and market benchmarks in annualized return, volatility, and Sharpe ratio. These findings underscore the potential of DRL-driven rebalancing for adaptive portfolio management.
Yu et al. (2025) studied this question.