PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 22, 2025Working Papers.1 citations

Application of Deep Reinforcement Learning to At-the-Money S&P 500 Options Hedging

View Full Paper
ZBZofia BrachaUniversity of WarsawJMJakub MichańkówKrakow University of EconomicsPSPaweł SakowskiCenter for Social and Economic Research

Key Points

  • The deep reinforcement learning agent demonstrated superior performance over traditional strategies, especially in volatile conditions.
  • Using the TD3 algorithm, the agent was assessed against the Black–Scholes strategy with metrics like Sharpe ratio and information ratio.
  • Performance evaluation covered nearly 17 years of data, showing a clear adaptability of the model across diverse market conditions.
  • Increased risk-awareness penalties negatively affected the agent's performance, revealing limits to its effectiveness.

Abstract

This paper explores the application of deep Q-learning to hedging at-the-money options on the S&P 500 index. We develop an agent based on the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, trained to simulate hedging decisions without making explicit model assumptions on price dynamics. The agent was trained on historical intraday prices of S&P 500 call options across years 2004 to 2024, using a single time series of six predictor variables: option price, underlying asset price, moneyness, time to maturity, realized volatility, and current hedge position. A walk-forward procedure was applied for training, which lead to nearly 17 years of out-of-sample evaluation. The performance of the deep reinforcement learning (DRL) agent is benchmarked against the Black–Scholes delta hedging strategy over the same time period. We assess both approaches using metrics such as annualized return, volatility, information ratio, and Sharpe ratio. To test models’ adaptability, we performed simulations across varying market conditions and added constraints such as transaction costs and risk-awareness penalties. Our results show that the DRL agent can outperform traditional hedging methods, particularly in volatile or high-cost environments, highlighting its robustness and flexibility in practical trading contexts. While the agent consistently outperforms delta hedging, its performance deteriorates when the risk-awareness parameter is higher. We also observed that the longer the time interval used for volatility estimation, the more stable the results.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bracha et al. (2025) studied this question.

synapsesocial.com/papers/68f83311d24b29c969481736https://doi.org/10.33138/2957-0506.2025.25.488
Ask AI
Helpful
Bookmark
Share
View Full Paper