PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 28, 2026Mathematics2 citationsOpen Access

Optimizing Crypto-Trading Performance: A Comparative Analysis of Innovative Reward Functions in Reinforcement Learning Models

View Full Paper
EKErgashevich Halimjon KhujamatovKIKobuljon IsmanovOMOybek Usmankulovich Mallaev

Key Points

  • This research aims to improve cryptocurrency trading performance by evaluating innovative reward functions in reinforcement learning models.
  • Introduces five new reward functions aimed at improving risk management and market adaptability.
  • Evaluates performance using three reinforcement learning algorithms: Deep Q-Network, Proximal Policy Optimization, and Advantage Actor–Critic.
  • Uses Bitcoin hourly trading data from 2018–2022 across four market regimes: bull, bear, high volatility, and recovery.
  • Adaptive Risk Control reward function achieved a Sharpe ratio of 2.47 and a cumulative return of 26.4%.
  • The function's maximum drawdown during a bearish market was only 16.8%.
  • Performance varied significantly by regime, with Adaptive Risk Control excelling in high volatility (Sharpe ratio 3.21).

Abstract

Cryptocurrency trading presents significant challenges due to extreme market volatility, rapid regime transitions, and non-stationary dynamics that render traditional trading strategies ineffective. Existing reinforcement learning approaches for cryptocurrency trading typically employ simplistic profit-based reward functions that fail to adequately capture risk management considerations, market microstructure costs, temporal dependencies, and regime-specific optimal behaviors. This limitation often results in strategies that perform well during favorable market conditions but suffer catastrophic losses during downturns. This paper introduces five novel reward functions grounded in economic utility theory, market microstructure, behavioral finance, adaptive risk management, and regime-conditional optimization. We systematically evaluate these reward functions across three reinforcement learning algorithms (Deep Q-Network, Proximal Policy Optimization, and Advantage Actor–Critic) and four distinct market regimes (bull, bear, high volatility, and recovery), using Bitcoin hourly data from 2018–2022. Our comprehensive experimental evaluation demonstrates that the Adaptive Risk Control reward function achieves exceptional performance, with a Sharpe ratio of 2.47, cumulative return of 26.4%, and maximum drawdown of only 16.8% during the predominantly bearish 2022 test period. Critically, regime-specific analysis reveals substantial performance heterogeneity: Adaptive Risk Control excels during high volatility (Sharpe ratio 3.21), while Temporal Coherence and Asymmetric Market-Conditional rewards dominate in trending and bear markets, respectively. These findings establish that sophisticated, theory-grounded reward engineering—rather than algorithmic innovations alone—constitutes the primary lever for improving RL trading systems, enabling positive risk-adjusted returns even during severe market downturns.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Khujamatov et al. (2026) studied this question.

synapsesocial.com/papers/69a286720a974eb0d3c0156chttps://doi.org/10.3390/math14050794
Ask AI
Helpful
Bookmark
Share
View Full Paper