This paper develops a deep reinforcement learning framework for cryptocurrency portfolio management in which transaction costs are derived from the Riemannian geometry of the underlying volatility model rather than assumed constant. A Proximal Policy Optimisation agent is trained on a reward function grounded in non-equilibrium thermodynamics: we use the free-energy Bellman equation, in which transaction costs are the geodesic slippage on the Fisher information manifold of a maximum-entropy Markov-switching GARCH model, and regime-transition costs are the Wasserstein-2 distance between the calm and turbulent return distributions. A thermodynamic Carnot bound on portfolio efficiency is established and empirically validated. Five hypotheses are tested across Bitcoin, Ethereum, Ripple, Litecoin, and Bitcoin Cash over January 2017 to March 2026. The geometric-cost agent achieves statistically superior Sharpe ratios relative to flat-fee baselines on four of five assets; portfolio turnover is reduced by 56 to 83 percent relative to signal-following; the thermodynamic friction point at which the agent prefers no-trade is asset-specific and ordered by turbulent half-life; a joint topological and geometric circuit breaker reduces Maximum Drawdown by 28 to 38 percent; and ablation confirms that every component of the observation vector contributes a statistically significant performance gain. The framework requires liquid cryptocurrency markets with validated parametric volatility models; transferability to other asset classes requires upstream recalibration.
Ntebogang Dinah Moroke (2026) studied this question.