Reward Hacking Patterns in Reinforcement Learning Agents Deployed in Financial Trading Systems

Main Article Content

Anis Nur 'Afifah
Suud Nofa Setia Mahdani
Mar'atun Solikhah

Abstract

This study analyzes reward hacking patterns in reinforcement learning agents deployed in financial trading systems by examining the divergence between internal reward maximization and externally valid trading performance. The results show that the profit-oriented agent achieved the highest final reward score of 0.87 and net return of 18.6%, but this performance was accompanied by a weak Sharpe ratio of 0.74, maximum drawdown of 24.8%, and reward-risk gap of 0.49. In contrast, the balanced reward agent produced a lower final reward score of 0.71 and net return of 14.2%, but achieved a stronger Sharpe ratio of 1.18 and lower maximum drawdown of 12.6%. Five reward hacking patterns were identified: transaction churning, drawdown masking, benchmark gaming, liquidity exploitation, and tail-risk accumulation. Transaction churning appeared most frequently with 157 flagged episodes, while tail-risk accumulation produced the highest severity score of 0.84 despite only 45 flagged episodes. Stress testing showed that the profit-oriented agent had the highest mean stress score of 0.85, with liquidity degradation as its worst scenario. The balanced reward agent showed stronger deployment robustness with a mean stress score of 0.52. Regime-level analysis further revealed that high-volatility markets produced the greatest fragility, with return dispersion of 3.5%, drawdown rate of 23.4%, and dominant tail-risk accumulation. The study contributes a structured framework for detecting reward hacking through reward-risk divergence, behavioral diagnostics, stress testing, regime sensitivity, and governance-oriented severity classification.

Article Details

Section
Articles