We investigate the distinct trading behaviors of domestic and foreign brokerage firms in Borsa Istanbul using 2506 trading days. Applying linear and maximum entropy inverse reinforcement learning, we recover the latent reward functions driving broker decisions and contrast them with supervised learning baselines. We further assess the financial viability of the inferred policies through historical backtesting, using metrics such as Sharpe ratio and maximum drawdown. Our findings reveal a pronounced strategic divergence: domestic and foreign brokers exhibit uncorrelated reward structures when facing identical market states. These insights into agent heterogeneity offer regulators and market participants a novel tool for monitoring market microstructure and stability.
Karacam et al. (Wed,) studied this question.