Abstract. In this paper, a multicriteria reward function is developed and empirically validated within a market-neutral portfolio management framework based on deep reinforcement learning. This function formalizes the optimal capital allocation problem as a Markov decision process in a nonstationary environment. The reward function simultaneously optimizes excess return per unit of risk, market neutrality via an explicit penalty on the correlation with a target market index, and transaction costs. As a result, systemic risk control and the cost of portfolio rebalancing are embedded into the modeling and training procedure. The model is trained and analyzed using cross-validation to mitigate the bias and adapt to regime shifts. According to the experimental validation results on leading equity indices of the USA, Europe, and Asia over a ten-year horizon, the above framework is effective, particularly showing consistent improvements in key metrics: higher Sharpe ratios, lower maximum drawdowns, and a substantially reduced correlation with the market index relative to classical optimization algorithms and reinforcement learning approaches with a single criterion.
Keywords: market-neutral strategy, portfolio optimization, reinforcement learning, market-neutral portfolio, deep learning, portfolio management.
Acknowledgments. The author is grateful to HSE colleagues for their valuable discussions, constructive feedback, and continued support of this work.
Belyakov, B.E., Application of a Multi-Objective Reward Function to Dynamic Risk Management in Market-Neutral Portfolios. Control Sciences 3, 42–52 (2026).