Multi-Agent Deep Reinforcement Learning for Decentralized Portfolio Management Under Stochastic Market Conditions.
Volume 5
Abstract
Decentralized portfolio management under stochastic market conditions is a complex sequential decision-making problem characterized by uncertain asset returns, dynamic correlations, transaction costs, liquidity constraints, volatility clustering, and rapidly changing risk preferences. Traditional portfolio optimization models often depend on centralized decision structures, fixed assumptions about return distributions, and static risk-return trade-offs, which limit their ability to adapt to heterogeneous investor objectives and unstable market regimes. This paper proposes a multi-agent deep reinforcement learning framework for decentralized portfolio management in stochastic financial environments. The proposed framework models each trading agent as an autonomous decision-maker responsible for learning adaptive asset allocation, rebalancing, and risk-control policies through continuous interaction with a simulated market environment. Multiple deep reinforcement learning algorithms, including Multi-Agent Deep Deterministic Policy Gradient, Multi-Agent Proximal Policy Optimization, Independent Deep Q-Networks, and actor-critic-based architectures, are employed to coordinate decentralized trading decisions while preserving agent-level autonomy. The state space incorporates historical asset prices, log returns, trading volume, volatility indicators, moving averages, momentum signals, correlation measures, macro-financial variables, and portfolio-specific risk exposures, while the action space represents dynamic portfolio weights, buy-sell decisions, and rebalancing strategies across multiple asset classes. A comprehensive preprocessing pipeline was applied, including missing value treatment, feature normalization, rolling-window state construction, covariance estimation, transaction cost modeling, and chronological train-validation-test splitting to prevent look-ahead bias. The reward function was designed to optimize long-term risk-adjusted performance by integrating portfolio return, Sharpe ratio, downside risk, maximum drawdown, turnover penalty, and diversification constraints. Experimental results demonstrate that the proposed multi-agent framework achieves superior performance compared with traditional benchmarks, including buy-and-hold, equal-weighted allocation, mean-variance optimization, risk parity, momentum-based strategies, and single-agent reinforcement learning models. The proposed approach records higher cumulative returns, improved Sharpe and Sortino ratios, lower maximum drawdown, reduced portfolio volatility, and stronger resilience during stochastic market shocks. Furthermore, inter-agent coordination analysis reveals that decentralized agents can learn complementary investment behaviors, adaptive diversification patterns, and regime-sensitive allocation strategies under uncertain market dynamics. The findings indicate that multi-agent deep reinforcement learning provides a scalable, adaptive, and robust solution for decentralized portfolio optimization, offering practical value for quantitative asset management, automated trading systems, robo-advisory platforms, and financial institutions operating in volatile and uncertain investment environments.
Top of Form
Keywords
References
- [1] Lopes, J. M., Pinho, C. S., Martins, R. C., & Bandeira, T. (2026). Sustainability Gets Smarter: Competitive Pressure and AI L eading the
- [2] Radwan, M., Ibrahim, A., Abdelsalam, M. M., Alhussan, A. A., Mattar, E. A., & El -Kenawy, E. S. M. (2026). Optimizing solar and wind
- [3] Alhussan, A. A., El -Kenawy, E. S. M., Eid, M. M., & Khodadadi, N. (2026). Hybrid Al -Biruni and Puma Optimization (BERPO) for
- [4] El-Kenawy, E. S. M., Ibrahim, A., Alhussan, A. A., Khafaga, D. S., Ahmed, A. E., & Eid, M. M. (2026). Smart city electricity loa d
- [5] El-Kenawy, E. S. M., Khodadadi, N., Mirjalili, S., Zaki, A. M., Ibrahim, A., Alhussan, A. A., ... & Eid, M. M. (2026). Glider sn ake
- [6] Mekaret, F., Rabehi, A., Zebentout, B., Tizi, S., Douara, A., Bellucci, S., ... & Alhussan, A. A. (2024). A comparative study of Schottky
- [7] KOUADRI, Ali, RABEHI, Abdelhalim, BENZIANE, Ali, et al. A Robust Multi -Transform Watermarking Scheme for Medical Images
- [8] Tibermacine, I. E., Russo, S., Scarano, G., Tedesco, G., Rabehi, A., Alhussan, A. A., ... & Napoli, C. (2025). Conditional VA E for
- [9] Kouadri, A., Benziane, A., Rabehi, A., Rabehi, A., Alhussan, A. A., Khafaga, D. S., & El -Kenawy, E. S. M. (2025). A novel hybrid
- [10] Russo, S., Tibermacine, I. E., Randieri, C., Rabehi, A., Alharbi, A. H., El -Kenawy, E. S. M., & Napoli, C. (2025). Exploiting facial
- [11] Bentegri, H., Rabehi, M., Kherfane, S., Nahool, T. A., Rabehi, A., Guermoui, M., ... & El -Kenawy, E. S. M. (2025). Assessment of
- [12] Tibermacine, I. E., Russo, S., Citeroni, F., Mancini, G., Rabehi, A., Alharbi, A. H., ... & Napoli, C. (2025). Adversarial de noising of EEG
- [13] Ouahabi, M. S., Benyounes, A., Barkat, S., Ihammouchen, S., Rekioua, T., Rabehi, A., ... & Alharbi, A. H. (2025). Real -time sensor fault
- [14] Belaid, A., Guermoui, M., Khelifi, R., Arrif, T., Chekifi, T., Rabehi, A., ... & Alhussan, A. A. (2024). Assessing suitable a reas for PV
- [15] Mehallou, A., M’hamdi, B., Amari, A., Teguar, M., Rabehi, A., Guermoui, M., ... & Khafaga, D. S. (2025). Optimal multiobjecti ve
- [16] Rabehi, A., El -Hadi, M., Benmahmoud, S., Rabehi, A., Alharbi, A. H., & El -Kenawy, E. S. M. (2025). SOCA -CFAR Processor in A
- [17] Bakria, D., Beladel, A., Korich, B., Teta, A., Mohammedi, R. D., Laouid, A. A., ... & El -kenawy, E. S. (2025). A novel enhanced Grey
- [18] Towfek, S. K., & Alhussan, A. A. (2024). Potato Production Forecasting Based on Balance Dynamic Biruni Earth Radius Algorithm for
- [19] Mishra, P., Alhussan, A. A., Khafaga, D. S., Lal, P., Ray, S., Abotaleb, M., ... & El -Kenawy, E. S. M. (2024). Forecasting production of
- [20] El-Kenawy, E. S. M., Mirjalili, S., Abdelhamid, A. A., Ibrahim, A., Khodadadi, N., & Eid, M. M. (2022). Meta -heuristic optimization
- [21] Abdelhamid, A. A., El -Kenawy, E. S. M., Khodadadi, N., Mirjalili, S., Khafaga, D. S., Alharbi, A. H., ... & Saber, M. (2022).
- [22] Alharbi, A. H., Towfek, S. K., Abdelhamid, A. A., Ibrahim, A., Eid, M. M., & Khafaga, D. S. & Saber, M.(2023). Diagnosis of
- [23] Alhussan, A. A., & Towfek, S. K. (2024). 5G Resource Allocation Using Feature Selection and Greylag Goose Optimization Algori thm.
- [24] Towfek, S. K., & Alhussan, A. A. (2024). Potato Production Forecasting Based on Balance Dynamic Biruni Earth Radius Algorithm for
- [25] Abdelhamid, A. A., Alhussan, A. A., Qenawy, A. S. T., Osman, A. M., Elshewey, A. M., & Eed, M. (2024). Potato harvesting pred iction
- [26] Eed, M., Alhussan, A. A., Qenawy, A. S. T., Osman, A. M., Elshewey, A. M., & Arnous, R. (2024). Potato Consumption Forecastin g
- [27] Mahmood, S., Sun, H., Iqbal, A., Alhussan, A. A., & El -kenawy, E. S. M. (2024). Green finance, sustainable infrastructure, and green
- [28] El-Kenawy, E. S. M., Alhussan, A. A., Khodadadi, N., Mirjalili, S., & Eid, M. M. (2024). Predicting potato crop yield with machi ne
- [29] Radwan, M., Alhussan, A. A., Ibrahim, A., & Tawfeek, S. M. (2025). Potato leaf disease classification using optimized machine learning
- [30] Ghasemi, M., Khodadadi, N., Trojovský, P., Li, L., Mansor, Z., Abualigah, L., ... & El -Kenawy, E. S. M. (2025). Kirchhoff’s law
- [31] Yassen, M. A., El -Kenawy, E. S. M., Abdel -Fattah, M. G., Ismail, I., & Mostafa, H. E. D. S. (2025). Explainable artificial intelligence
- [32] Radwan, M., Alhussan, A. A., Ibrahim, A., & Tawfeek, S. M. (2025). Potato leaf disease classification using optimized machine learning
- [33] Mozhdehi, A. T., Khodadadi, N., Aboutalebi, M., El -Kenawy, E. S. M., Hussien, A. G., Zhao, W., ... & Mirjalili, S. (2025). Divine
- [34] El-kenawy, E. S. M., Alhussan, A. A., Mattar, E. A., & Radwan, M. (2026). Feature selection and hyperparameter tuning in transfo rmer -
- [35] Chen, L., Xu, C., Lim, W. H., Sharma, A., Tiang, S. S., Chong, K. S., ... & Khafaga, D. S. (2025). Transparent and reliable c onstruction
- [36] Khodadadi, N., Towfek, S. K., Zaki, A. M., Alharbi, A. H., Khodadadi, E., Khafaga, D. S., ... & Eid, M. M. (2025). Predicting
- [37] Alhussan, A. A., El -Kenawy, E. S. M., Khafaga , D. S., Alharbi, A. H., & Eid, M. M. (2025). Groundwater resource prediction and
- [38] Mozhdehi, A. T., Khodadadi, N., Aboutalebi, M., El -Kenawy, E. S. M., Hussien, A. G., Zhao, W., ... & Mirjalili, S. (2025). Divine
