Deep Reinforcement Learning for Personalized Cancer Treatment Planning and Radiotherapy Dose Optimization.
Volume 5
Abstract
Personalized cancer treatment planning is a complex clinical decision-making process that requires the optimal selection of therapeutic strategies while balancing tumor control, patient-specific anatomical constraints, treatment toxicity, and long-term survival outcomes. Radiotherapy remains a major treatment modality for many cancer types; however, designing high-quality radiation treatment plans is often time-consuming and highly dependent on expert clinical judgment. Conventional treatment planning methods typically rely on manual parameter adjustment, predefined dose constraints, and iterative optimization procedures, which may limit consistency, scalability, and adaptability across heterogeneous patient populations. To address these challenges, this paper proposes a Deep Reinforcement Learning (DRL)-based framework for personalized cancer treatment planning and radiotherapy dose optimization.
The proposed framework formulates cancer treatment planning as a sequential decision-making problem, where an intelligent DRL agent learns to recommend patient-specific treatment actions based on evolving clinical states. These states include tumor characteristics, imaging-derived anatomical features, dose–volume histogram parameters, organ-at-risk constraints, patient demographics, staging information, prior treatment response, and relevant molecular or clinical biomarkers where available. The action space includes adaptive radiotherapy parameter adjustment, beam configuration refinement, fractionation strategy selection, dose escalation or de-escalation decisions, and optimization priority tuning for tumor coverage and organ-at-risk sparing. A reward function is designed to maximize tumor control probability while minimizing normal tissue complication probability, cumulative toxicity, and violation of clinically accepted dose constraints.
The proposed architecture integrates advanced DRL algorithms, including Deep Q-Networks, Proximal Policy Optimization, Actor–Critic models, and Deep Deterministic Policy Gradient, to support both discrete and continuous treatment optimization scenarios. To enhance patient-specific adaptation, the framework incorporates multimodal data from computed tomography, magnetic resonance imaging, positron emission tomography, electronic health records, pathology reports, and radiotherapy planning systems. A comprehensive preprocessing pipeline is applied, including image registration, tumor and organ-at-risk segmentation, feature extraction, dose normalization, missing data imputation, and construction of individualized treatment-state trajectories. The DRL agent is trained using historical treatment plans, simulated dose distributions, and clinically validated planning objectives to learn optimal policies that reflect expert decision-making while improving planning efficiency.
The framework is evaluated using retrospective radiotherapy datasets and simulated treatment planning environments across selected cancer types, including lung, head-and-neck, prostate, and cervical cancer. Model performance is assessed using clinically meaningful metrics such as planning target volume coverage, dose conformity index, homogeneity index, organ-at-risk dose reduction, tumor control probability, normal tissue complication probability, cumulative toxicity score, and overall plan quality score. Comparative analysis is conducted against conventional inverse planning, rule-based optimization, manual expert planning, and non-reinforcement learning-based deep learning models. Experimental evaluation is expected to demonstrate that the proposed DRL-based system can generate clinically acceptable and patient-specific radiotherapy plans with improved consistency, reduced planning time, and better balance between tumor irradiation and healthy tissue protection.
To improve clinical transparency and physician trust, the proposed framework incorporates explainability mechanisms, including policy visualization, reward decomposition, dose–volume contribution analysis, and patient-level decision trajectory interpretation. These tools allow radiation oncologists and medical physicists to understand why specific dose adjustments or treatment actions are recommended for an individual patient. The proposed system offers a scalable and adaptive decision-support platform for precision oncology by combining reinforcement learning, medical imaging, radiotherapy physics, and patient-specific clinical evidence. Its integration into clinical workflows has the potential to support more efficient treatment planning, personalized dose optimization, reduced toxicity risk, and improved quality of cancer care.
Keywords
References
- [1] El-Kenawy, E. S. M., Khodadadi, N., Mirjalili, S., Abdelhamid, A. A., Eid, M. M., & Ibrahim, A. (2024). Greylag goose optimizati on:
- [2] Abdollahzadeh, B., Khodadadi, N., Barshandeh, S., Trojovský, P., Gharehchopogh, F. S., El -kenawy, E. S. M., ... & Mirjalili, S. (2024).
- [3] El-Kenawy, E. S. M., Ibrahim, A., Mirjalili, S., Eid, M. M., & Hussein, S. E. (2020). Novel feature selection and voting classif ier
- [4] El-Kenawy, E. S. M., Eid, M. M., Saber, M., & Ibrahim, A. (2020). MbGWO -SFS: Modified binary grey wolf optimizer based on
- [5] El-Kenawy, E. S., & Eid, M. (2020). Hybrid gray wolf and particle swarm optimization for feature selection. Int. J. Innov. Compu t. Inf.
- [6] El-Kenawy, E. S. M., Mirjalili, S., Ibrahim, A., Alrahmawy, M., El -Said, M., Zaki, R. M., & Eid, M. M. (2021). Advanced meta -
- [7] Khodadadi, N., Abualigah, L., El -Kenawy, E. S. M., Snasel, V., & Mirjalili, S. (2022). An archive -based multi -objective arithmetic
- [8] Ibrahim, A., Mirjalili, S., El -Said, M., Ghoneim, S. S., Al -Harthi, M. M., Ibrahim, T. F., & El -Kenawy, E. S. M. (2021). Wind speed
- [9] Abdelhamid, A. A., El -Kenawy, E. S. M., Alotaibi, B., Amer, G. M., Abdelkader, M. Y., Ibrahim, A., & Eid, M. M. (2022). Robust
- [10] El-Kenawy, E. S. M., Mirjalili, S., Alassery , F., Zhang, Y. D., Eid, M. M., El -Mashad, S. Y., ... & Abdelhamid, A. A. (2022). Novel
- [11] Hassib, E. M., El -Desouky, A. I., Labib, L. M., & El -Kenawy, E. S. M. (2020). WOA+ BRNN: An imbalanced big data classification
- [12] Eid, M. M., El -kenawy, E. S. M., & Ibrahim, A. (2021, March). A binary sine cosine -modified whale optimization algorithm for feature
- [13] El-Sayed Towfek, M. (2018). El -kenawy. Trust Model for Dependable File Exchange in Cloud Computing. International Journal of
- [14] Abdelhamid, A. A., Towfek, S. K., Khodadadi, N., Alhussan, A. A., Khafaga, D. S., Eid, M. M., & Ibrahim, A. (2023). Waterwhee l plant
- [15] El-Kenawy, E. S. M., Abdelhamid, A. A., Ibrahim, A., Mirjalili, S., Khodadad, N., Alduailij, M. A., ... & Khafaga, D. S. (2023). Al-
- [16] Alhussan, A. A., Abdelhamid, A. A., El -Kenawy, E. S. M., Ibrahim, A., Eid, M. M., Khafaga, D. S., & Ahmed, A. E. (2023). A binary
- [17] Alhussan, A. A., Khafaga, D. S., Abotaleb, M., Mishra, P., & El -Kenawy, E. S. M. (2024). Global Potato Production Forecasting Based
- [18] Towfek, S. K., & Alhussan, A. A. (2024). Potato Production Forecasting Based on Balance Dynamic Biruni Earth Radius Algorithm for
- [19] Mishra, P., Alhussan, A. A., Khafaga, D. S., Lal, P., Ray, S., Abotaleb, M., ... & El -Kenawy, E. S. M. (2024). Forecasting production of
- [20] El-Kenawy, E. S. M., Mirjalili, S., Abdelhamid, A. A., Ibrahim, A., Khodadadi, N., & Eid, M. M. (2022). Meta -heuristic optimization
- [21] Abdelhamid, A. A., El -Kenawy, E. S. M., Khodadadi, N., Mirjalili, S., Khafaga, D. S., Alharbi, A. H., ... & Saber, M. (2022).
- [22] Alharbi, A. H., Towfek, S. K., Abdelhamid, A. A., Ibrahim, A., Eid, M. M., & Khafaga, D. S. & Saber, M.(2023). Diagnosis of
- [23] Alhussan, A. A., & Towfek, S. K. (2024). 5G Resource Allocation Using Feature Selection and Greylag Goose Optimization Algori thm.
- [24] Towfek, S. K., & Alhussan, A. A. (2024). Potato Production Forecasting Based on Balance Dynamic Biruni Earth Radius Algorithm for
- [25] Abdelhamid, A. A., Alhussan, A. A., Qenawy, A. S. T., Osman, A. M., Elshewey, A. M., & Eed, M. (2024). Potato harvesting pred iction
- [26] Eed, M., Alhussan, A. A., Qenawy, A. S. T., Osman, A. M., Elshewey, A. M., & Arnous, R. (2024). Potato Consumption Forecastin g
- [27] Mahmood, S., Sun, H., Iqbal, A., Alhussan, A. A., & El -kenawy, E. S. M. (2024). Green finance, sustainable infrastructure, and green
- [28] El-Kenawy, E. S. M., Alhussan, A. A., Khodadadi, N., Mirjalili, S., & Eid, M. M. (2024). Predicting potato crop yield with machi ne
- [29] Radwan, M., Alhussan, A. A., Ibrahim, A., & Tawfeek, S. M. (2025). Potato leaf disease classification using optimized machine learning
- [30] Ghasemi, M., Khodadadi, N., Trojovský, P., Li, L., Mansor, Z., Abualigah, L., ... & El -Kenawy, E. S. M. (2025). Kirchhoff’s law
- [31] Yassen, M. A., El -Kenawy, E. S. M., Abdel -Fattah, M. G., Ismail, I., & Mostafa, H. E. D. S. (2025). Explainable artificial intelligence
- [32] Radwan, M., Alhussan, A. A., Ibrahim, A., & Tawfeek, S. M. (2025). Potato leaf disease classification using optimized machine learning
- [33] Mozhdehi, A. T., Khodadadi, N., Aboutalebi, M., El -Kenawy, E. S. M., Hussien, A. G., Zhao, W., ... & Mirjalili, S. (2025). Divine
- [34] El-kenawy, E. S. M., Alhussan, A. A., Mattar, E. A., & Radwan, M. (2026). Feature selection and hyperparameter tuning in transfo rmer -
- [35] Chen, L., Xu, C., Lim, W. H., Sharma, A., Tiang, S. S., Chong, K. S., ... & Khafaga, D. S. (2025). Transparent and reliable c onstruction
- [36] Khodadadi, N., Towfek, S. K., Zaki, A. M., Alharbi, A. H., Khodadadi, E., Khafaga, D. S., ... & Eid, M. M. (2025). Predicting
- [37] Alhussan, A. A., El -Kenawy, E. S. M., Khafaga , D. S., Alharbi, A. H., & Eid, M. M. (2025). Groundwater resource prediction and
- [38] Mozhdehi, A. T., Khodadadi, N., Aboutalebi, M., El -Kenawy, E. S. M., Hussien, A. G., Zhao, W., ... & Mirjalili, S. (2025). Divine
- [39] Lopes, J. M., Pinho, C. S., Martins, R. C., & Bandeira, T. (2026). Sustainability Gets Smarter: Competitive Pressure and AI L eading the
- [40] Radwan, M., Ibrahim, A., Abdelsalam, M. M., Alhussan, A. A., Mattar, E. A., & El -Kenawy, E. S. M. (2026). Optimizing solar and wind
- [41] Alhussan, A. A., El -Kenawy, E. S. M., Eid, M. M., & Khodadadi, N. (2026). Hybrid Al -Biruni and Puma Optimization (BERPO) for
- [42] El-Kenawy, E. S. M., Ibrahim, A., Alhussan, A. A., Khafaga, D. S., Ahmed, A. E., & Eid, M. M. (2026). Smart city electricity loa d
- [43] El-Kenawy, E. S. M., Khodadadi, N., Mirjalili, S., Zaki, A. M., Ibrahim, A., Alhussan, A. A., ... & Eid, M. M. (2026). Glider sn ake
- [44] Mekaret, F., Rabehi, A., Zebentout, B., Tizi, S., Douara, A., Bellucci, S., ... & Alhussan, A. A. (2024). A comparative study of Schottky
- [45] KOUADRI, Ali, RABEHI, Abdelhalim, BENZIANE, Ali, et al. A Robust Multi -Transform Watermarking Scheme for Medical Images
- [46] Tibermacine, I. E., Russo, S., Scarano, G., Tedesco, G., Rabehi, A., Alhussan, A. A., ... & Napoli, C. (2025). Conditional VA E for
- [47] Kouadri, A., Benziane, A., Rabehi, A., Rabehi, A., Alhussan, A. A., Khafaga, D. S., & El -Kenawy, E. S. M. (2025). A novel hybrid
- [48] Russo, S., Tibermacine, I. E., Randieri, C., Rabehi, A., Alharbi, A. H., El -Kenawy, E. S. M., & Napoli, C. (2025). Exploiting facial
- [49] Bentegri, H., Rabehi, M., Kherfane, S., Nahool, T. A., Rabehi, A., Guermoui, M., ... & El -Kenawy, E. S. M. (2025). Assessment of
- [50] Tibermacine, I. E., Russo, S., Citeroni, F., Mancini, G., Rabehi, A., Alharbi, A. H., ... & Napoli, C. (2025). Adversarial de noising of EEG
- [51] Ouahabi, M. S., Benyounes, A., Barkat, S., Ihammouchen, S., Rekioua, T., Rabehi, A., ... & Alharbi, A. H. (2025). Real -time sensor fault
- [52] Belaid, A., Guermoui, M., Khelifi, R., Arrif, T., Chekifi, T., Rabehi, A., ... & Alhussan, A. A. (2024). Assessing suitable a reas for PV
- [53] Mehallou, A., M’hamdi, B., Amari, A., Teguar, M., Rabehi, A., Guermoui, M., ... & Khafaga, D. S. (2025). Optimal multiobjecti ve
- [54] Rabehi, A., El -Hadi, M., Benmahmoud, S., Rabehi, A., Alharbi, A. H., & El -Kenawy, E. S. M. (2025). SOCA -CFAR Processor in A
- [55] Bakria, D., Beladel, A., Korich, B., Teta, A., Mohammedi, R. D., Laouid, A. A., ... & El -kenawy, E. S. (2025). A novel enhanced Grey
