DOI: 10.1049/syb2.70083 ISSN: 1751-8849

Double Q‐Learning for Intelligent Multi‐Drug Scheduling in Cancer Chemotherapy Optimisation

Behnoush Alizade, Ahmad Hajipour

ABSTRACT

Chemotherapy scheduling poses a challenging control problem due to the need to suppress tumour growth whilst maintaining systemic toxicity within clinically acceptable limits. This study develops a double Q‐learning–based controller for optimising daily dosing of a three‐drug regimen consisting of cisplatin, docetaxel and irinotecan. A pharmacokinetics–pharmacodynamics (PK/PD) tumour model with eight resistance states is used as the simulation environment. The controller aims to minimise tumour burden whilst enforcing strict toxicity constraints aligned with clinical dosing guidelines. Simulation results show that double Q‐learning substantially outperforms classical Q‐learning, achieving near‐complete tumour suppression within the simulation framework, corresponding to a residual tumour fraction on the order of (approximately six orders of magnitude reduction) whilst maintaining toxicity within predefined constraints. Robustness analyses under physiological parameter variations of up to and under abrupt disturbance events further demonstrate that the double Q‐learning policy preserves stable closed‐loop behaviour within the simulation environment and exhibits strong resilience to uncertainty. Overall, the results indicate that double Q‐learning provides a proof‐of‐concept framework for adaptive chemotherapy optimisation, with potential for future investigation in reinforcement learning‐based chemotherapy optimisation frameworks.

More from our Archive