DOI: 10.1002/mp.70592 ISSN: 0094-2405

Reference‐Free large language model agents for physician‐guided radiotherapy treatment planning

Dongrong Yang, Xin Wu, Yibo Xie, Xinyi Li, Qiuwen Wu, Q. Jackie Wu, Yang Sheng

Abstract

Background

Large language models (LLMs) have recently demonstrated exceptional capabilities, offering the potential to streamline workflows and enhance efficiency across diverse tasks. However, their application in domain‐specific areas, such as radiotherapy treatment planning, remains challenging due to the lack of publicly available, specialized knowledge required for these tasks.

Purpose

To investigate the critical architectural and functional components necessary to adapt an off‐the‐shelf large language model into a clinically viable planning agent for inverse treatment planning in intensity‐modulated radiation therapy (IMRT).

Materials/methods

Twenty head‐and‐neck (HN) cancer patients who received IMRT at our institution were retrospectively collected under IRB approval. The LLM agent was implemented to directly interact with the clinical treatment planning system (TPS) to iteratively extract intermediate plan states and propose new constraint values to guide inverse optimization. Its decision‐making was informed by real‐time plan evaluations and prior optimization outcomes, enabling adaptive refinement of planning strategies across iterations. The agent was equipped with three core capabilities: comprehension of clinical planning objectives, contextual understanding of the optimization environment, and arithmetic proficiency for quantitative reasoning. Treatment planning was conducted in a reference‐free inference setting, wherein the LLM operated without prior exposure to manually generated treatment plans and without any fine‐tuning or task‐specific training. LLM‐generated treatment plans were compared with clinically approved plans created by certified dosimetrists. Key dosimetric endpoints were evaluated and statistically analyzed. Paired comparisons between LLM‐generated and clinical plans were performed using the Wilcoxon signed‐rank test.

Results

LLM‐generated plans achieved comparable organ‐at‐risk (OAR) sparing relative to clinical plans, while demonstrating improved hot spot control ( D max : 106.5% vs. 108.8%, p  < 0.05) and superior conformity (conformity index: 1.18 vs. 1.39, p  < 0.05 for boost PTV; 1.82 vs. 1.88, p  = 0.47 for primary PTV).

Conclusions

This study demonstrates the feasibility of a reference‐free, LLM‐driven workflow for automated IMRT treatment planning in a commercial TPS. The proposed approach provides a generalizable solution that could reduce planning variability and support broader adoption of AI‐based planning strategies.

More from our Archive