DOI: 10.3390/nu18193205 ISSN: 2072-6643

Nutrient Calculation Accuracy, Expert-Rated Clinical Appropriateness and Safety Indicators, and Reproducibility of Large Language Model-Generated Diet Plans for Polyendocrine Metabolic Ovarian Syndrome: A Real-Case-Based Evaluation

Ayşenur Çalık, Pınar Ece Demiray, Ayşe Betül Bilen, Gülen Ecem Kalkan, Muazzez Garipağaoğlu

Background/Objectives: This study evaluated the nutrient calculation accuracy, reproducibility, and expert-rated clinical appropriateness and safety indicators of three-day diet plans generated for women with polycystic ovary syndrome (PCOS), recently proposed to be termed polyendocrine metabolic ovarian syndrome (PMOS). Methods: Anonymized data from six women with heterogeneous PMOS profiles, purposively selected from a single clinic, were converted into standardized Turkish prompts and submitted to ChatGPT-4, Claude Sonnet 4.6, Gemini 2.5 Flash, and Mistral 2.0 through free web interfaces. Three independent generations per case–model combination yielded 216 daily menus. LLM-reported nutrient values were compared with BeBiS 9.0 calculations. Reproducibility was assessed using intraclass correlation coefficients (ICCs). Two experts independently evaluated the first generated plan for each case–model combination using a study-specific clinical appropriateness rubric and red-flag checklist. Results: No model achieved consistently low error across nutrients. Claude Sonnet 4.6 showed the lowest percentage errors for energy, carbohydrate, and fiber; ChatGPT-4 for protein; and Mistral 2.0 for fat. Gemini 2.5 Flash showed the largest percentage errors across all nutrients and a mean energy bias of +531.1 kcal. ChatGPT-4 showed comparatively higher plan-level reproducibility, while the remaining models generally showed poor reproducibility. Claude Sonnet 4.6 had the highest descriptive clinical appropriateness scores and Gemini 2.5 Flash the lowest, but no pairwise differences remained significant after Holm correction. Gemini 2.5 Flash generated the most confirmed study-specific red flags (n = 28), with all six cases meeting the criterion for the “masked low-energy” indicator. Conclusions: LLM-generated diet plans showed model-dependent calculation errors and limited reproducibility, and the present findings do not support their independent or unsupervised use as individualized clinical nutrition recommendations.