Accuracy and Safety of Language Model-Generated Breastfeeding Counseling Responses: An Expert-Based Comparative Evaluation of ChatGPT-3.5, ChatGPT-4, and BreastfeedGPT
Seda Serhatlıoğlu, Yeşim YeşilAim:
This study aims to compare the responses generated by ChatGPT-3.5, ChatGPT-4, and BreastfeedGPT, a customized domain-specific GPT, to frequently asked questions related to breastfeeding counseling.
Design and Methods:
Ten breastfeeding-related questions were selected based on international guidelines and expert validation. Each model generated responses to the same questions, which were anonymized and evaluated by eight health professionals with expertise in breastfeeding counseling using structured Likert scales. Evaluation dimensions included (1) scientific accuracy and scope, (2) tone, language, and motivation, and (3) reliability via mDISCERN scoring. Friedman and Wilcoxon signed-rank tests with Bonferroni correction were used to determine statistical significance.
Results:
BreastfeedGPT model outperformed ChatGPT-3.5 and ChatGPT-4 across all domains. It received the highest mean scores in scientific accuracy (3.72 ± 0.23) and motivational tone (3.78 ± 0.42), with statistically significant differences (
Conclusion:
BreastfeedGPT may offer advantages over general-purpose AI tools in providing breastfeeding-related information. However, because human expert responses were not evaluated in parallel, its role as a complementary resource should be interpreted cautiously and confirmed in future direct comparisons. Such tools should support, rather than replace, professional breastfeeding care.