Information quality, transparency, and readability of generative AI chatbot responses to parenting practice questions for infants and toddlers aged 0–3 years: A cross-sectional comparative study
Xiao-Na Sun, Yi Liu, Xiu-Fen Wu, Yan Kong, Zhi-Hui Li, Gui-Ling YuBackground
Generative artificial intelligence (AI) chatbots are increasingly used by caregivers to obtain parenting information, but the quality, transparency, and readability of their responses remain uncertain.
Objective
To compare responses generated by six publicly available chatbots to parenting practice questions for infants and toddlers aged 0–3 years.
Methods
In this cross-sectional comparative study, 53 standardized English questions across five domains were submitted in single-turn conversations to ChatGPT, Gemini, Perplexity, Copilot, DeepSeek, and Doubao through their official web interfaces in Qingdao, China, from 4 to 14 March 2026. The first complete response to each question was evaluated using DISCERN, Ensuring Quality Information for Patients (EQIP), the JAMA benchmark criteria, the Global Quality Score (GQS), and six readability indices. Differences were assessed using Friedman tests and paired Wilcoxon signed-rank tests with Benjamini–Hochberg correction.
Results
A total of 318 responses were analyzed. Copilot generally achieved the highest information-quality and transparency scores, whereas Doubao scored lower on DISCERN, EQIP, and GQS. JAMA scores were low across all systems. ChatGPT had the lowest reading burden on most grade-level indices, while DeepSeek had the highest Flesch Reading Ease Score. No system met the prespecified readability benchmarks.
Conclusions
Publicly available chatbots may provide preliminary parenting information; however, their limited source transparency and high reading burden indicate that they should be used only as supplementary tools and should not be relied upon as independent or primary sources of parenting health advice.