DOI: 10.1145/3816980 ISSN: 2573-0142
One Question, Four Voices: How Advice for Alzheimer's Caregiving Differs Between Caregivers, Clinicians, and Large Language Models
Congning Ni, Jinkyung Katie Park, Yang Li, Sarvech Qadir, Sean S. Huang, James S. Powers, Julia A. Hiner, Vania Leung, Patricia Commiskey, Lijun Song, Qingxia Chen, Pamela J. Wisniewski, Bradley A. Malin, Zhijun Yin
Online caregiver forums are a critical support infrastructure for Alzheimer’s disease and related dementias (ADRD), where people seek both emotional reassurance and practical guidance. Prior CSCW and HCI work has examined peer and clinician support in these communities and, separately, explored support-oriented chatbots and AI-assisted helping; however, we still lack systematic evidence on how AI responses compare with peer and clinician roles when answering the same caregiver questions in situ. We address this gap through a systematic, side-by-side comparison of responses to 85 real caregiver questions from
ALZConnected
across four responder groups:
peer caregivers
,
physicians
,
ChatGPT
, and a retrieval-augmented
CareGPT
grounded in peer discussions. We evaluate responses using 28 measures spanning linguistic form (psycholinguistic cues, structure, and style), emotional support, informational support, and topic scope, including LLM-based scoring for selected measures that we audit and calibrate against double-coded human ratings. Across analyses, we observe a consistent division of communicative labor: peers provided time-anchored narratives and community grounding; clinicians provided concise, advice-forward, clinically bounded guidance; and LLMs produced a formal, verbose, highly structured “consultation” voice. Notably, CareGPT shows minimal measurable separation from ChatGPT, indicating that retrieval-plus-prompt-injection alone is insufficient to shift the model’s communicative stance. Quantitatively, LLM outputs score higher than peers on politeness and emotion-support markers, and are comparable to clinicians on informativeness and physician-alignment measures; yet, they are less narrative and less contextually situated. Topic modeling further shows that LLM responses concentrate narrowly on advice-oriented themes, while peers cover a broader range (e.g., emotional processing, legal/financial concerns, and daily cares). These insights show that AI systems can complement, but not replace, human contributions to ADRD caregiver support. We advocate for hybrid systems that combine AI’s structured delivery with the distinct contributions of human voices to strengthen emotional and informational caregiving support infrastructures, while future work should examine whether similar patterns hold in other caregiving domains.