DOI: 10.3390/app16167904 ISSN: 2076-3417

The HEART Framework for LLM-Enabled Socially Assistive Robots in Healthcare: A PRISMA-Informed Structured Review

Tihomir Orehovački

Large language models (LLMs) are expanding the capabilities of socially assistive robots (SARs) through natural dialogue, personalisation, multimodal reasoning, retained interaction context, and adaptive behaviour in healthcare. Integrating generative language models into robots, however, complicates evaluation because fluent output may exaggerate perceived competence and increase the risks of hallucination, overtrust, privacy exposure, relationship dependency, and unsafe reliance on advice or actions. This PRISMA-informed review synthesises healthcare robotics, human–robot interaction, LLM-enabled systems, ethics, implementation, and care delivery. Database searches returned 128 records, of which 110 were unique after deduplication. Supplementary retrieval and assessment yielded 85 substantive sources spanning background mapping, primary analysis, and governance. Studies focused mainly on feasibility, usability, acceptability, dialogue quality, and short-term engagement, whereas longitudinal safety, governance of retained interaction context, comparative effectiveness, workflow integration, and sustained healthcare value received limited attention. These gaps indicate that evaluation of LLM-enabled SARs must account for physical presence, social role, interaction memory, and potential actions rather than focus on conversational performance alone. The review therefore proposes HEART, a healthcare-specific evaluative architecture comprising Human-Centred Communication, Ethical and Trustworthy Deployment, Adaptive and Embodied Intelligence, Relationship Continuity, and Translational Healthcare Value. HEART uses boundary rules, operational indicators, qualitative labels, and non-additive deployment gates to separate evaluative domains, define assessable outcomes, summarise reported support, and prevent strengths in one area from masking critical safety or governance failures. Future research should validate HEART through longitudinal and comparative assessment of hallucination severity, language-to-action safety, long-term effects, equity, and post-deployment monitoring.

More from our Archive