Measuring the Impact of AI-Driven Well-Being Apps—Instrument Development and Pilot Evidence from the Malu Prototype
Sarah Hatfield, Jeanette TammThe aim of the present study was to develop an instrument that enables evaluation of AI-based mental health apps, which are promising digital interventions for promoting psychological well-being. The instrument was used to conduct an initial evaluation of an early pilot stage of the well-being app MALU. As part of a non-representative hypothesis-testing longitudinal study, N = 11 participants aged 18 to 34 used the app over a period of two weeks. The participants were surveyed at three points regarding perceived stress (Perceived Stress Scale), sleep problems (short version of the Insomnia Severity Index), and chatbot usability (Chatbot Usability Scale). The results showed a significant decrease in perceived stress between the first and third measurement points (Z = −2.31, p = 0.01), as well as for perceived sleep problems between the second and third measurement points (Z = −1.86, p = 0.03). Perceived chatbot usability increased significantly over the course of the study (Z = 2.37, p = 0.01). The results suggest potential effectiveness of the app in reducing stress and sleep problems as well as an improvement in the user experience regarding the chatbot interaction over time. The evaluation instrument proved suitable for use in early development phases.