DOI: 10.1177/20552076261487359 ISSN: 2055-2076

Symptom checkers: Gaps in their evidence base and suggestions for accuracy evaluation

Norbert Donner-Banzhoff, Helmut Schaefer, Nadine Schlicker, Annette Becker, Joerg Haasenritter, Veronika van der Wardt

Symptom checkers (SCs) are mobile healthcare applications for the general population. Based on symptoms entered by the user, SCs provide a list of likely diagnoses and/or triage advice. Studies conducted so far show limited accuracy of current SCs with designs carrying a high risk of bias. In this narrative review we discuss three major methodological shortcomings of previous studies: 1) wrong study setting, 2) inappropriate reference standards, and 3) too small samples. To reflect the home setting, participants of an accuracy study must be prospectively recruited to document symptomatic episodes and health care used. Studies based on vignettes or samples drawn from healthcare facilities are not sufficient. A panel reviewing follow-up data must determine diagnoses for each episode and appropriate (non-)utilization of healthcare (delayed-type reference standard). Even when outcomes are collapsed into four groups, the sample size requirements for investigating triage appropriateness are considerable (n=62,000). The severity and treatability of rare conditions must be accommodated. Because of the defensive tendency of available SCs (with sensitivity maximized), the specificity will necessarily be limited to 20 to 30% at most, irrespective of the findings of an empirical study. We discuss the prospects for technological improvements, such as Large Language Models. For investigations of their accuracy, the same design requirements apply as for established models. Accuracy evaluation should be guided by the function of the intervention under consideration independently of the technology used. Given the limitations of our knowledge regarding SCs’ accuracy, their place in the health care system is far from clear.