How to Train Your Chatbot: Information‐Theoretic Foundations of Diagnostic Questioning in Inborn Errors of Immunity
Saul O. Lugo Reyes, Estefanía Vásquez Echeverri, Juan Carlos Bustamante Ogando, Lina M. Castano‐Jaramillo, Natalia Vélez Tirado, Edna Venegas Montoya, Alejandro Tarango García, Héctor Gómez Tello, Alejandro Palma, Matías Oleastro, Selma Cecilia Scheffler Mendoza, Sara Elva Espinosa‐Padilla, Eduardo Guaní Guerra, Hanadys Ale, Marco A. Yamazaki‐Nakashimada, Kathleen E. Sullivan, Chiharu MurataABSTRACT
Background
Navigating the more than 550 inborn errors of immunity (IEI) requires efficient diagnostic reasoning. Information theory suggests questions should be prioritized by their capacity to reduce diagnostic uncertainty (entropy); yet whether experts or large language models (LLMs) optimize for information gain remains unquantified.
Objective
We compared expert clinician and LLM diagnostic prioritization strategies using an information‐theoretic framework.
Methods
Fifteen immunologists and six LLMs (ChatGPT, Claude, Gemini, Grok, DeepSeek, and Llama) ranked 35 diagnostic questions by efficiency. Shannon's entropy was used to estimate expected information gain (EIG) for each question. Agreement was assessed via Spearman correlations, consensus ranking, and principal components analysis (PCA).
Results
Clinician consensus rankings strongly correlated with estimated information gain (Spearman ρ = −0.71, p < 0.001). “Age at onset?” ranked first by clinicians, provided the highest information gain (2.29 bits), reducing diagnostic uncertainty by 80%. Clinicians and LLMs showed strong agreement on top‐tier discriminators (Spearman ρ = 0.73, p < 0.001). However, PCA revealed a distinct LLM cluster; clinicians prioritized bedside/history questions, whereas LLMs favored syndromic and laboratory features. Optimal questioning reached diagnostic confidence in 4–5 steps, approaching the theoretical minimum.
Conclusions
Expert clinicians implicitly approximate information‐theoretic optimization in IEI diagnostics. While LLMs share a core heuristic for high‐yield questions, divergence in mid‐sequence reasoning suggests a shift from experiential heuristics to probabilistic data‐matching. This framework provides a principled basis for training AI‐assisted tools that mirror expert diagnostic logic.