DOI: 10.1002/qub2.70058 ISSN: 2095-4689

The evaluation crisis in artificial intelligence‐driven computational biology: Pitfalls and solutions

Tsung Fei Khang, Joshua W. K. Ho, Lin Hou, Jiangning Song, Limsoon Wong

Abstract

Artificial intelligence (AI) is rapidly transforming computational biology. However, closer scrutiny reveals that many reported results do not hold under rigorous evaluation. This paper identifies prevalent, recurring issues in the field, including narrow datasets, hidden confounders that make supposedly independent test data effectively indistinguishable from the training data (the doppelgänger effect), misleading performance metrics, and over‐reliance on overly easy test cases. We argue that meaningful assessment must go beyond metric‐centric model evaluation to encompass biological realism, data diversity, confounding control, and interpretability. Crucially, AI studies in computational biology should deliver not only predictive performance but also genuine biological insight and transferable methodological understanding. To address these systemic limitations, we propose a unifying framework, VALID, comprising the following five guiding principles for rigorous AI in biology: Veracious biological grounding and novelty, artifact control and methodological rigor, legitimate generalization, inspectable preprocessing, and deployment‐aware performance. Together, these principles serve as a diagnostic lens for structuring our discussion throughout this paper. We further illustrate common pitfalls through representative examples and provide actionable guidance for authors, reviewers, and editors aimed at improving the rigor and scientific value of AI‐driven computational biology.