From Medical Records to AI-Ready Datasets: A Practical Guide for Clinical Researchers
Catalin Anghel, Andreea Alexandra Anghel, Marian Viorel Craciun, Simona Moldovanu, Adina Cocu, Diana-Elena Vulpe, Calina Maier, Vasile Potop, Christiana Diana Maria Dragosloveanu, Constantin Adrian Andrei, Serban Dragosloveanu, Cristian ScheauBackground: Medical artificial intelligence (AI), machine learning (ML), and deep learning (DL) studies frequently begin with datasets collected for routine care rather than for computational modeling. Such datasets may contain inconsistent variables, heterogeneous measurement time points, unexplained NaN values, poorly defined outcomes, missing metadata, and insufficient documentation, which can compromise model development before any algorithm is selected. Methods: This Technical Note proposes a physician-facing Clinical AI-Readiness Guide for preparing medical datasets before AI-based analysis. The guide was developed as a practical framework organized around pre-modeling decisions, including the clinical task, cohort, minimum common dataset, outcome definition, predictor variables, measurement timing, missing-data logic, standardization, non-tabular data linkage, data dictionary, and validation readiness. Results: The proposed guide translates AI-readiness principles into concrete data-collection rules for clinical, laboratory, imaging, physiological-signal, textual, follow-up, and multimodal data. It emphasizes clinically consistent data acquisition, reliable target labeling, explicit missing-data logic, patient-level linkage, structured metadata, and validation feasibility. A structured checklist and scoring approach are also proposed as practical pre-modeling assessment tools to classify datasets as not ready, exploratory only, ML-ready with limitations, or AI-ready for model development. Conclusions: Medical AI-readiness should be established before model development begins. By helping physicians collect, structure, and document data more consistently, the proposed guide may improve collaboration between clinical and technical teams and reduce preventable dataset-related failures in medical AI research.