DOI: 10.3390/bioengineering13101098 ISSN: 2306-5354

A Bias-Aware Analytic Framework for Real-World Data with AI-Audited Workflows: Application to Opioid Ordering in Pediatric Oncology

Tricia Morphew, Hannah Pease, Michelle A. Fortier, Zeev N. Kain, Lois W. Sayrs

Background: The growing use of computational modeling of real-world data (RWD) in clinical research introduces significant risks for data scientists and AI trainers already managing inconsistency, unexamined bias, and analytical opacity. Building upon more than 600 guideline-based checklists and pre-planning templates for data quality and transparency reporting in observational RWD, we propose a run-time operational execution loop to embed reproducibility analytics within the workflow. Our approach acts as a run-time gatekeeper mapping prespecified analytics goals to a closed-loop LLM semantic code auditor, checking for both logic, bias, and vulnerabilities during code execution. Our framework was designed to bridge the gap between passive disclosure, ensuring proposed requirements, and post-hoc validation. Our structured hierarchical approach coupled with AI auditing is the first phase in the development of full-stack automation of bias-aware RWD analytic validation. Methods: Our analytic framework spans 10 domains across three analytic phases. Each domain includes prespecified criteria governing verification with potential biases, mechanisms, and mitigation strategies systematically mapped throughout. Phase I entails specification (Domains 1–4: objectives and estimands, data inputs and outputs, cohort and filters, and variable coding); Phase II: modeling (Domains 5–7: model testing, diagnostics, and reasonableness checks); and Phase III: robustness and interpretation (Domains 8–10: missing data handling, temporal and stratified checks, and scope boundaries). We utilize specific prompt-directed large language model (LLM) auditing of model elements to ensure data quality. The methodology is illustrated with examples from an empirical characterization of inpatient opioid data from the Cerner Real-World Data (2014–2021). GLMMs and corresponding outputs were generated in R v4.4.2. Curated AI-auditing of first-pass code drafting was performed using ChatGPT 5.2 (OpenAI), with cross-LLM portability assessed using Claude Sonnet 5 (Anthropic). Results: The Cerner RWD database contained 16,734 pediatric cancer patients (76,395 encounters) across 55 U.S. health systems. During the specification phase, total daily dose could not be estimated with sufficient validity from the Cerner RWD, and “as-needed” (PRN) orders, warranted empirically supported exclusion. The modeling phase revealed residual variability in opioid ordering attributable primarily to patient-level differences within health systems (ICC = 19.9%) rather than between health systems (ICC = 2.0%). Temporal analysis revealed divergent opioid-ordering trajectories from 2017 onward, with decreasing morphine and increasing fentanyl ordering. We then utilized curated prompting to enforce an AI-generated coding requirement to produce valid likelihood ratio tests across the prespecified model sequence accounting for temporal trends. After harmonizing the model specification and estimation settings, both LLM workflows produced identical random-effect variance estimates, ICCs, and AIC values evaluated through the third decimal place, with corresponding fixed-effect estimates likewise matching. Conclusions: The 10-domain framework provides a structured, reproducible mechanism for governing RWD analyses by using a closed-loop semantic code auditor to embed run-time analytics directly into the data science workflow, supporting transparent exposure definition, systematic bias assessment, and prespecified analytic decision-making. By bridging the gap between passive transparency checklists and active automation, the framework operationalizes existing data quality guidelines through an executable, bias-aware code-auditing process. Cross-LLM concordance in identifying problematic logic and reproducing complex workflows supports the portability of this run-time gatekeeper. Ultimately, this approach establishes a scalable pathway for embedding expert statistical decision-making into the automation of safe, reliable real-world evidence (RWE) generation.