DOI: 10.1111/xen.70171 ISSN: 0908-665X

Reference Composition Biases Automated Cell Annotation and Obscures Biological Signals in Xenotransplantation Single‐cell Atlases

Lisha Mou, Ying Lu, Zijing Wu, Qishan Zhang, Zuhui Pu

ABSTRACT

Artificial intelligence is transforming single‐cell biology, but automated annotation depends not only on model architecture but also on reference composition. Whether compositional biases within reference datasets influence AI‐assisted cell identity assignment remains poorly understood. Here, we establish a harmonized single‐cell resource integrating 254 245 cells from five datasets of genetically engineered pig kidney grafts in human and non‐human primate (NHP) recipients. We evaluate representative annotation strategies, including the foundation model TranscriptFormer, the autonomous biomedical agent Biomni, and conventional approaches. Across frameworks, reference composition, represented as the empirical cell‐type composition, emerged as a major factor associated with inferred cell identity. In donor‐pig kidney grafts, abundant collecting‐duct (CD) references promoted widespread CD assignment and suppressed distal tubular identities. Under the empirical reference composition, 97.0% of cells carrying target‐native thick ascending limb (TAL) annotations and 97.2% of cells carrying target‐native distal convoluted tubule (DCT) annotations in a primary pig‐to‐NHP kidney xenograft dataset were reassigned to the CD compartment, revealing profound reference‐driven label collapse. This distortion extended to compartment‐level analyses, obscuring biological programs. Notably, an endothelial interferon‐stimulated gene signature in the primary target remained detectable under target‐native annotation but was masked by reference‐driven label transfer. We establish a harmonized xenotransplantation single‐cell resource and identify reference‐composition auditing as a general framework for evaluating AI‐assisted biological interpretation.