DOI: 10.1111/tpj.71076 ISSN: 0960-7412

Biology‐informed neural networks learn nonlinear representations from omics data to improve genomic prediction and biological discovery

Katiana Kontolati, Rini Jasmine Gladstone, Ian W. Davis, Ethan M. Pickering

SUMMARY

Traditional genotype‐to‐phenotype models depend heavily on direct mappings that achieve only modest accuracy, forcing breeders to conduct large, costly field trials to maintain or marginally improve genetic gain. Models that incorporate intermediate molecular phenotypes can achieve higher predictive fit, but remain impractical since such data are unavailable at deployment or design time. Biology‐informed neural networks (BINNs) overcome this limitation by encoding pathway‐level inductive biases and leveraging multi‐omics data only during training, while using genotype data alone during inference. Here, we extend BINNs for genomic prediction and selection in crops by integrating thousands of single‐nucleotide polymorphisms with multi‐omics measurements and prior biological knowledge. By directly embedding omics‐derived priors, BINN outperforms conventional models in low‐data ( n < p ) regimes and enables sensitivity analyses that expose biologically meaningful traits. Applied to maize gene expression and multi‐environment field trial data, BINN improves rank correlation accuracy within and across most subpopulations under sparse data conditions and nonlinearly identifies genes that GWAS/transcriptome‐wide association studies may fail to uncover. With complete domain knowledge for a synthetic metabolomics benchmark, BINN substantially reduces prediction error relative to conventional neural nets and correctly identifies the most important nonlinear pathway. Importantly, both cases show that highly sensitive BINN latent variables correlate with the experimental quantities they represent, despite not being trained on them. This suggests that BINNs learn biologically relevant representations, nonlinear or linear, from genotype to phenotype. Together, BINNs establish a framework for improved genomic prediction accuracy and biological discovery that can guide genomic selection, candidate gene selection, pathway enrichment, and gene‐editing prioritization.

More from our Archive