Systematic prioritization of candidate genes in camptothecin biosynthesis using multi‐omics and deep learning
Shenqi Wang, Xing Wu, Maria Moreno, Joonseok Oh, Stephen L. Dellaporta, Jason M. Crawford, Farren J. IsaacsAbstract
Camptothecin (CPT), a plant‐derived monoterpene indole alkaloid first identified in Camptotheca acuminata , is a drug precursor widely used for cancer chemotherapeutics. However, the full set of genes responsible for CPT biosynthesis remains unclear, hindering efforts to elucidate the complete pathway or establish biosynthetic production of CPT in heterologous hosts. In this study, we engineered an experimental callus system for inducible production of CPT, which enabled multi‐omics and deep learning analyses to identify candidate genes in CPT biosynthesis. We first generated an improved genome assembly and gene annotation for C. acuminata . We then leveraged the natural variation of CPT levels in C. acuminata tissues and performed transcriptomic analysis of multiple callus and tissue types to shortlist candidate enzymes responsible for CPT biosynthesis. Finally, we conducted large‐scale deep learning–enabled protein–ligand complex structure prediction to prioritize 117 candidate enzymes for studies that map their roles in CPT biochemical reactions. By integrating experimental, genomic, transcriptomic, and deep learning approaches, this study provides a valuable foundation for the complete elucidation of the CPT biosynthetic pathway.