CoCoRV-nf: a powerful and cost-effective tool for rare variant analysis leveraging external biobank sequence data identified new candidate predisposition genes in amyotrophic lateral sclerosis and neuroblastoma
Saima Sultana Tithi, Johnathan Cooper-Knock, Michael Benatar, Joanne Wuu, J Paul Taylor, Gang Wu, Wenan ChenAbstract
Although sequencing costs have steadily decreased with advances in technology, they remain high for large scale studies. The design of traditional individual-disease sequencing studies is either case only or cases with relatively few controls, resulting in potential loss of statistical power for discovery of disease associated genes. Here we show that for a given number of sequenced cases, a large control sample size is critical to maximize power for rare variant burden analysis. Furthermore, we have developed an end-to-end workflow based tool (CoCoRV-nf) to facilitate the use of external biobank sequence resources as controls. The modules include consistent variant QC, variant annotation, ancestry population prediction, and gene based burden analysis using summary genotype information, and combined analysis from multiple independent results. The tool supports exomes and genomes from gnomAD and All of Us as controls with preprocessed datasets. We apply the tool in two rare neurological diseases: amyotrophic lateral sclerosis and neuroblastoma. For each disease, two case cohorts are paired with gnomAD and All of Us data, respectively, followed by a combined analysis. Not only did we recapture known genes, but also, we identified new candidate genes for both diseases. By leveraging multiple large external biobank sequence data, we demonstrate the feasibility of using our tool to maximize statistical power to identify new disease predisposition genes.