DOI: 10.1093/bioadv/vbag222 ISSN: 2635-0041

AmpliPhy improves gene trees by adding homologous sequences without affecting alignments

Dongwook Kim, Manuel Gil, Kazutaka Katoh, Christophe Dessimoz

Abstract

Motivation

In phylogenomics, gene tree reconstruction depends on multiple sequence alignment and tree inference, and ongoing work continues to improve inference quality. Denser taxon sampling has been associated with improved gene tree inference, suggesting that adding homologs could be a practical route to higher accuracy as sequence databases continue to expand. However, adding sequences can influence multiple steps of typical inference pipelines, and little is known on its specific effect on the multiple sequence alignment, tree reconstruction, and rooting steps.

Results

We performed a large-scale empirical and simulated benchmarks to quantify how homolog enrichment affects alignment and phylogenetic inference. Using an enrichment–impoverishment design and a measure of tree accuracy based on taxonomic congruence, we found that enrichment consistently improves tree inference quality, while effects on alignment quality are marginal. We show that this improvement is associated with, but not restricted to accurate root placement on enriched trees when sensitive homolog search is accompanied. Notably, much of the benefit can be retained with relatively compact alignments produced by sequence addition. Building on these observations, we provide a tool, AmpliPhy, which efficiently improves phylogenetic reconstruction of protein families through homolog enrichment.

Availability and Implementation

The AmpliPhy open-source pipeline software is available at https://github.com/DessimozLab/ampliphy. The scripts used to generate the data and figures are available from https://github.com/DessimozLab/ampliphy-analysis.

More from our Archive