Long‐Read Metabarcoding With Optimized Bioinformatic Analysis Outperforms Short‐Read Metabarcoding for Assessing Soil Protist Diversity and Ecology
Aline Adler, David Singer, Sarah Wegmüller, Sarah Scotton, Thierry J. Heger, Alexandre KuhnABSTRACT
Metabarcoding is a powerful tool for assessing microbial community composition in various ecosystems. It involves the amplification of a marker gene followed by high‐throughput sequencing and bioinformatic analysis to identify species. Second‐generation sequencing has democratized biodiversity studies by allowing high‐throughput sequencing of short DNA fragments. Long‐read sequencing now allows for the use of longer markers, potentially offering improved taxonomic resolution. Until recently, this advantage was partly offset by the higher error rate of long‐read sequencing. Here we demonstrate that with current Oxford Nanopore Technologies sequencing (Kit V14 chemistry and R10.4.1 flow cells) combined with an appropriate analysis pipeline, sequencing errors have a negligible impact on the accuracy of long‐read metabarcoding. Using DNA extracted from two protist cultures, we estimated taxonomic misassignment of individual long‐reads at genus level between 0% and 0.01% with long‐read metabarcoding using the 18S rRNA gene and between 0% and 0.01% with short‐read Illumina metabarcoding relying on the V4 region only. We also propose an optimized long‐read clustering procedure that incorporates pre‐sorting reads by quality. When applied to the same cultures, it produced fewer but larger sequence clusters and increased the average similarity to the reference sequences. Applied to environmental DNA extracted from 26 vineyard soil samples, this method identified stronger correlations between protist communities and environmental variables compared to short‐read metabarcoding. Notably, taxonomic assignment of individual long‐reads (without clustering) further increased sensitivity to environmental patterns. These results support the reliability of long‐read metabarcoding and highlight its strong potential for ecological research in general. Broader adoption of this approach may improve the accuracy of biodiversity assessments and in turn, future studies will benefit from the expanded representation of long sequences in public databases.