Computational Pipeline Reveals Nature’s Untapped Reservoir of Halogenating Enzymes
Judit Szenei, Ashleigh Burke, Anne Liong, Aleksandra Korenskaia, April L. Lukowski, Nadine Ziemert, Pablo I. Nikel, Pedro N. Leão, Bradley S. Moore, Tilmann Weber, Kai BlinAbstract
Microbial halogenated natural products (hNPs) hold ecological, agricultural, and biomedical relevance. The hNP-producing potential of an organism can be assessed by the precise prediction of halogenating enzymes, yet detailed annotations of halogenases are often missing from genomic and metagenomic data. We created a manually curated database (https://halogenases.secondarymetabolites.org/) containing information on the halide specificity, role, and position of verified catalytic residues and the results of mutagenesis studies of more than 120 experimentally validated or in silico inferred halogenases. The collection of experimental data supports a computational pipeline that allows family-, substrate-, and halide-scope-level annotation of halogenating enzymes by relying on functionally important residues, conserved motifs, and profile hidden Markov models (pHMMs). Our analysis with sequence similarity networks (SSNs) highlighted several underexplored clusters in the UniRef50 database. We further investigated a cluster of vanadium-dependent haloperoxidases because a halogenase from Rhodopirellula baltica (RhobaVHPO), previously labeled as a hypothetical chloroperoxidase, clustered apart from the known chloroperoxidases and bromoperoxidases. The monochlorodimedone assay confirmed the chlorination activity of RhobaVHPO and showed its preference for bromide. Our database and workflow provide extensive and scalable solutions for the systematic and precise annotation of halogenating enzymes in genomic and metagenomic data sets. The in-depth categorization of halogenases will improve the chemical structure prediction of microbial hNPs, supporting ecological assessments and natural product discovery.