DOI: 10.1093/bioadv/vbag224 ISSN: 2635-0041

Machine-learning-based analysis of host-depleted k -mer profiles for early detection and abundance estimation of coffee pathogens

Seunghyun Lim, Ezekiel Ahn, Dapeng Zhang, Lyndel W Meinhardt, Sunchung Park

Abstract

Motivation

Early detection of plant pathogens is essential for timely disease management, but diagnosis at low infection levels remains difficult because pathogen-derived sequences are often masked by abundant host DNA. K-mer-based machine learning offers a potentially sensitive, alignment-free feature-based approach for detecting infection and estimating pathogen abundance from sequencing data, but its performance across infection levels, biological backgrounds, and preprocessing strategies remains insufficiently characterized.

Results

We evaluated a k-mer-based machine-learning framework using simulated Illumina short-read data from two Coffea arabica cultivars (ET39 and Typica) infected with Hemileia vastatrix or Fusarium xylarioides across a broad range of infection rates. Among ten models tested on raw and host-depleted k-mer profiles, logistic regression (LR) and linear support vector machine (LSVM) were the most sensitive. Host depletion with KrakenUniq substantially improved low-infection-rate detection, lowering the practical detection threshold to 0.05%. Models trained at lower infection rates performed better when tested across different infection rates than models trained at higher rates, although performance declined when the test infection rate fell below 0.05% because infected samples were increasingly misclassified as healthy. External validation showed limited biological transferability: low-infection-rate performance depended strongly on host background, whereas high-infection-rate performance depended more on pathogen identity. To improve weak-signal detection, we developed a two-stage framework combining elastic-net logistic regression for detection with ridge regression for infection rate prediction. This framework maintained near-perfect detection at 0.05% and above, improved detection at 0.01% under mixed-infection-rate training, and showed strong agreement between true and predicted infection rates (Spearman’s ρ = 0.95). Predictive k-mers were extensively shared between LR and LSVM and showed distinct compositional differences between healthy- and infected-associated features.

Availability and implementation

Code is available at GitHub (https://github.com/TropicalBreeding/-coffee-pathogen-kmer-ml)

More from our Archive