DOI: 10.3390/genes17101173 ISSN: 2073-4425

A Frozen Plant DNA Language Model for Computational Prioritization and Functional Annotation of Melon (Cucumis melo) Variants

Guosheng Sun, Yanping Wei, Entong Li, Mengting Xiao, Zhilin Zhang, Zhenchao Zhang, Zhongliang Dai, Zhihu Ma, Changwei Zhang

Background/Objectives: Fruit quality in melon is determined by complex traits including sugar accumulation, aroma, flesh color, ripening behavior and disease resistance. Conventional genomic selection provides limited mechanistic interpretation, especially for regulatory and de novo variants. We adapted the EVEE interpretable embedding–probing framework to melon using PlantCAD2 as a frozen DNA language model backbone. Methods: The model was not fine-tuned or further pre-trained. Instead, lightweight probes were trained on fixed embeddings derived from sequence windows surrounding genetic variants. The pipeline integrated supervised variant-effect prediction, functional perturbation profiling, training-free masked log-likelihood ratio scoring, and cis-regulatory element generation. We applied the framework to 32,268 public genotyping-by-sequencing (GBS) SNPs from a melon diversity panel. A total of 30,052 variants were successfully encoded and analyzed. Results: Of these, 31.9% were classified as high predicted functional effect under the GWAS-proximal probe. We adopted a rigorous independent split based on chromosome and locus blocks to ensure test variants and training variants occupy distinct linkage-disequilibrium blocks. In this setting, the covariance probe yields an AUROC of approximately 0.51 with a 95% confidence interval covering 0.50, indicating no predictive advantage over random guessing. The full-cohort AUROC value of 0.748 reported earlier is attributed to linkage-disequilibrium leakage and overlapping positive samples between datasets, and thus does not reflect true independent generalization. The 17-class structural annotation probe reached 95.97% per-position accuracy. Functional perturbation signals were enriched in intronic, coding, and proximal promoter regions, consistent with known regulatory architectures in melon. Using the masked language model head with Gibbs sampling, we generated 76 candidate 400 bp cis-regulatory elements associated with sugar metabolism pathways. A posterior consistency check against 18 published GWAS/QTL loci showed effect-score patterns broadly concordant with reported loci; this is a consistency analysis against known loci, not an independent prospective validation. Variants in the CmTST2 region exhibited elevated mean effect scores of 0.7922 compared with the genome-wide background of 0.406. Conclusions: Overall, this study provides a computational, hypothesis-generating framework in which a frozen plant DNA language model combined with lightweight probes prioritizes melon variants and generates candidate regulatory sequences; all outputs are unverified in-silico predictions requiring experimental validation before any functional or breeding use. Because every variant is scored independently from its local sequence context, the framework necessarily treats genetically linked, network-dependent traits as single-locus effects and does not model epistasis, pathway-level interactions, or genotype-by-environment effects; the outputs are sequence-level priors, not phenotype predictions. The results support a generalizable framework for model-guided functional genomics and precision breeding in Cucurbitaceae crops.