DOI: 10.1515/sagmb-2025-0078 ISSN: 2194-6302

A hybrid deep learning framework for WT or mutant peptide prediction using p53 mutation data

Manisha R. Patil, Anand Bihari

Abstract

p53 is a tumor suppressor protein that maintains genome integrity. Single amino acid variations (SAVs), especially hotspot mutations (such as R175 and R248), are closely associated with oncogenic transformation and weaken DNA-binding and transcriptional regulatory capabilities. To identify deleterious SAVs for understanding disease mechanisms and key cellular processes, including apoptosis, DNA repair, and cell cycle regulation. This study used quantitative biochemical and microbial descriptors of p53 peptide sequences, along with a CNN and a 2-layer bidirectional long short-term memory (Bi-LSTM) network architecture with attention mechanisms and ESM-2-based embeddings, to classify sequences as wild-type or mutant. A systematic analysis of the p53 mutation dataset was performed using molecular weight, instability index, hydrophobicity, motif enrichment, and amino acid substitution patterns to identify mutation-centered peptides. Using stratified 5-fold cross-validation, the proposed model was tested, achieving an accuracy of 0.98, an area under the ROC curve (AUROC) of 0.98, and a precision-recall AUC of 0.99. Furthermore, SHAP-based interpretation identified the key amino acid residues and biochemical factors that contribute to the model’s predictive performance. The proposed approach provides insights into the structural, sequential, and biochemical effects of variants, an interpretable, robust framework for evaluating the functional consequences of p53 hotspot and other variants, and a computationally efficient tool for prioritizing high-risk p53 variants.

More from our Archive