DOI: 10.1002/smtd.70949 ISSN: 2366-9608

Iterative Self‐Supervised Signal Curation Enables High‐Fidelity Peptide Profiling in Parallel Nanopore Sensing

Hailin Pan, Fengqin Luo, Yishuo Zhang, Shizhe Jiang, Wenzhi Wang, Huixian Zeng, Yunfeng Zhang, Ou Wang, Wenbing Qin, Lingxuan Zhang, Dapeng Wang, Junyi Chen, Ji Wang

ABSTRACT

The recent repurposing of nanopore arrays for proteomics generates extensive high‐throughput datasets, but stochastic molecular translocations and sensor heterogeneity inevitably introduce massive non‐informative signal artifacts. Therefore, an automated and unbiased data curation method is essential. Here, we present NanoCurator, an iterative, self‐supervised CNN‐LSTM autoencoder that automatically extracts high‐fidelity translocation fingerprints from complex raw data. By evaluating reconstruction error and Lempel‐Ziv complexity via adaptive thresholding, NanoCurator effectively segregates genuine signals from diverse noise profiles in both simulated and empirical datasets. Crucially, this curation significantly elevates classification accuracy by 1.8% for a 15‐peptide panel, 2.3% for post‐translational modifications, and 1.5% for heterogeneous mixtures, while reducing data volume requirements, thereby establishing a robust, highly generalizable paradigm for automated signal quality control. Ultimately, this versatile framework accelerates precise peptide identification and enables the reliable, large‐scale application of massively parallel nanopore sensing in proteomics.

More from our Archive