Comparative Assessment of Similarity Metrics for 1 H NMR Spectral Matching: Evidence for the Robustness of the Normalized Square Dot Product
Fabien Torralba, Guillaume Hoffmann, Emmanuel Cassin, Gwladys Charpentier, Séverine Sechet, Audrey Gaucher, Guillaume Rousselot, Olivier Aroule, Christophe Morell, Asma Bourafai‐AziezABSTRACT
Automated comparison of 1 H NMR fingerprints is increasingly used for the quality control and authentication of plant extracts; however, there is insufficient guidance on the most reliable similarity metrics for large and heterogeneous spectral datasets. This paper presents a systematic comparison of similarity metrics for 1 H NMR spectra using a large experimental dataset. For this purpose, a large database of 4574 1 H NMR spectra was constructed (available on request). Spectra were examined both over the full chemical shift range and within three windows: aliphatic (0–3 ppm), aromatic (6–10 ppm) and combined (0–3 + 6–10 ppm). Classical vector‐based metrics (Pearson correlation, cosine similarity, Euclidean distance, dot products, normalized area by sum), binary indices (Jaccard, Dice, Russell‐Rao) and three recently introduced measures, including the normalized square dot product (NSDP), were applied to all spectrum pairs. Performance was assessed using retrieval statistics (Top 1/Top 10) and bootstrap resampling. Cosine similarity, Pearson correlation, NSDP and a subtraction‐based metric yielded high mean similarity values (typically ≥ 0.94), low variance and robust Top 1/Top 10 rankings across all spectral windows, whereas Euclidean distance, raw or squared dot products and binary indices were strongly affected by intensity scaling, threshold choice and peak overlap. The aliphatic region provided the strongest discrimination, whereas combining aliphatic and aromatic regions improved overall robustness of spectral matching. These results support NSDP, cosine similarity, Pearson correlation and subtraction similarity as default metrics for automated 1 H NMR spectral matching in plant extract quality control, dereplication and library‐search workflows.