DOI: 10.1021/acs.jpca.6c03296 ISSN: 1089-5639

Bridging the Computational–Experimental Domain Gap: An Autoencoder-Based Machine Learning Workflow for Raman Spectra Characterization of Amino Acid Mixtures

Sheng-Hsuan Hung, Yu-Huan Huang, Zong-Rong Ye, Berlin Chen, Yi-Hsin Liu, Ming-Kang Tsai

Abstract

Raman spectroscopy offers rapid, nondestructive molecular characterization, but machine-learning analysis of experimental Raman spectra is often limited by the cost and difficulty of constructing sufficiently large, reproducible experimental data sets. Theoretical spectra generated by density functional theory provide an attractive alternative; however, discrepancies in peak positions, intensities, line shapes, and background signals create a substantial domain gap between computed and measured spectra. This study develops an autoencoder-based workflow that enables machine-learning classifiers trained exclusively on computational Raman spectra to identify the dominant amino acid component in experimental mixtures. Raman spectra of the 20 common proteinogenic amino acids were calculated using density functional theory over the 400–1800 cm1 region. Binary mixtures were generated by randomly combining two amino acids at variable ratios, producing 20,000 computational spectra for training Random Forest and XGBoost classifiers. Both classifiers achieved approximately 99.5% accuracy on held-out computational spectra, confirming that the dominant component could be identified from theoretical spectral features. Nevertheless, direct application to experimental phenylalanine–glutamate mixtures resulted in essentially 0% accuracy, demonstrating that baseline correction alone was insufficient to overcome the computational–experimental discrepancy. To address this limitation, experimental spectra were first processed using asymmetrically reweighted penalized least-squares baseline correction and subsequently transformed toward their corresponding computational representations using variational autoencoder, standard autoencoder, and Gaussian-mixture variational autoencoder architectures. Combining baseline correction with the variational autoencoder constrained predictions to the two amino acids present in the mixtures but yielded only 50% accuracy. The Gaussian-mixture variational autoencoder produced no meaningful further improvement. In contrast, removal of latent-space Gaussian regularization allowed the standard autoencoder to minimize the direct reconstruction discrepancy between experimental inputs and computational targets, increasing classification accuracy to 99.645% for Random Forest and 99.395% for XGBoost. These results demonstrate that baseline correction combined with deterministic autoencoder-based domain transformation can effectively bridge computational and experimental Raman spectra. The proposed workflow substantially reduces dependence on large experimental training data sets and provides a practical framework for low-cost Raman-based characterization of amino acid mixtures, with potential applications in biomedical analysis, clinical screening, and chemically informed spectroscopic machine learning.

More from our Archive