DOI: 10.3390/molecules31162836 ISSN: 1420-3049

Linear and Nonlinear Learning from Spectroelectrochemical Data: Interrogation of PLS and CNN Behavior Under Experimental Scarcity

Abderrahman Atifi

Spectroelectrochemistry (SEC) provides unique information-rich datasets by coupling molecular spectroscopic fingerprints with electrochemical activity. Despite its richness, direct machine-learning (ML) analysis of SEC data under realistic experimental constraints and scarcity remains unexplored. This work examines how much spectroscopically encoded electrochemical information can be learned from minimal SEC training data in a chemically reversible two-electron redox system, and how model choice interacts with limited experimental diversity across scan rate. Using purely experimental SEC datasets collected at four scan rates (2, 3, 5, and 7 mV/s), partial least squares (PLS) and convolutional neural networks (CNNs) regressors are evaluated under several SEC-level training, validation, and test configurations. Model performance is assessed across three targets of increasing physical complexity, including species concentrations, derivative cyclic voltabsorptometry (DCVA) current, and experimental cyclic voltammetry (CV) current. Under single-SEC training, both models achieve the expected near-quantitative concentration prediction (R2 ~0.99), while performance decreases for DCVA (R2 ~0.93) and most substantially for CV current (R2 ~0.78), reflecting the progressively weaker and more indirect encoding of these targets within the absorbance data. Introducing minimal experimental diversity with only two distinct training SEC datasets enables both PLS and CNN models to generalize strongly to an unseen third SEC dataset, achieving maximum CV R2 values approaching ~0.98 in the most favorable configurations. CNN models extend the apparent linear performance ceiling observed for PLS by capturing localized, scan-rate-conditioned nonlinear correlations between spectral evolution and the experimentally measured CV response, yielding improved waveform reconstruction and greater robustness to training SEC dataset selection. These results demonstrate that, within the present chemically reversible and spectroscopically well-resolved SEC system, high-fidelity prediction of electrochemical targets can be achieved without large datasets when limited but strategically selected electrochemical diversity is introduced. SEC dataset linearity is further shown to be target-dependent and becomes operationally meaningful only when scan-rate space is sufficiently sampled. More broadly, this work establishes a controlled framework for investigating ML-enabled SEC dataset analysis under experimentally scarce conditions and provides guidance for experimental design and calibration in low-data spectroelectrochemical settings.

More from our Archive