DOI: 10.1002/jrs.70206 ISSN: 0377-0486

High‐Fidelity Synthetic Raman Spectra Generation for Sinter Basicity Prediction Using β‐Variational Autoencoders

Marjorie Ariele Pereira, Daniel Cruz Cavalieri, Adilson Ribeiro Prado, Cassius Zanetti Resende, Erica Simões Rodrigues

ABSTRACT

Raman spectroscopy has emerged as a powerful analytical tool across diverse industrial sectors, owing to its nondestructive nature, high chemical specificity, and ability to provide unique molecular “fingerprints.” In the steelmaking industry, this technique offers a promising route for the rapid and precise characterization of critical materials such as sinter, a porous agglomerate of iron ore fines essential for blast furnace charging. However, practical deployment is often hindered by the scarcity of labeled spectral data. This work investigates the use of a ‐variational autoencoder (‐VAE) for the generation of high‐fidelity synthetic Raman spectra from a limited set of real sinter samples (321 spectra). With and latent codes sampled from a full‐covariance Gaussian fitted to the aggregate posterior, the model produces genuinely novel spectra whose global distributional fidelity matches that of SMOTE (Fréchet spectral distance: 4.55 vs. 4.41), despite SMOTE being constructed by interpolation of existing spectra; the mean intensity correlation with real spectra is 0.930 0.038. These synthetic spectra were used to augment supervised regression models for basicity () prediction. The LightGBM model achieved (95% CI: [0.72, 0.84]) in cross‐validation with original data; no augmentation method produced a statistically significant improvement (Wilcoxon signed‐rank, all ), and on the held‐out test set (), no augmented model exceeded the original‐only baseline (; S‐VAE: 0.809; Interpolação: 0.786; ‐VAE: 0.780; SMOTE: 0.779). The operational value of generative augmentation emerges instead in a downstream classification experiment: classifiers trained with increasing fractions of synthetic spectra sustain higher macro F1 at 100% synthetic fraction (‐VAE  0.66, S‐VAE  0.60) than SMOTE‐augmented classifiers, which degrade to 0.58. These results establish that the value of VAE‐based generation lies in producing distributionally faithful yet novel spectra that remain useful when consumed directly by downstream tasks such as spectral library expansion and classifier training, rather than in regression gains.

More from our Archive