DOI: 10.3390/foods15193375 ISSN: 2304-8158

Prediction of Chemical Composition of Alfalfa Blocks Using FTIR Spectroscopy and Machine Learning Models

Haijun Du, Yanhua Ma, Liying Cao, Hongzhou Liu, He Su, Xiangbo Liu

Alfalfa blocks are an important commercial roughage, and their crude protein (CP), starch, fat, neutral detergent fibre (NDF), and acid detergent fibre (ADF) contents are key indicators of nutritional quality. This study investigated Fourier-transform infrared (FTIR) spectroscopy combined with machine learning for rapid quantitative prediction of these constituents. Different spectral preprocessing strategies and variable-selection methods were compared using partial least squares regression (PLSR), random forest regression (RFR), and back-propagation (BP) neural networks. The effects of spectral preprocessing and preceding baseline correction varied among constituents, indicating that baseline correction did not uniformly improve predictive performance. Competitive adaptive reweighted sampling (CARS) reduced spectral dimensionality, and the resulting PLSR models outperformed the RFR and BP models for all five constituents in the reported comparisons, with values reported as mean ± standard deviation across 20 repeated data partitions. For CP, the PLSR model combining first-derivative Savitzky–Golay preprocessing with CARS (SG1D-CARS-PLSR) achieved a prediction-set coefficient of determination (R2) of 0.961 ± 0.021 and a residual predictive deviation (RPD) of 5.629 ± 1.245. For starch, SG1D-CARS-PLSR achieved an R2 of 0.877 ± 0.023 and an RPD of 2.959 ± 0.291. For NDF, the PLSR model combining adaptive iteratively reweighted penalised least-squares baseline correction with CARS (BC-airPLS-CARS-PLSR) achieved an R2 of 0.986 ± 0.006 and an RPD of 9.355 ± 1.947. For ADF, BC-airPLS-CARS-PLSR achieved an R2 of 0.839 ± 0.031 and an RPD of 2.591 ± 0.287. For fat, BC-MSC-CARS-PLSR achieved an R2 of 0.697 ± 0.058 and an RPD of 2.1.883 ± 0.164, and did not deliver the expected stable prediction performance as predefined. These findings highlight the importance of constituent-specific preprocessing and variable selection and demonstrate that the optimised FTIR-based models provide a rapid for estimating the five constituents from a single spectrum, with an analysis time considerably shorter than that of the reference wet-chemistry methods.