DOI: 10.1021/acs.jcim.6c01878 ISSN: 1549-9596

Unimodal vs Multimodal Learning: A Systematic Evaluation of Fusion Strategies and Model Design for Molecular Property Prediction and Uncertainty Quantification

Joseph Wasswa, George William Kajjumba, Bharath Ramsundar

Abstract

Accurate prediction of molecular properties is fundamental to environmental chemistry, yet remains challenging when experimental data are limited. Multimodal fusion provides a promising strategy for integrating complementary molecular representations; however, the relative contributions of molecular representation, fusion strategy, and learning algorithm to the predictive accuracy and uncertainty remain poorly understood. Five molecular modalities (RDKit descriptors, Mol2Vec embeddings, graph neural network embeddings, SMILES representations, and MS2 fragmentation spectra) were evaluated by using early and late fusion strategies with four learning algorithms (LightGBM, RF, AttentiveFP, and DMPNN). Across 14 physicochemical properties, multimodal models exhibited modest numerical improvements over the best unimodal models, although these differences were generally not statistically significant. In contrast, uncertainty quantification revealed clearer distinctions among the modeling strategies. Multimodal integration significantly improved the alignment between prediction error and estimated uncertainty. The fusion strategy had a modest influence on epistemic uncertainty, with significant early versus late differences observed only for selected modality combinations, whereas meta-learner selection had the greatest effect on uncertainty calibration. Ablation, grouped SHAP, and RDKit descriptor reduction analyses showed that RDKit descriptors remained consistently informative despite substantial descriptor reduction, while Mol2Vec, SMILES, GNN embeddings, and MS2 contributed in a property-dependent and partially redundant manner. Computational cost increased substantially with multimodal complexity, whereas predictive accuracy exhibited diminishing returns, indicating that intermediate multimodal configurations often provided the most favorable balance among computational efficiency, predictive performance, and uncertainty reliability. Overall, the results demonstrate that successful multimodal learning depends on the coordinated selection of complementary molecular representations, fusion strategy, and learning algorithm rather than simply increasing the number of integrated modalities. Multimodal integration may provide particular value by improving the reliability of uncertainty estimation.

More from our Archive