Integrating First-Principles Modeling with Explainable Machine Learning for Non-Isothermal Chromatography
Muhammad Bilal, Sirajul Haq, Muhammad Asif, Ali M. AlhartomiAbstract
Developing preparative chromatographic methods for multicomponent mixtures under nonisothermal conditions is difficult: mechanistic, physics-based models are accurate but computationally expensive, whereas data-driven surrogates are fast but often opaque and physically uninterpretable. We address this tension with an integrated framework that couples mechanistic simulation, ensemble machine learning, and explainable AI, privileging physical interpretability alongside predictive accuracy. We extend the nonisothermal general rate model (GRM) to three components with competitive Bi-Langmuir adsorption and solve it with a high-resolution finite-volume scheme, generating 1375 simulations that span 31 dimensionless input parameters. From the elution profiles, we extract 12 performance indicators covering retention, separation efficiency, productivity, and thermal effects, using standard chromatographic formulas. We benchmark six ensemble learning algorithms; XGBoost performs best, predicting cycle time with R2 = 0.982. We then use SHAP to interpret the surrogate and find its feature rankings to be consistent with the parametric sensitivities of the underlying GRM across every output examined─an internal consistency check that links the data-driven model back to first-principles transport behavior, operates independently of held-out accuracy, and requires no experimental data─which we corroborate with a correlation-robust permutation-importance cross-check. SHAP additionally quantifies the global and local influence of each parameter, exposing trade-offs and control mechanisms. Because the trained surrogate evaluates a new operating point in under a millisecond (about 1.6 × 107 times faster than a full GRM integration), parameter sweeps that are otherwise impractical become routine, so the framework is both fast and physically transparent. The framework therefore applies to computational process screening and mechanistic troubleshooting within the bounds of the training design, producing physically traceable, SHAP-auditable output that can inform─though not by itself constitute─a regulatory model package, which additionally requires experimental validation of the underlying model. We emphasize that the SHAP–GRM agreement establishes consistency between the surrogate and the simulation, not the physical fidelity of the GRM itself. Beyond chromatography, the approach transfers to other separation or reaction processes where simulation cost is the bottleneck.