DOI: 10.7731/kifse.b784d9bc ISSN: 2950-8991

A Preliminary Study on Machine Learning-Based Prediction of Vapor Pressure and Flash Point of Hazardous Materials

Joohyung Roh, Sukwon Shon, Jihyun Woo, Minsuk Kong

As a preliminary studdy, this study applied machine learning (ML) models to predict vapor pressure and flash point—key physicochemical properties of hazardous materials—and compared predictive performance and feature importance across different models. A dataset of 234 hazardous materials was constructed from the National Hazardous Materials Integrated Information System of the Republic of Korea's National Fire Agency using physicochemical properties such as molecular weight, boiling point, and viscosity as input features. Three regression algorithms—Random Forest, Extreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LightGBM)—were used, and their predictive performance was evaluated using the coefficient of determination (R²) and root mean square error (RMSE). Among the tested models, XGBoost demonstrated the highest predictive performance for vapor pressure (R² = 0.82, RMSE = 57.74 mmHg), whereas LightGBM achieved superior performance in flash point prediction (R² = 0.94, RMSE = 13.34 °C). By comparison, Random Forest exhibited lower predictive accuracy, particularly for vapor pressure, suggesting limitations in capturing complex nonlinear relationships. Feature importance analysis indicated that while XGBoost relied heavily on a few key features, LightGBM utilized multiple physicochemical properties more evenly. This difference in feature utilization was closely associated with differences in model performance, demonstrating that the appropriate ML model may vary depending on the target property. These findings highlight the significance of model selection and feature interpretation in hazardous material property prediction.

More from our Archive