More than Black Boxes: Machine Learning Models’ Capacity to Capture Housing Submarket Patterns
José Rojas-Quiroz, Carlos Marmolejo-DuarteMachine learning (ML) models like XGBoost have gained traction in housing valuation as they outperform classical hedonic models in predictive accuracy. However, their black box nature has limited their interpretability, until recent advances like SHAP values enabled deeper insights into variable contributions. This study explores whether XGBoost not only predicts housing prices more accurately but can also help identify latent submarket structures, without imposing predefined spatial or socioeconomic boundaries. Using a raw dataset of 7092 housing listings from Barcelona (reduced to 6527 unique listings and finally to 6112 observations after outlier removal), we employ a sequential methodological framework: first, OLS and Spatial Durbin Models identify statistically significant predictors while accounting for spatial dependence; second, an XGBoost model is trained with these validated variables to extract SHAP values quantifying each attribute’s model-attributed contribution to individual predicted prices; third, clustering these SHAP values reveals three distinct submarkets with systematic differences in model-attributed valuation patterns. We validate the resulting segments through five complementary procedures: out-of-sample cluster assignment verifying segment generalizability; nested OLS interaction models that formally test differences in model-attributed valuation patterns across clusters; spatial block bootstrap resampling that assesses cluster stability under geographic subsampling; non-parametric tests against the 2017 cadastral value—an independent administrative proxy for spatial value stratification; and spatial autocorrelation statistics confirming non-random geographic structure. Together, this multi-pronged validation provides convergent evidence that the identified segments capture stable and externally coherent patterns in the model’s learned valuation structure. Unlike prior studies focused solely on prediction, our approach bridges ML’s technical rigor with econometric interpretability. We argue that this capability may stem from XGBoost’s tree-based architecture, which naturally partitions the feature space to accommodate heterogeneous valuation patterns. Future research could extend this framework to more recent tree-based models, further advancing interpretable ML for housing market analysis.