DOI: 10.3390/buildings16153129 ISSN: 2075-5309

Semantic Segmentation and Spatial Feature Quantification of Interior Environmental Design Elements: A Deep Learning-Based Framework for Data-Driven Indoor Space Analysis

Yunda Shi, Hui Yu, Xin Dong, Chenyu Tang

Image-based analysis offers a scalable way to examine indoor environmental design, but semantic segmentation studies often stop at pixel-level recognition and provide limited design-oriented quantification. This study proposes the Interior Design Element Segmentation and Spatial Quantification Framework (IDESQ Framework) to convert indoor scene images into measurable spatial design indicators. Using ADEChallengeData2016, an indoor subset containing eight scene categories was constructed, and the original ADE semantic labels were re-mapped into twelve interior environmental design element categories. U-Net, DeepLabv3+, PSPNet, and SegFormer-B0 were evaluated under the same annotation system. DeepLabv3+ achieved the highest performance, with a mean Intersection over Union (mIoU) of 0.494, Pixel Accuracy of 0.755, and Mean Accuracy of 0.661 on the internal test split; on the independent ADE validation subset, its mIoU was 0.495. The predicted masks were then used to calculate area proportions, furniture density, functional facility ratio, soft decoration ratio, decorative object ratio, greenery ratio, visual complexity, and spatial distribution features. The quantified results showed scene-dependent patterns, including a high spatial envelope ratio and low visual complexity in corridors, higher functional facility ratios in kitchens and bathrooms, and richer decorative composition in living rooms. Ground-truth–prediction (GT–Pred) consistency analysis showed that scene-level aggregation improved agreement between prediction-derived and GT-derived indicators, with a Pearson correlation of 0.960 for the eight main indicators. These results indicate that IDESQ can support automated and interpretable comparison of indoor design element composition and spatial patterns across scene types.

More from our Archive