DOI: 10.3390/math14162941 ISSN: 2227-7390

Optimized Sparse Attention Regularized Transformer with Low-Rank Constraint for Efficient Urban Landscape Perception Modeling

Yawei Liu, Junming Chen, Hongji Yue

Deep neural networks increasingly convert street-level imagery into quantitative measures of urban perception, but the cost of transformer backbones limits repeated inference over city-scale image collections. This study proposes an optimized Sparse Attention Regularized Transformer with a Low-Rank constraint (SART-LR), a Siamese vision transformer in which softmax is replaced by learnable α-entmax attention and the attention and feed-forward projections are directly factorized. The evaluation assigns each Place Pulse 2.0 image to exactly one of the training, validation, or test partitions, thereby preventing the same image from entering multiple partitions through different comparisons. All models are tuned with an equal, architecture-specific validation budget and evaluated over five seeds. Under this image-disjoint evaluation, SART-LR reaches an average pairwise accuracy of 72.6%, 2.1 percentage points above ViT-B/16 and 1.2 points above Swin-T. The explicit layer-wise derivation gives 2.97 million trainable parameters, a 5.0-fold reduction relative to the depth- and width-matched dense backbone and a 29.1-fold difference from ViT-B/16; the latter comparison is reported only as an end-to-end model total because the architectures differ. Five repeated city-level folds give a cross-city accuracy of 67.0±1.3%, corresponding to a 5.6-point decrease from the image-disjoint result. Paired tests indicate that the advantage over Swin-T varies by attribute, and analyses stratified by rater agreement show lower accuracy for ambiguous comparisons. These findings support an accuracy–efficiency benefit on Place Pulse 2.0, while the absence of an independent urban-perception dataset and the geographic imbalance of the 56-city sample limit external-validity claims.

More from our Archive