Small-Sample Recognition and Classification of Ancient Architectural Forms in Chinese Local Gazetteer Maps Using Deep Learning Method
Lan Li, Feng Kang, Xue Li, Shujia Zhang, Zhongchao ZhouChinese local gazetteer maps contain visual evidence of historical cities and architecture, but their hand-drawn, low-texture, and small-sample characteristics make automatic recognition difficult. This study develops a transfer-learning workflow to classify gazetteer-derived architectural images and examines its usefulness for digital heritage image organization, indexing, and retrieval. A dataset of 685 images covering nine classification labels was constructed from Chinese local gazetteer maps. Five representative models, ResNet-18, MobileNetV3-Small, EfficientNet-B0, DenseNet-121, and ViT-B/16, were trained and compared under the same transfer-learning setting. Their performance was evaluated using accuracy, macro-average F1 score, class-level metrics, confusion matrices, difficult-sample analysis, and Grad-CAM visualization. CNN-based models outperformed the Vision Transformer on this task. EfficientNet-B0 achieved the best results, with a test accuracy of 87.38% and a macro-average F1 score of 86.60%; after disabling class weights, the macro-average F1 score increased to 86.81%. Misclassifications mainly occurred in visually similar or context-dependent category pairs, including gate tower and drum tower, bell tower and drum tower, and buildings inside and outside the city. The results show that compact CNNs can support classification, indexing, and retrieval of hand-drawn architectural images in gazetteer archives, although the workflow should be regarded as an auxiliary tool rather than direct evidence for specific extant buildings.