DOI: 10.3390/sym18101594 ISSN: 2073-8994

Research on Manchu Multi-Font Recognition Based on Structure-Aware Adversarial Reconstruction with Difficult Instance Masking

Hang Yu, Dadong Wang, Yu Zhou, Mingchen Sun, Tongtong Zhang, Zhenjiang Tan

Optical Character Recognition (OCR) for Manchu is hindered by large intra-domain variation caused by diverse fonts. Consequently, traditional end-to-end models perform poorly in zero-shot cross-domain scenarios. To address this bottleneck, we propose a two-stage decoupled paradigm, “Normalize-then-Recognize,” and construct the Target-Masked Adversarial Reconstruction Network (TMAR-Net) as a modular visual pre-filter. To address the continuous cursive axis and severe foreground–background imbalance of Manchu script, TMAR-Net utilizes a paired ConvNeXt generator equipped with bottleneck self-attention and introduces a Structure-Aware Hard-Mining Target-Masked Loss (SA-HMTM Loss). Through the joint application of pixel-level masking, instance-level hard mining, and feature-level structural anchoring, this approach reduces broken strokes, long-tail sample collapse, and topological hallucinations during reconstruction. Evaluated on a dataset containing 20,000 words, system-level validation shows that, after lightweight domain adaptation, the end-to-end recognition accuracies of single-font and multi-font baseline models rise to averages of 92.34% and 92.56%, respectively, peaking at 95.52%. This yields an average absolute improvement of over 90 percentage points compared with the direct single-font baseline, closely approaching the 99.81% recognition rate typically achieved under in-domain conditions. This work provides a robust, system-level solution for digitizing high-variance, low-resource minority multi-font documents.