Long‐Range Spatial Modeling for Handwritten Urdu Character Recognition Using a
CNN
–Feature Shift Attention (
FSA
) Network
Shahzad Ashraf, Ephrance‐Eunice Namugenyi, Muhammad Ahsan Javaid, Furqan Memon ABSTRACT
Handwritten Urdu script recognition remains one of the most demanding challenges in Optical Character Recognition (OCR) due to the fully cursive nature of the script, context‐sensitive allographic character forms, overlapping ligature structures, and semantically discriminative diacritical marks. Existing OCR systems, predominantly designed for Latin‐based non‐cursive scripts, fail to generalize effectively to Urdu handwriting under diverse writers, document conditions, and image quality variations. The proposed study is specifically designed to model diacritical variations and long‐range spatial dependencies within isolated Urdu characters through a Feature Shift Attention (FSA) architecture that integrates a three‐block convolutional frontend for fine‐grained local stroke and diacritical dot extraction with a four‐stage FSA backend employing shifted window multi‐head self‐attention across the full character image. A validated and class‐balanced training corpus of approximately 113,800 samples is aggregated from the UCOM, CENPARMI, and UNHD datasets and expanded to 400,000 samples through class‐conditional geometric and photometric augmentation. A two‐phase staged fine‐tuning protocol using the AdamW optimizer with cosine‐annealed scheduling has been employed to adapt pre‐trained FSA representations to the Urdu script domain without catastrophic forgetting. When evaluated against CNN‐only, SVM, and KNN baselines on a stratified blind held‐out test set spanning all 39 Urdu base character classes, the proposed system achieved a macro‐averaged accuracy of 96.3%, precision of 95.8%, recall of 96.1%, and F1 score of 95.9%, outperforming all competing systems and enabling robust discrimination between visually similar characters differentiated by dot patterns.