Cross-Lingual Transfer for Mammography Report Classification in Low-Resource Settings
Anuar Dosmaganbetov, Tomiris Zhaksylyk, Beibit AbdikenovLabeled clinical text is scarce in many non-English healthcare settings, limiting the development of robust clinical NLP systems. We tested whether supervision from Spanish mammography reports improved the classification of Russian-language reports from Kazakhstan. The study included 4279 Spanish and 495 Russian reports mapped to three BI-RADS-derived operational classes (routine, follow-up, and suspicious). Explicit BI-RADS identifiers were removed from the input, and Russian performance was assessed using grouped five-fold out-of-fold evaluation. We compared Russian-only XLM-R fine-tuning, Spanish zero-shot transfer, sequential Spanish-to-Russian XLM-R fine-tuning, and a matched word/character TF–IDF logistic-regression baseline. Under the fixed split, label budget, and four-epoch schedule tested here, sequential transfer exceeded Russian-only XLM-R in all three paired training seeds, increasing the mean macro-F1 from 0.2178 to 0.3239; zero-shot performance was less stable (mean 0.1898). However, the matched TF–IDF model achieved the highest macro-F1 (0.5214) and better performance on the rare suspicious class (F1 0.3306 versus 0.1139 for transferred XLM-R). Thus, Spanish initialization was compatible with subsequent low-resource Russian adaptation in this experiment, while the target-language lexical model remained stronger. These findings are limited to the two datasets, grouped split, XLM-R configuration, and small, imbalanced target-data regime evaluated here.