DOI: 10.1177/18758967261476359 ISSN: 1064-1246

Between accuracy and empathy: A comparative study of human and transformer-based multi-label emotion annotation

Ulfeta Marovac, Anida Vrcić Amar

Abstract

This study examines structural alignment between human and transformer-based multi-label emotion annotation in a low-resource language setting. A corpus of 408 Serbian narratives was manually annotated using soft-label scores mapped to the 28-category GoEmotions taxonomy. Inter-annotator agreement was moderate (micro F1 = 0.58), reflecting the inherent subjectivity of fine-grained emotion labeling. Four modeling paradigms were compared: a zero-shot English model applied to machine-translated texts, a domain-adapted fine-tuned model trained on the translated corpus, multilingual XLM-RoBERTa models trained under hard- and soft-label supervision, and regional South-Slavic models trained on original Serbian texts. The highest overall performance was achieved by the domain-adapted GoEmotions model fine-tuned on the translated corpus (micro F1 = 0.50), substantially outperforming zero-shot cross-lingual transfer. Among multilingual and regional models not originally trained on emotion-specific data, XLM-R-BERTić achieved the strongest micro-level performance (micro F1 = 0.36), indicating the contribution of regional language pre-training to human–machine alignment. Polarity analysis conducted on the pre-trained baseline model indicated low disagreement (≈6%), whereas divergences remained pronounced at the fine-grained category level. The intensity-weighted soft-label annotations reflect the fuzzy boundaries of affective categories, highlighting the importance of domain adaptation and emotion-specific training for improving human–machine alignment.

More from our Archive