Multi-Metric Evaluation of Translation-Based Cross-Lingual Sentiment Consistency Using Large Language Models and Neural Machine Translation
Esra Duruoglu Cetin, Cagri SahinIn today’s globalized and digitally connected world, individuals increasingly share emotions, opinions, and experiences across multiple languages, making accurate translation essential for cross-lingual sentiment analysis. Although machine translation (MT) is widely used in multilingual applications, the relationships among translation quality, semantic similarity, and sentiment consistency remain insufficiently understood. This study investigates the performance of six LLM-based systems (GPT-4o-mini, Gemini 2.5 Flash-Lite, Qwen 2.5, Llama 3.1, Mistral 7B, and NiuTrans LMT) and four NMT-based systems (Google Translate, Microsoft Translator, NLLB-200, and LibreTranslate-v1.5) in maintaining classifier-mediated sentiment consistency across twelve translation directions involving English, Spanish, French, and Chinese. Experiments were conducted on the Multilingual Amazon Reviews Corpus (MARC), comprising 84,000 randomly sampled user reviews. A multidimensional evaluation framework was used, combining sentiment-consistency metrics (accuracy, weighted F1, MCC, and SSR), translation-quality estimation (COMET-QE), and semantic-similarity assessment (LaBSE). Statistical significance was examined using the Friedman and Nemenyi post hoc tests. The results show that GPT-4o-mini, Gemini 2.5 Flash-Lite, and Google Translate consistently ranked among the strongest systems across multiple evaluation dimensions. Performance differences were particularly pronounced in translation directions involving Chinese, highlighting the influence of language-specific structural characteristics. Furthermore, semantic similarity and translation quality exhibited only moderate relationships with sentiment consistency, indicating that high semantic similarity does not necessarily guarantee strong sentiment consistency. Overall, the findings demonstrate the importance of multidimensional and statistically grounded evaluation frameworks for assessing cross-lingual sentiment consistency and provide practical insights into the strengths and limitations of contemporary MT systems.