ArabicOpinion: A Five-Dimensional Multi-Granular Evaluation Framework for Arabic Opinion Summarization
Bayan Aldashnan, Abdulrahman Alothaim, Ahmed AlsanadProgress in Arabic opinion summarization has been limited by the absence of standardized evaluation frameworks, despite Arabic being the fifth most widely spoken language globally. We present ArabicOpinion, the first large-scale, human-evaluated benchmark and automated evaluation framework for Arabic opinion summarization. The framework introduces entity-aware evaluation that assesses named entities alongside opinions, addressing a key gap in existing work. It includes evaluation metrics covering five dimensions: relevance, faithfulness, sentiment preservation, abstractiveness, and genericity, with a bidirectional multi-granular architecture that measures the first two independently. We benchmark GPT-4o, Claude, and JAIS. The automated metrics show strong alignment with human judgments. Opinion-level Matryoshka embeddings perform best for relevance and faithfulness, while AraBERT-Restaurant-Sentiment achieves the highest sentiment preservation accuracy. These results provide a standardized foundation for evaluating Arabic opinion summarization and offer insights into the performance of large language models on Arabic content.