DOI: 10.3390/electronics15194420 ISSN: 2079-9292

A Multi-Agent Framework for Arabic Scam Analysis: Dataset-Specific mT5 Classification and Cultural Annotations

Saleh Alqahtani, Priyadarsi Nanda, Qiang Wu, Raddad Faqihi, Bashair Alrashed

This study presents a multi-agent framework for Arabic scam analysis, with roles for classification, behavioural interpretation and educational guidance. A total of 3000 messages were translated and enriched from English-language collections associated with the Short Message Service (SMS) Spam Collection, Enron and SmishTank. The experiments examine whether dataset-specific multilingual Text-to-Text Transfer Transformer (mT5) training improves classification and whether cultural annotations and an additional cultural training objective improve performance. The training, validation and test partitions contained 2400, 300 and 300 messages. Across ten random seeds, dataset-specific classifiers achieved 88.17% mean test accuracy and 88.07% macro-averaged F1 score (macro-F1); the 95% confidence interval for mean accuracy was 87.36–88.97%. Accuracy exceeded combined-data mT5 by 1.90 percentage points (Holm-adjusted p=0.0250). Mean accuracies were 98.50%, 92.90% and 73.10% for SMS, email and SmishTank-derived messages. In exploratory cultural comparisons, supplying annotations increased encoder–decoder accuracy by 1.80 points, while adding the cultural objective increased it by 1.23 points; neither difference was statistically significant after adjustment. The 150-message agent assessment yielded 84.7% mean integrated threat accuracy. The results support dataset-specific classification within the evaluated collections and show limited benefits from the tested cultural formulation. Independent Arabic data and user studies are priorities for extending the framework.