DOI: 10.3390/electronics15163612 ISSN: 2079-9292

Dialect Bias in Arabic Toxicity Detection Under Controlled Distributional Shifts

Maha Alamri

Toxicity detection models can behave inconsistently across Arabic dialects, yet the level of training-data imbalance at which dialect-conditioned bias becomes statistically detectable remains unquantified. This paper presents a controlled framework for measuring such bias under class-conditional distributional shift. Four transformer encoders (AraBERT, MARBERT, CAMeLBERT-DA, and XLM-RoBERTa) are fine-tuned on a balanced toxicity corpus comprising texts in the Saudi, Egyptian, and Tunisian dialects across five random seeds, after which the representation of Saudi-dialect texts in the toxic class is experimentally increased at three injection levels, λ∈{0.10,0.20,0.30}, under two designs: one that jointly alters toxic-class composition and class balance, and one that isolates toxic-class dialectal composition. Bias emergence is quantified through a statistical bias-sensitivity threshold BSTB, the lowest injection level at which the Saudi-centered false-positive-rate gap is positive and statistically significant under a permutation test. False-negative-rate disparities are additionally evaluated, while ERASER-based comprehensiveness and sufficiency measures and SHAP scenario analysis serve as exploratory attribution diagnostics. Under the first design, BSTB is reached at λ=0.20 or λ=0.30 in most model–seed runs. Under the second design, no model–seed run reaches BSTB within the tested range. The matched-control analysis detects no significant marker-specific comprehensiveness effect.

More from our Archive