DOI: 10.3390/electronics15163496 ISSN: 2079-9292

PEMTCL: A Prompt-Enhanced Multi-Task Contrastive Learning Framework for Fine-Grained Toxic Language Detection

Shan Jin, Xiaochao Fan, Zhenzhen He, Peng Chen

Fine-grained toxic language detection aims to assess the harmfulness, expression style, and attack target of textual content, supporting refined content moderation on online platforms. Most existing methods model each subtask independently, failing to exploit semantic correlations among tasks, while conventional contrastive learning is prone to introducing false negative samples under multi-task, multi-label settings. To address these limitations, we propose a Prompt-Enhanced Multi-Task Contrastive Learning framework (PEMTCL). The framework employs a dynamic prompt generation mechanism to construct instance-level soft prompts, enabling the encoder to adaptively capture fine-grained semantic cues required by different subtasks. It then jointly optimizes four subtasks—toxicity identification, toxic type discrimination, expression type detection, and targeted group detection—achieving inter-task knowledge complementarity within a shared representation space. Furthermore, a label-aware multi-head contrastive learning module leverages multi-task label relationships for reliable negative sample selection and incorporates a sample reweighting mechanism, effectively mitigating false negatives and class imbalance. Experiments on two Chinese datasets, ToxiCN and FG-COLD, demonstrate that PEMTCL achieves the best performance on most subtasks, with notable improvements on fine-grained tasks. Ablation studies, training data scale analysis, and case studies further validate the effectiveness and robustness of the proposed model.

More from our Archive