Quantifying Social Biases in Language Model Classifiers is Domain-Dependent
Tamara Quiroga, Felipe Bravo-Marquez, Valentin BarriereAs natural language processing (NLP) systems increasingly operate on heterogeneous data, faithful social bias evaluation requires benchmarks that are both domain-sensitive and scalable. However, existing resources still face a core trade-off: template-based benchmarks scale easily but often lack linguistic authenticity, whereas naturally occurring examples (NOEs) reflect real usage but are expensive to curate and often provide limited coverage across domains and social groups. To bridge this gap, we investigate whether large language models (LLMs) can automatically adapt template-based bias datasets to specific domains using zero-shot prompting.
We evaluate the effectiveness of LLMs along three dimensions: bias consistency, which measures agreement between bias estimates derived from adapted templates and NOEs; content preservation, which assesses whether the original semantic intent is maintained; and domain adherence, which evaluates how well the generated text reflects the linguistic characteristics of the target domain. Across the Equity Evaluation Corpus (EEC) and the Identity Phrase Templates Test Set (IPTTS), and across three domains—IMDb, Twitter, and Wikipedia Talk Pages—using both automatic and human evaluations, we show that domain-adapted templates capture real-world bias patterns more faithfully than standard templates. Overall, our results highlight the importance of domain variation in fairness research and position LLM-based adaptation as a scalable and reproducible framework for robust social bias assessment in NLP.