Language as data, hate as task: The epistemic normalisation of hate speech in AI
Stefanie UllmannAutomated hate speech detection is now a large technical field, but engages unevenly with linguistic and social-scientific work on how hate and discrimination are communicated. This article uses corpus-assisted discourse analysis of 322 publications (299 computer science, 23 social sciences and humanities) to compare how language, context, bias and hate speech are conceptualised. Computer science texts construct language as data and hate speech as a classification task, developing intricate taxonomies of model-side bias while largely sidelining racism, power and discourse context. Social sciences and humanities texts frame hate speech as an evolving semiotic and socio-political practice, foregrounding racism, intersectionality and governance. These configurations form an ‘epistemic pipeline’ in which language is translated into features and scores, normalising some forms of hate while obscuring others. I argue that linguistically and socially informed detection would require reconfiguring what is being modelled rather than simply adding more data or architectures.