Multi-Agent GenAI for Harm Reduction: A Psychologist’s (or Social Scientist’s) Tutorial on Classifying Harmful Workplace Language
S. Gabe Hatch, Bryar Topham, Cristen Dalessandro, Zachary T. Goodman, Jason T. Martineau, Alexander G. LovellBackground: The American Psychological Association has recently acknowledged that generative artificial intelligence has the potential to cause considerable harm if not properly managed and supervised. The current tutorial situates psychologists, and social scientists, as uniquely positioned to take part in the development of these models given their ethical background, understanding of human behavior, and ability to conduct meaningful research. Method: In an effort to cause the least amount of harm during the training process, the current work proposes multi-agent generative artificial intelligence simulations as a way to build initial safeguards into deployable models using an accessible and real-life use case: detecting harmful language (e.g., profanity) in employee recognition messages. Results: We provide replicable R code on how to run simulations, which results in models with high levels of classification accuracy on simulated data (i.e., 97.5%). Conclusions: Finally, strengths, limitations, and future directions for research are discussed.