The label effect in speech perception: Comparing AI-generated and human voices
Kerstin Heidi Kipp, Janis Hanenberg, Kyara Klinar, Emilian Tersek, Annemarie Klemm, Hannah WehrumAs synthetic speech increasingly approaches human-like naturalness, understanding how labelling as ‘human’ or ‘AI-generated’ shapes listeners’ perceptions is crucial for communication research, voice studies and AI transparency policy. This study examines how explicit and sometimes misleading labels influence evaluations of highly natural-sounding AI-generated voice clones vs. human speakers. In an online experiment ( N = 200), participants rated short novel excerpts spoken by humans or generated with ElevenLabs Multilingual v2, each paired with either accurate or intentionally incorrect labels. Two consistent effects emerged: human voices were rated more positively than AI-generated voices and recordings labelled ‘human’ received higher ratings than those labelled as ‘AI-generated’, regardless of the actual voice type. Notably, AI voices labelled as ‘human’ were rated more positively than human voices labelled as ‘AI-generated’, indicating that labelling can partly outweigh auditory impressions. Age moderated this effect – older participants rated AI-labelled voices more positively and showed smaller differences between label conditions. Women rated voices more positively than men, while professional background in communication and prior experience with AI voices had no effect. Although limited by the restricted voice set and non-interactive design, the findings illustrate how symbolic framing shapes auditory perception and underscore the importance of transparent, communication-sensitive labelling practices.