"I Want to Publicize My Stutter": Community-led Collection and Curation of Chinese Stuttered Speech Data
Qisheng Li, Shaomei WuThis paper documents the process undertaken by StammerTalk , a grassroots community of Chinese-speaking people who stutter, to autonomously collect and curate stuttered speech data for more inclusive speech AI models. While people with disabilities are often excluded or treated merely as the subjects of AI data collection, our work introduces a new model for disability data collection in which the disability community exerts agency and control over their personal data and data-driven experiences. Our ethnographic data show that community-led data collection not only produces data needed to represent the community in AI systems, but also empowers the community and its members, by embracing - rather than concealing - stuttering and stutterer identity, and strengthening the social bonds of the community. Recognizing the lack of adequate socio-technical infrastructure for community-led, grassroots data collection, we discuss practical challenges, as well as the strategies and factors for communities to succeed in similar endeavors.