Investigating the cochleagram preprocessing and the role of interfered spin-wave-based reservoir computing in speech recognition
Sota Hikasa, Wataru Namiki, Daiki Nishioka, Ryo Iguchi, Jiaxuan Chen, Maki Nishimura, Kazuya Terabe, Takashi TsuchiyaRecently, artificial intelligence (AI) technologies have been applied to a wide range of real-world applications. Speech recognition is one of the most important AI tasks and is regarded as a key application for edge-AI systems. Accordingly, speech recognition has been widely used as a benchmark task for evaluating the performance of physical reservoir computing (PRC). However, most PRC studies rely on frequency-feature-extraction methods, such as cochleagram, as input preprocessing. Although these preprocessing methods provide highly discriminative features, they make it difficult to isolate the contribution of the PRC itself. Consequently, the capability required for PRC-only speech recognition remains unclear. In this study, we investigated the processing capability required for PRC-only speech recognition by systematically comparing the presence and absence of cochleagram preprocessing and PRC. As the physical reservoir, we employed a nonlinear interfered spin-wave PRC exhibiting input-frequency-dependent responses. Two speech-recognition tasks, spoken-digit recognition and speaker classification, were evaluated under all combinations of cochleagram preprocessing and PRC processing. When cochleagram preprocessing was employed, both tasks achieved high-recognition accuracy. In contrast, using only the spin-wave PRC yielded approximately 84.6% accuracy for speaker classification, whereas spoken-digit recognition remained at 29.0%. Furthermore, the accuracy of spoken-digit recognition improved when the speaker was fixed, whereas speaker classification maintained relatively high accuracy even when the spoken digit was fixed. These results indicate that speaker classification and spoken-digit recognition require different types of processing in the present spin-wave PRC and provide useful insight into the processing capability required for PRC-only speech recognition.