DOI: 10.1515/cdbme-2026-0130 ISSN: 2364-5504

Targeted Cepstral Jitter Modeling Using Lightweight Neural Networks for Real-time Auditory Feedback Modulation

H. Eikens, A. Culcu, K. Hollborn, I. S. Schiller, G. Schmidt

Abstract

Real-time auditory feedback modulation requires signal representations that allow perceptually effective voicequality control under strict latency constraints. This work investigates a cepstrum-based approach to modulating perceived vocal roughness, a central aspect of hoarseness, in sustained vowels. The proposed method introduces frame-wise cepstral peak shifts to impose jitter-like temporal instability and uses a lightweight gated recurrent unit (GRU) to generate short-term modulation trajectories. The model is motivated by an analysis of imitated hoarse vowel recordings showing a characteristic jitter distribution and a pronounced negative autocorrelation at lag 1. Analytical evaluation shows that the proposed system reproduces these global short-term jitter properties well and induces peak-based cepstral changes consistent with imitated hoarse speech. A listening study further shows that the modulation increases perceived hoarseness and roughness, but also reduces naturalness and increases technical artifacts relative to unmodified speech. The results suggest that the neural model captures the intended local jitter dynamics, while the main limitations arise from the cepstral real-time representation and reconstruction constraints.