HoloDiff: Holistic Talking Human Animation Via Latent Diffusion
Xinmu Wang, Fumi Honda, Xiang Gao, Junqi Huang, Wei Chen, Xianfeng David GuABSTRACT
Holistic talking human animation aims to synthesize coordinated facial expressions and body poses that are consistent with spoken content and human communication patterns. This task requires both realistic motion generation and accurate temporal alignment between speech and motion, which remain challenging for existing methods that model facial and bodily motion separately or operate in high‐dimensional motion spaces. We present HoloDiff, a latent diffusion framework for holistic talking human animation. Given an input speech signal, HoloDiff encodes facial and bodily dynamics into a compact and structured latent representation that preserves anatomical structure and cross‐part interactions. A speech‐conditioned diffusion process is then applied in this latent space to model long‐range temporal dynamics and generate motion sequences aligned with audio input. By performing diffusion in a structured latent space, HoloDiff reduces modeling complexity while improving temporal coherence and audio–motion alignment. Extensive experiments on the SHOW dataset demonstrate that HoloDiff produces more realistic and expressive talking human animation and achieves better multimodal correspondence between speech and motion compared to baselines.