DOI: 10.1126/sciadv.aed7511 ISSN: 2375-2548

Predictive processing as a scalable computational principle for embodied multitask intelligence

Hayato Idei, Tamon Miyake, Tetsuya Ogata, Yuichi Yamashita

Humans exhibit remarkable flexibility in adapting to diverse and uncertain environments—a hallmark arising from the brain’s ability to integrate multimodal sensory streams into coherent predictive models. Drawing on this principle, we introduce a scalable hierarchical multimodal recurrent neural network grounded in predictive processing under the free-energy principle, capable of directly integrating more than 30,000-dimensional visuo-proprioceptive inputs without dimensionality reduction or handcrafted preprocessing. Using sensory data from teleoperation of a full-scale physical humanoid robot performing two caregiving-related tasks—rigid-body repositioning and flexible-towel wiping—the model learns to predict high-dimensional visuo-proprioceptive streams end to end. In open-loop adaptive inference experiments, the framework exhibits three emergent properties: (i) self-organized hierarchical latent dynamics governing task transitions, uncertainty, and occlusion inference; (ii) robustness to degraded vision via multimodal integration; and (iii) asymmetric interference in multitask learning. Although evaluated in simulations, the framework is extensible to closed-loop robot control, with proprioceptive predictions driving action, thereby establishing a generalizable computational foundation bridging brain theory, artificial intelligence, and embodied robotics.

More from our Archive