DOI: 10.1145/3847120 ISSN: 1551-6857

Advancing Egocentric Video Dialogue: A Contextual Reasoning Approach with New Benchmark Dataset

Jiayi Zou, Chaofan Chen

Egocentric video reasoning capability is essential for advancing the development of first-person wearable devices. However, most existing egocentric video understanding tasks are limited to single-instance text input, overlooking the model's capability to reason about context throughout the entire duration of the video. To bridge this gap, we propose a novel task named Egocentric Video Dialogue, which refers to answering questions by considering both the video content and dialogue history. To support this task, we design a new dataset, EgoDialogue, which constructs continuous dialogues based on video question answering datasets. Egocentric videos and dialogue histories present two challenges, requiring the model to understand core interactions in visual data and the dialogue history in text, respectively. To address these issues, we introduce an EgoReason framework, consisting of an interaction exploration module and a dialogue reasoning module. The former module first reorganizes patches that correspond to the same spatial position across different frames and then sorts these blocks according to their frame order. In this way, it models spatiotemporal dynamics to help explore interactions. The latter module improves the understanding of dialogues by gradually integrating the questions and answers from each dialogue round into the fusion layer. In the experiments, our method shows superior performance compared to several VideoQA models and vision language models, achieving state-of-the-art results in prediction accuracy and various machine translation metrics.