CGCA-TOD: Coordinating Global Context and Task-Relevant Evidence Through Cross-Granularity Alignment for End-to-End Task-Oriented Dialogue
Jiawei Huang, Inwhee JoeEnd-to-end task-oriented dialogue systems must preserve information from the dialogue history while identifying evidence relevant to the current task. Existing approaches typically encode the complete multi-turn history into a unified representation. Although this preserves information accumulated across dialogue turns, it can combine contextual components with different functions and contextual scopes. To address this issue, we propose CGCA-TOD, a cross-granularity context alignment framework that coordinates global contextual information with task-focused context modeling. CGCA-TOD constructs global, interaction-focused, and current task-focused representations with increasingly focused contextual scopes and progressively refines the global representation through two alignment stages guided by the focused representations. This process adjusts the contribution of global contextual information according to the current interaction and task. We further introduce an auxiliary contrastive regularization objective to constrain the relationships among representations produced at different alignment stages. Experiments on MultiWOZ show that CGCA-TOD achieves competitive performance under both full-data and low-resource settings.