DOI: 10.3390/app16199710 ISSN: 2076-3417

CineScope-Fuse: Cross-Scale Semantic Fusion for Cinematic Aesthetic Assessment of Multimodal-LLM-Generated Videos

Yanxi Dan, Yuheng Li

Multimodal large language models and diffusion transformers have made text-to-video generation accessible to film production, advertising and social-media editing, but the assessment of generated video aesthetics remains poorly aligned with how viewers diagnose cinematic failure. A video may be sharp at the frame level while failing because of jitter, implausible object motion or mismatch between the prompt and the evolving event. We introduce CineScope-Fuse, a cross-scale semantic fusion network for cinematic aesthetic assessment of multimodal-LLM-generated videos. The method decomposes video quality into spatial fidelity, temporal stability and prompt–event alignment and then recomposes them through a semantic-aware module, a unified cross-attention network and a cross-scale spatio-temporal fusion block. The design turns the three common failure families of generated videos into explicit learning targets rather than treating them as unstructured regression noise. On the proposed Cine-AIGV benchmark and five public auxiliary benchmarks, CineScope-Fuse achieved the strongest overall correlation with human ratings, reaching 0.889 SRCC and 0.884 PLCC on Cine-AIGV. Ablations further showed that semantic alignment learning, cross-dimensional attention and multi-scale temporal fusion contributed complementary gains. These results indicate that cinematic video assessment benefits from jointly modeling low-level fidelity, temporal continuity and high-level narrative alignment, especially for synthetic clips whose visual realism and semantic plausibility fail in different ways.