DOI: 10.1515/dsll-2026-0033 ISSN: 2943-0607

Human-Annotated or MLLM-Assisted? A Comparative Evaluation of Filmic Metaphor Identification Using FILMIP-AI

Lorena Bort-Mir, Mohammad Saeid Miri

Abstract

This study introduces FILMIP-AI, an operationalization of the Filmic Metaphor Identification Procedure (FILMIP) for multimodal large language models (MLLMs) by converting its structured interpretive steps into chain-of-thought (CoT) prompts. We evaluate FILMIP-AI using a sample of 30 TV commercials across three native video-processing MLLMs (Gemini 2.5 Pro, Gemini 3.1 Pro Preview, and Gemini 3.5 Flash) under several k-shot prompting strategies from zero-to ten-shot. The results demonstrate that CoT prompting with annotated examples significantly improves the models’ performance, with mid-range strategies (4–8 shots) yielding the best performance in this pilot sample. Within this sample, Gemini 3.1 Pro Preview achieved the highest accuracy, yielding an average F1 score of 0.83 in the 4-shot setup. Despite these strengths, qualitative error analysis reveals persistent vulnerabilities: models frequently misclassify non-metaphorical content as metaphorical, struggle to integrate cross-modal information, fail with culturally embedded metaphors, and hallucinate invalid conceptual hierarchies. These findings align with prior research on MLLM limitations in multimodal metaphor identification. We conclude that while FILMIP-AI can significantly reduce manual annotation labor, it is best used as a pre-annotation tool under rigorous human expert supervision.