DOI: 10.3390/computers15090632 ISSN: 2073-431X

Reasoning Together: Designing and Evaluating MLLM Team Strategies for Multimodal Quiz Questions

Anastasia Kotelnikova, Viktor Byzov, Maria Dolzhenkova, Evgeny Kotelnikov

Multimodal large language models (MLLMs) often struggle with open-ended questions requiring the integration of visual evidence, indirect textual clues, and background knowledge. We investigate whether team-based inference improves performance on Russian-language multimodal What? Where? When? questions. We introduce a nine-dimensional design space and a seven-stage reasoning pipeline for MLLM teams, instantiate six collaboration strategies, and evaluate six heterogeneous MLLMs on a new multimodal dataset of 1170 questions. Matched comparisons systematically assess performance differences associated with majority voting, captain-based aggregation, rationale sharing, and an additional discussion round. All team configurations outperformed the corresponding individual models. Rationale sharing produced the strongest and most consistent gains, while additional interaction was beneficial mainly when rationales were exchanged. The best strategy, Talkative Debate with Gemma 4 as captain, achieved 31.2% accuracy, while team-based inference yielded improvements of up to 20 percentage points over the corresponding individual models. An analysis of whether correct candidate answers appeared, persisted, or disappeared between the initial and post-discussion rounds showed that debate improved performance mainly by preserving or introducing correct hypotheses before final aggregation. Thus, rationale-aware collaboration substantially improves answer accuracy on the Russian-language multimodal quiz questions studied here, although visual interpretation, knowledge gaps, implicit clues, and answer-format constraints remain major limitations.