UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks

arXiv cs.CV / 4/28/2026

📰 NewsIdeas & Deep AnalysisModels & Research

共有:

Key Points

The paper proposes UpstreamQA, a modular approach to make Video Question Answering (VideoQA) use explicit multi-step reasoning rather than opaque implicit reasoning in many large multimodal models.
UpstreamQA first applies multimodal large reasoning models to generate object identification and scene context, then feeds the resulting enriched reasoning traces into downstream LMMs for final VideoQA.
Experiments on the OpenEQA and NExTQA datasets using LRMs (o4-mini, Gemini 2.5 Pro) and LMMs (GPT-4o, Gemini 2.5 Flash) show that explicit reasoning can improve both performance and interpretability.
The authors also find that adding explicit reasoning may reduce performance in cases where the baseline model already performs strongly, indicating the approach is scenario-dependent.
Overall, UpstreamQA provides a framework for combining explicit reasoning with native multimodal understanding in VideoQA to improve results and diagnostic transparency.

Abstract

Video Question Answering (VideoQA) demands models that jointly reason over spatial, temporal, and linguistic cues. However, the task's inherent complexity often requires multi-step reasoning that current large multimodal models (LMMs) perform implicitly, leaving their internal decision process opaque. In contrast, large reasoning models (LRMs) explicitly generate intermediate logical steps that enhance interpretability and can improve multi-hop reasoning accuracy. Yet, these models are not designed for native video understanding, as they typically rely on static frame sampling. We propose UpstreamQA, a modular framework that disentangles and evaluates core video reasoning components through explicit upstream reasoning modules. Specifically, we employ multimodal LRMs to perform object identification and scene context generation before passing enriched reasoning traces to downstream LMMs for VideoQA. We evaluate UpstreamQA on the OpenEQA and NExTQA datasets using two LRMs (o4-mini, Gemini 2.5 Pro) and two LMMs (GPT-4o, Gemini 2.5 Flash). Our results demonstrate that introducing explicit reasoning can significantly boost performance and interpretability of downstream VideoQA, but can also lead to performance degradation when baseline performance is sufficiently high. Overall, UpstreamQA offers a principled framework for combining explicit reasoning and multimodal understanding, advancing both performance and diagnostic transparency in VideoQA in several scenarios.