Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: zhang, Kaixin, Li, Xiaohe, Li, Jiahao, Wu, Haohua, Zhao, Xinyu, Fan, Zide, Wang, Lei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910054190415872
author zhang, Kaixin
Li, Xiaohe
Li, Jiahao
Wu, Haohua
Zhao, Xinyu
Fan, Zide
Wang, Lei
author_facet zhang, Kaixin
Li, Xiaohe
Li, Jiahao
Wu, Haohua
Zhao, Xinyu
Fan, Zide
Wang, Lei
contents Multi-modal Large Language Models (MLLMs) have significantly advanced video reasoning, yet Video Question Answering (VideoQA) remains challenging due to its demand for temporal causal reasoning and evidence-grounded answer generation. Prevailing end-to-end MLLM frameworks lack explicit structured reasoning between visual perception and answer derivation, causing severe hallucinations and poor interpretability. Existing methods also fail to address three core gaps: faithful visual clue extraction, utility-aware clue filtering, and end-to-end clue-answer alignment. Inspired by hierarchical human visual cognition, we propose ClueNet, a clue-aware video reasoning framework with a two-stage supervised fine-tuning paradigm without extensive base model modifications. Decoupled supervision aligns clue extraction and chain-based reasoning, while inference supervision with an adaptive clue filter refines high-order reasoning, alongside lightweight modules for efficient inference. Experiments on NExT-QA, STAR, and MVBench show that ClueNet outperforms state-of-the-art methods by $\ge$ 1.1%, with superior generalization, hallucination mitigation, inference efficiency, and cross-backbone compatibility. This work bridges the perception-to-generation gap in MLLM video understanding, providing an interpretable, faithful reasoning paradigm for high-stakes VideoQA applications.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15008
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
zhang, Kaixin
Li, Xiaohe
Li, Jiahao
Wu, Haohua
Zhao, Xinyu
Fan, Zide
Wang, Lei
Computer Vision and Pattern Recognition
Multi-modal Large Language Models (MLLMs) have significantly advanced video reasoning, yet Video Question Answering (VideoQA) remains challenging due to its demand for temporal causal reasoning and evidence-grounded answer generation. Prevailing end-to-end MLLM frameworks lack explicit structured reasoning between visual perception and answer derivation, causing severe hallucinations and poor interpretability. Existing methods also fail to address three core gaps: faithful visual clue extraction, utility-aware clue filtering, and end-to-end clue-answer alignment. Inspired by hierarchical human visual cognition, we propose ClueNet, a clue-aware video reasoning framework with a two-stage supervised fine-tuning paradigm without extensive base model modifications. Decoupled supervision aligns clue extraction and chain-based reasoning, while inference supervision with an adaptive clue filter refines high-order reasoning, alongside lightweight modules for efficient inference. Experiments on NExT-QA, STAR, and MVBench show that ClueNet outperforms state-of-the-art methods by $\ge$ 1.1%, with superior generalization, hallucination mitigation, inference efficiency, and cross-backbone compatibility. This work bridges the perception-to-generation gap in MLLM video understanding, providing an interpretable, faithful reasoning paradigm for high-stakes VideoQA applications.
title Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.15008