Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pu, Bowei, Liu, Chuanbin, Ge, Yifan, Zhou, Peicheng, Sun, Yiwei, Lu, Zhiying, Hu, Zhangchi, Xie, Hongtao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910050092580864
author Pu, Bowei
Liu, Chuanbin
Ge, Yifan
Zhou, Peicheng
Sun, Yiwei
Lu, Zhiying
Hu, Zhangchi
Xie, Hongtao
author_facet Pu, Bowei
Liu, Chuanbin
Ge, Yifan
Zhou, Peicheng
Sun, Yiwei
Lu, Zhiying
Hu, Zhangchi
Xie, Hongtao
contents Video Large Language Models improve reasoning over complex videos by generating intermediate reasoning text. However, reliable reasoning depends on accurate video perception. In existing approaches, perception evidence is intertwined with reasoning text, making it difficult to directly supervise the perception process. We argue that reliable supervision requires explicitly separating perception evidence from reasoning so that perception can be verified independently. To supervise perception directly, we propose Decoupled Perception and Logic (DPL), which represents perception as fixed-format evidence units containing timestamps and visual descriptions. This structured representation enables direct extraction of perception content and simplifies alignment between video segments and reward evaluation. Building on DPL, we introduce a perception reward that encourages both hallucination resistance and perception-based reasoning. An Factual-Aware Evaluator (FAE) provides anti-hallucination scores and achieves hallucination evaluation performance comparable to GPT-4o. In addition, we validate reasoning consistency by feeding perception results and questions into a reference model. Experiments show that, by providing reliable process rewards, Video-DPL consistently improves post-training performance at both 3B and 7B scales, while delivering higher data efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18463
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
Pu, Bowei
Liu, Chuanbin
Ge, Yifan
Zhou, Peicheng
Sun, Yiwei
Lu, Zhiying
Hu, Zhangchi
Xie, Hongtao
Computer Vision and Pattern Recognition
I.4
Video Large Language Models improve reasoning over complex videos by generating intermediate reasoning text. However, reliable reasoning depends on accurate video perception. In existing approaches, perception evidence is intertwined with reasoning text, making it difficult to directly supervise the perception process. We argue that reliable supervision requires explicitly separating perception evidence from reasoning so that perception can be verified independently. To supervise perception directly, we propose Decoupled Perception and Logic (DPL), which represents perception as fixed-format evidence units containing timestamps and visual descriptions. This structured representation enables direct extraction of perception content and simplifies alignment between video segments and reward evaluation. Building on DPL, we introduce a perception reward that encourages both hallucination resistance and perception-based reasoning. An Factual-Aware Evaluator (FAE) provides anti-hallucination scores and achieves hallucination evaluation performance comparable to GPT-4o. In addition, we validate reasoning consistency by feeding perception results and questions into a reference model. Experiments show that, by providing reliable process rewards, Video-DPL consistently improves post-training performance at both 3B and 7B scales, while delivering higher data efficiency.
title Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
topic Computer Vision and Pattern Recognition
I.4
url https://arxiv.org/abs/2511.18463