Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Pu, Bowei, Liu, Chuanbin, Ge, Yifan, Zhou, Peicheng, Sun, Yiwei, Lu, Zhiying, Hu, Zhangchi, Xie, Hongtao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hallucination Mitigation Prompts Long-term Video Understanding
by: Sun, Yiwei, et al.
Published: (2024)
by: Sun, Yiwei, et al.
Published: (2024)
Decoupled Sensitivity-Consistency Learning for Weakly Supervised Video Anomaly Detection
by: Zheng, Hantao, et al.
Published: (2026)
by: Zheng, Hantao, et al.
Published: (2026)
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
by: Li, Yinglu, et al.
Published: (2025)
by: Li, Yinglu, et al.
Published: (2025)
From Evaluation to Defense: Advancing Safety in Video Large Language Models
by: Sun, Yiwei, et al.
Published: (2025)
by: Sun, Yiwei, et al.
Published: (2025)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
by: Li, Yayuan, et al.
Published: (2025)
by: Li, Yayuan, et al.
Published: (2025)
Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution
by: Shi, Shijun, et al.
Published: (2025)
by: Shi, Shijun, et al.
Published: (2025)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
by: Lee, Kyuho, et al.
Published: (2025)
by: Lee, Kyuho, et al.
Published: (2025)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
by: Deng, Pei, et al.
Published: (2025)
by: Deng, Pei, et al.
Published: (2025)
MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics
by: Li, Jing, et al.
Published: (2025)
by: Li, Jing, et al.
Published: (2025)
Video-Based Human Pose Regression via Decoupled Space-Time Aggregation
by: He, Jijie, et al.
Published: (2024)
by: He, Jijie, et al.
Published: (2024)
Understanding Identity Continuity in Thermal Video through Scene-Level Consistency
by: Sun, Wei-Chieh, et al.
Published: (2026)
by: Sun, Wei-Chieh, et al.
Published: (2026)
A Simple Baseline for Streaming Video Understanding
by: Shen, Yujiao, et al.
Published: (2026)
by: Shen, Yujiao, et al.
Published: (2026)
AVadCLIP: Audio-Visual Collaboration for Robust Video Anomaly Detection
by: Wu, Peng, et al.
Published: (2025)
by: Wu, Peng, et al.
Published: (2025)
HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding
by: Jin, Haopeng, et al.
Published: (2026)
by: Jin, Haopeng, et al.
Published: (2026)
Human-Centric Perception for Child Sexual Abuse Imagery
by: Laranjeira, Camila, et al.
Published: (2026)
by: Laranjeira, Camila, et al.
Published: (2026)
GraphiContact: Pose-aware Human-Scene Robust Contact Perception for Interactive Systems
by: Lin, Xiaojian, et al.
Published: (2026)
by: Lin, Xiaojian, et al.
Published: (2026)
DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion
by: Liu, Jinyuan, et al.
Published: (2025)
by: Liu, Jinyuan, et al.
Published: (2025)
Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception
by: Ou, Yiwei, et al.
Published: (2026)
by: Ou, Yiwei, et al.
Published: (2026)
Neural Fields for 3D Tracking of Anatomy and Surgical Instruments in Monocular Laparoscopic Video Clips
by: Gerats, Beerend G. A., et al.
Published: (2024)
by: Gerats, Beerend G. A., et al.
Published: (2024)
TF-Lane: Traffic Flow Module for Robust Lane Perception
by: Xie, Yihan, et al.
Published: (2026)
by: Xie, Yihan, et al.
Published: (2026)
Understanding colors of Dufaycolor: Can we recover them using historical colorimetric and spectral data?
by: Hubička, Jan, et al.
Published: (2025)
by: Hubička, Jan, et al.
Published: (2025)
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
by: Gautam, Sushant, et al.
Published: (2026)
by: Gautam, Sushant, et al.
Published: (2026)
An Immersive Multi-Elevation Multi-Seasonal Dataset for 3D Reconstruction and Visualization
by: Liu, Xijun, et al.
Published: (2024)
by: Liu, Xijun, et al.
Published: (2024)
Frequency-Decomposed INR for NIR-Assisted Low-Light RGB Image Denoising
by: Shi, Ligen, et al.
Published: (2026)
by: Shi, Ligen, et al.
Published: (2026)
Synthetic Image Detection with CLIP: Understanding and Assessing Predictive Cues
by: Willi, Marco, et al.
Published: (2026)
by: Willi, Marco, et al.
Published: (2026)
Perception-to-Pursuit: Track-Centric Temporal Reasoning for Open-World Drone Detection and Autonomous Chasing
by: Oruganti, Venkatakrishna Reddy
Published: (2026)
by: Oruganti, Venkatakrishna Reddy
Published: (2026)
Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors
by: Sun, Jiachen, et al.
Published: (2024)
by: Sun, Jiachen, et al.
Published: (2024)
Rethinking Video Deblurring with Wavelet-Aware Dynamic Transformer and Diffusion Model
by: Rao, Chen, et al.
Published: (2024)
by: Rao, Chen, et al.
Published: (2024)
RCooper: A Real-world Large-scale Dataset for Roadside Cooperative Perception
by: Hao, Ruiyang, et al.
Published: (2024)
by: Hao, Ruiyang, et al.
Published: (2024)
GStex: Per-Primitive Texturing of 2D Gaussian Splatting for Decoupled Appearance and Geometry Modeling
by: Rong, Victor, et al.
Published: (2024)
by: Rong, Victor, et al.
Published: (2024)
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
by: Shen, Meng, et al.
Published: (2026)
by: Shen, Meng, et al.
Published: (2026)
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
by: Hu, Yangliu, et al.
Published: (2025)
by: Hu, Yangliu, et al.
Published: (2025)
VidPanos: Generative Panoramic Videos from Casual Panning Videos
by: Ma, Jingwei, et al.
Published: (2024)
by: Ma, Jingwei, et al.
Published: (2024)
Detecting AI-Generated Videos with Spiking Neural Networks
by: Jang, Minsuk, et al.
Published: (2026)
by: Jang, Minsuk, et al.
Published: (2026)
EDSNet: Efficient-DSNet for Video Summarization
by: Prasad, Ashish, et al.
Published: (2024)
by: Prasad, Ashish, et al.
Published: (2024)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
by: Wu, Jason, et al.
Published: (2026)
by: Wu, Jason, et al.
Published: (2026)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
by: Wang, Yiming, et al.
Published: (2026)
by: Wang, Yiming, et al.
Published: (2026)
Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
by: Rajapaksha, Uchitha, et al.
Published: (2024)
by: Rajapaksha, Uchitha, et al.
Published: (2024)
Semantic Similarity is a Spurious Measure of Comic Understanding: Lessons Learned from Hallucinations in a Benchmarking Experiment
by: Driggers-Ellis, Christopher, et al.
Published: (2026)
by: Driggers-Ellis, Christopher, et al.
Published: (2026)
Relightable and Dynamic Gaussian Avatar Reconstruction from Monocular Video
by: Choi, Seonghwa, et al.
Published: (2025)
by: Choi, Seonghwa, et al.
Published: (2025)
Similar Items
-
Hallucination Mitigation Prompts Long-term Video Understanding
by: Sun, Yiwei, et al.
Published: (2024) -
Decoupled Sensitivity-Consistency Learning for Weakly Supervised Video Anomaly Detection
by: Zheng, Hantao, et al.
Published: (2026) -
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
by: Li, Yinglu, et al.
Published: (2025) -
From Evaluation to Defense: Advancing Safety in Video Large Language Models
by: Sun, Yiwei, et al.
Published: (2025) -
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
by: Li, Yayuan, et al.
Published: (2025)