Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jingze, Zhang, Quan, Suo, Hongfei, Cai, Zeqiang, Chen, Hongbo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification
by: Zhang, Quan, et al.
Published: (2026)
by: Zhang, Quan, et al.
Published: (2026)
Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification
by: Zhang, Quan, et al.
Published: (2026)
by: Zhang, Quan, et al.
Published: (2026)
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
by: Lu, Chaochao, et al.
Published: (2024)
by: Lu, Chaochao, et al.
Published: (2024)
Navigate Beyond Shortcuts: Debiased Learning through the Lens of Neural Collapse
by: Wang, Yining, et al.
Published: (2024)
by: Wang, Yining, et al.
Published: (2024)
Reinforcing Video Reasoning with Focused Thinking
by: Dang, Jisheng, et al.
Published: (2025)
by: Dang, Jisheng, et al.
Published: (2025)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
by: Xiao, Kelaiti, et al.
Published: (2025)
by: Xiao, Kelaiti, et al.
Published: (2025)
Video-R1: Reinforcing Video Reasoning in MLLMs
by: Feng, Kaituo, et al.
Published: (2025)
by: Feng, Kaituo, et al.
Published: (2025)
Causal Debiasing for Visual Commonsense Reasoning
by: Zou, Jiayi, et al.
Published: (2025)
by: Zou, Jiayi, et al.
Published: (2025)
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
by: Ou, Linyu, et al.
Published: (2025)
by: Ou, Linyu, et al.
Published: (2025)
The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
by: Anvekar, Tejas, et al.
Published: (2025)
by: Anvekar, Tejas, et al.
Published: (2025)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
by: Yilmaz, Nilay, et al.
Published: (2025)
by: Yilmaz, Nilay, et al.
Published: (2025)
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
by: Guo, Longteng, et al.
Published: (2026)
by: Guo, Longteng, et al.
Published: (2026)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
Reinforcing Consistency in Video MLLMs with Structured Rewards
by: Quan, Yihao, et al.
Published: (2026)
by: Quan, Yihao, et al.
Published: (2026)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems
by: Chen, Shuhang, et al.
Published: (2025)
by: Chen, Shuhang, et al.
Published: (2025)
The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
by: Zhang, Jiabei, et al.
Published: (2025)
by: Zhang, Jiabei, et al.
Published: (2025)
Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection
by: Wang, Yihui, et al.
Published: (2026)
by: Wang, Yihui, et al.
Published: (2026)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
Think 360°: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth
by: Chen, Mingrui, et al.
Published: (2026)
by: Chen, Mingrui, et al.
Published: (2026)
Causal-Inspired Multitask Learning for Video-Based Human Pose Estimation
by: Chen, Haipeng, et al.
Published: (2025)
by: Chen, Haipeng, et al.
Published: (2025)
MILO: A Lightweight Perceptual Quality Metric for Image and Latent-Space Optimization
by: Çoğalan, Uğur, et al.
Published: (2025)
by: Çoğalan, Uğur, et al.
Published: (2025)
Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior
by: Lin, Jiaying, et al.
Published: (2024)
by: Lin, Jiaying, et al.
Published: (2024)
Training-Free Reasoning and Reflection in MLLMs
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
A Causality-Inspired Model for Intima-Media Thickening Assessment in Ultrasound Videos
by: Gao, Shuo, et al.
Published: (2025)
by: Gao, Shuo, et al.
Published: (2025)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
by: Han, Su Ho, et al.
Published: (2025)
by: Han, Su Ho, et al.
Published: (2025)
G-ZAP: A Generalizable Zero-Shot Framework for Arbitrary-Scale Pansharpening
by: Yang, Zhiqi, et al.
Published: (2026)
by: Yang, Zhiqi, et al.
Published: (2026)
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
by: Zhang, Bob, et al.
Published: (2025)
by: Zhang, Bob, et al.
Published: (2025)
Motion-Adaptive Temporal Attention for Lightweight Video Generation with Stable Diffusion
by: Hong, Rui, et al.
Published: (2026)
by: Hong, Rui, et al.
Published: (2026)
A Perceptually Inspired Variational Framework for Color Enhancement
by: Palma-Amestoy, Rodrigo, et al.
Published: (2025)
by: Palma-Amestoy, Rodrigo, et al.
Published: (2025)
VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
by: Zhang, Ruifei, et al.
Published: (2025)
by: Zhang, Ruifei, et al.
Published: (2025)
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
by: Guo, Haiyang, et al.
Published: (2025)
by: Guo, Haiyang, et al.
Published: (2025)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
by: Tong, Jintao, et al.
Published: (2025)
by: Tong, Jintao, et al.
Published: (2025)
VersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment
by: Meng, Shibei, et al.
Published: (2026)
by: Meng, Shibei, et al.
Published: (2026)
Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
by: Zhao, Xinkui, et al.
Published: (2025)
by: Zhao, Xinkui, et al.
Published: (2025)
SAMIC: A Lightweight Semantic-Aware Mamba for Efficient Perceptual Image Compression
by: Zhang, Jiaqian, et al.
Published: (2026)
by: Zhang, Jiaqian, et al.
Published: (2026)
Beyond Semantic Priors: Mitigating Optimization Collapse for Generalizable Visual Forensics
by: Liu, Jipeng, et al.
Published: (2026)
by: Liu, Jipeng, et al.
Published: (2026)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
by: Fang, Bo, et al.
Published: (2025)
by: Fang, Bo, et al.
Published: (2025)
Similar Items
-
View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification
by: Zhang, Quan, et al.
Published: (2026) -
Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification
by: Zhang, Quan, et al.
Published: (2026) -
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
by: Lu, Chaochao, et al.
Published: (2024) -
Navigate Beyond Shortcuts: Debiased Learning through the Lens of Neural Collapse
by: Wang, Yining, et al.
Published: (2024) -
Reinforcing Video Reasoning with Focused Thinking
by: Dang, Jisheng, et al.
Published: (2025)