Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Jingze, Zhang, Quan, Suo, Hongfei, Cai, Zeqiang, Chen, Hongbo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification
por: Zhang, Quan, et al.
Publicado: (2026)
por: Zhang, Quan, et al.
Publicado: (2026)
Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification
por: Zhang, Quan, et al.
Publicado: (2026)
por: Zhang, Quan, et al.
Publicado: (2026)
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
por: Lu, Chaochao, et al.
Publicado: (2024)
por: Lu, Chaochao, et al.
Publicado: (2024)
Navigate Beyond Shortcuts: Debiased Learning through the Lens of Neural Collapse
por: Wang, Yining, et al.
Publicado: (2024)
por: Wang, Yining, et al.
Publicado: (2024)
Reinforcing Video Reasoning with Focused Thinking
por: Dang, Jisheng, et al.
Publicado: (2025)
por: Dang, Jisheng, et al.
Publicado: (2025)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
por: Xiao, Kelaiti, et al.
Publicado: (2025)
por: Xiao, Kelaiti, et al.
Publicado: (2025)
Video-R1: Reinforcing Video Reasoning in MLLMs
por: Feng, Kaituo, et al.
Publicado: (2025)
por: Feng, Kaituo, et al.
Publicado: (2025)
Causal Debiasing for Visual Commonsense Reasoning
por: Zou, Jiayi, et al.
Publicado: (2025)
por: Zou, Jiayi, et al.
Publicado: (2025)
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
por: Ou, Linyu, et al.
Publicado: (2025)
por: Ou, Linyu, et al.
Publicado: (2025)
The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
por: Anvekar, Tejas, et al.
Publicado: (2025)
por: Anvekar, Tejas, et al.
Publicado: (2025)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
por: Yilmaz, Nilay, et al.
Publicado: (2025)
por: Yilmaz, Nilay, et al.
Publicado: (2025)
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
por: Guo, Longteng, et al.
Publicado: (2026)
por: Guo, Longteng, et al.
Publicado: (2026)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
por: Ouyang, Kun, et al.
Publicado: (2025)
por: Ouyang, Kun, et al.
Publicado: (2025)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
por: Liu, Yuanxin, et al.
Publicado: (2025)
por: Liu, Yuanxin, et al.
Publicado: (2025)
Reinforcing Consistency in Video MLLMs with Structured Rewards
por: Quan, Yihao, et al.
Publicado: (2026)
por: Quan, Yihao, et al.
Publicado: (2026)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
por: Guo, Hao, et al.
Publicado: (2026)
por: Guo, Hao, et al.
Publicado: (2026)
MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems
por: Chen, Shuhang, et al.
Publicado: (2025)
por: Chen, Shuhang, et al.
Publicado: (2025)
The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
por: Zhang, Jiabei, et al.
Publicado: (2025)
por: Zhang, Jiabei, et al.
Publicado: (2025)
Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection
por: Wang, Yihui, et al.
Publicado: (2026)
por: Wang, Yihui, et al.
Publicado: (2026)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
por: Zhu, Rui, et al.
Publicado: (2026)
por: Zhu, Rui, et al.
Publicado: (2026)
Think 360°: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth
por: Chen, Mingrui, et al.
Publicado: (2026)
por: Chen, Mingrui, et al.
Publicado: (2026)
Causal-Inspired Multitask Learning for Video-Based Human Pose Estimation
por: Chen, Haipeng, et al.
Publicado: (2025)
por: Chen, Haipeng, et al.
Publicado: (2025)
MILO: A Lightweight Perceptual Quality Metric for Image and Latent-Space Optimization
por: Çoğalan, Uğur, et al.
Publicado: (2025)
por: Çoğalan, Uğur, et al.
Publicado: (2025)
Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior
por: Lin, Jiaying, et al.
Publicado: (2024)
por: Lin, Jiaying, et al.
Publicado: (2024)
Training-Free Reasoning and Reflection in MLLMs
por: Wei, Hongchen, et al.
Publicado: (2025)
por: Wei, Hongchen, et al.
Publicado: (2025)
A Causality-Inspired Model for Intima-Media Thickening Assessment in Ultrasound Videos
por: Gao, Shuo, et al.
Publicado: (2025)
por: Gao, Shuo, et al.
Publicado: (2025)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
por: Wang, Qi, et al.
Publicado: (2025)
por: Wang, Qi, et al.
Publicado: (2025)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
por: Han, Su Ho, et al.
Publicado: (2025)
por: Han, Su Ho, et al.
Publicado: (2025)
G-ZAP: A Generalizable Zero-Shot Framework for Arbitrary-Scale Pansharpening
por: Yang, Zhiqi, et al.
Publicado: (2026)
por: Yang, Zhiqi, et al.
Publicado: (2026)
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
por: Zhang, Bob, et al.
Publicado: (2025)
por: Zhang, Bob, et al.
Publicado: (2025)
Motion-Adaptive Temporal Attention for Lightweight Video Generation with Stable Diffusion
por: Hong, Rui, et al.
Publicado: (2026)
por: Hong, Rui, et al.
Publicado: (2026)
A Perceptually Inspired Variational Framework for Color Enhancement
por: Palma-Amestoy, Rodrigo, et al.
Publicado: (2025)
por: Palma-Amestoy, Rodrigo, et al.
Publicado: (2025)
VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
por: Zhang, Ruifei, et al.
Publicado: (2025)
por: Zhang, Ruifei, et al.
Publicado: (2025)
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
por: Guo, Haiyang, et al.
Publicado: (2025)
por: Guo, Haiyang, et al.
Publicado: (2025)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
por: Tong, Jintao, et al.
Publicado: (2025)
por: Tong, Jintao, et al.
Publicado: (2025)
VersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment
por: Meng, Shibei, et al.
Publicado: (2026)
por: Meng, Shibei, et al.
Publicado: (2026)
Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
por: Zhao, Xinkui, et al.
Publicado: (2025)
por: Zhao, Xinkui, et al.
Publicado: (2025)
SAMIC: A Lightweight Semantic-Aware Mamba for Efficient Perceptual Image Compression
por: Zhang, Jiaqian, et al.
Publicado: (2026)
por: Zhang, Jiaqian, et al.
Publicado: (2026)
Beyond Semantic Priors: Mitigating Optimization Collapse for Generalizable Visual Forensics
por: Liu, Jipeng, et al.
Publicado: (2026)
por: Liu, Jipeng, et al.
Publicado: (2026)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
por: Fang, Bo, et al.
Publicado: (2025)
por: Fang, Bo, et al.
Publicado: (2025)
Ejemplares similares
-
View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification
por: Zhang, Quan, et al.
Publicado: (2026) -
Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification
por: Zhang, Quan, et al.
Publicado: (2026) -
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
por: Lu, Chaochao, et al.
Publicado: (2024) -
Navigate Beyond Shortcuts: Debiased Learning through the Lens of Neural Collapse
por: Wang, Yining, et al.
Publicado: (2024) -
Reinforcing Video Reasoning with Focused Thinking
por: Dang, Jisheng, et al.
Publicado: (2025)