Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Ruina, Wang, Chen, Wei, Lai, Bai, Jionghao, Yu, Bin, Huang, Weiran, Wang, Kai, Wang, Yue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
von: Wei, Lai, et al.
Veröffentlicht: (2025)
von: Wei, Lai, et al.
Veröffentlicht: (2025)
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
von: Wei, Lai, et al.
Veröffentlicht: (2025)
von: Wei, Lai, et al.
Veröffentlicht: (2025)
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
von: Ma, Dongsheng, et al.
Veröffentlicht: (2026)
von: Ma, Dongsheng, et al.
Veröffentlicht: (2026)
Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation
von: Luo, Weiqing, et al.
Veröffentlicht: (2026)
von: Luo, Weiqing, et al.
Veröffentlicht: (2026)
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
von: Wei, Lai, et al.
Veröffentlicht: (2026)
von: Wei, Lai, et al.
Veröffentlicht: (2026)
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2023)
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2023)
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
von: Braun, Tobias, et al.
Veröffentlicht: (2024)
von: Braun, Tobias, et al.
Veröffentlicht: (2024)
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
Open-Vocabulary Federated Learning with Multimodal Prototyping
von: Zeng, Huimin, et al.
Veröffentlicht: (2024)
von: Zeng, Huimin, et al.
Veröffentlicht: (2024)
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
von: Ma, Chuang, et al.
Veröffentlicht: (2026)
von: Ma, Chuang, et al.
Veröffentlicht: (2026)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
MultiST: A Cross-Attention-Based Multimodal Model for Spatial Transcriptomic
von: Wang, Wei, et al.
Veröffentlicht: (2026)
von: Wang, Wei, et al.
Veröffentlicht: (2026)
WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
von: Jiang, Yiwen, et al.
Veröffentlicht: (2025)
von: Jiang, Yiwen, et al.
Veröffentlicht: (2025)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
von: Wei, Lai, et al.
Veröffentlicht: (2023)
von: Wei, Lai, et al.
Veröffentlicht: (2023)
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2025)
Continual SFT Matches Multimodal RLHF with Negative Supervision
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
VITED: Video Temporal Evidence Distillation
von: Lu, Yujie, et al.
Veröffentlicht: (2025)
von: Lu, Yujie, et al.
Veröffentlicht: (2025)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
von: Liu, Wenjie, et al.
Veröffentlicht: (2026)
von: Liu, Wenjie, et al.
Veröffentlicht: (2026)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
von: Wang, Yabing, et al.
Veröffentlicht: (2024)
von: Wang, Yabing, et al.
Veröffentlicht: (2024)
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
von: Guo, Jarvis, et al.
Veröffentlicht: (2024)
von: Guo, Jarvis, et al.
Veröffentlicht: (2024)
DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination
von: Gong, Xuan, et al.
Veröffentlicht: (2024)
von: Gong, Xuan, et al.
Veröffentlicht: (2024)
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools
von: Qi, Ji, et al.
Veröffentlicht: (2023)
von: Qi, Ji, et al.
Veröffentlicht: (2023)
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and Disinformation
von: Yuan, Xin, et al.
Veröffentlicht: (2023)
von: Yuan, Xin, et al.
Veröffentlicht: (2023)
MSG-Chart: Multimodal Scene Graph for ChartQA
von: Dai, Yue, et al.
Veröffentlicht: (2024)
von: Dai, Yue, et al.
Veröffentlicht: (2024)
Look Twice: Training-Free Evidence Highlighting in Multimodal Large Language Models
von: Morini, Marco, et al.
Veröffentlicht: (2026)
von: Morini, Marco, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
von: Tian, Changyuan, et al.
Veröffentlicht: (2026) -
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
von: Wei, Lai, et al.
Veröffentlicht: (2025) -
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026) -
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
von: Wei, Lai, et al.
Veröffentlicht: (2025) -
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
von: Ma, Dongsheng, et al.
Veröffentlicht: (2026)