Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Jinda, Wu, Junkang, Li, Jinghan, Huang, Kexin, Yang, Shuo, Chen, Mingzhu, Wu, Jiancan, Liu, Kuien, Wang, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
by: Lu, Jinda, et al.
Published: (2026)
by: Lu, Jinda, et al.
Published: (2026)
AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
by: Lu, Jinda, et al.
Published: (2025)
by: Lu, Jinda, et al.
Published: (2025)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
by: Lu, Jinda, et al.
Published: (2025)
by: Lu, Jinda, et al.
Published: (2025)
Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model
by: Wang, Yuan, et al.
Published: (2026)
by: Wang, Yuan, et al.
Published: (2026)
Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era
by: Hong, Qiuhe, et al.
Published: (2026)
by: Hong, Qiuhe, et al.
Published: (2026)
Pay Attention to Where You Looked
by: Berian, Alex, et al.
Published: (2026)
by: Berian, Alex, et al.
Published: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
by: Fuller, Anthony, et al.
Published: (2025)
by: Fuller, Anthony, et al.
Published: (2025)
Rethinking Visual Content Refinement in Low-Shot CLIP Adaptation
by: Lu, Jinda, et al.
Published: (2024)
by: Lu, Jinda, et al.
Published: (2024)
Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization
by: Li, Jinghan, et al.
Published: (2026)
by: Li, Jinghan, et al.
Published: (2026)
ClickAttention: Click Region Similarity Guided Interactive Segmentation
by: Xu, Long, et al.
Published: (2024)
by: Xu, Long, et al.
Published: (2024)
Aligning Multimodal LLM with Human Preference: A Survey
by: Yu, Tao, et al.
Published: (2025)
by: Yu, Tao, et al.
Published: (2025)
Boosting Few-Shot Learning via Attentive Feature Regularization
by: Zhu, Xingyu, et al.
Published: (2024)
by: Zhu, Xingyu, et al.
Published: (2024)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Enhancing Tree Type Detection in Forest Fire Risk Assessment: Multi-Stage Approach and Color Encoding with Forest Fire Risk Evaluation Framework for UAV Imagery
by: Zhang, Jinda
Published: (2024)
by: Zhang, Jinda
Published: (2024)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
by: Shen, Yuxiang, et al.
Published: (2026)
by: Shen, Yuxiang, et al.
Published: (2026)
Learning Procedural-aware Video Representations through State-Grounded Hierarchy Unfolding
by: Zhao, Jinghan, et al.
Published: (2025)
by: Zhao, Jinghan, et al.
Published: (2025)
GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory Prediction
by: Pei, Muleilan, et al.
Published: (2025)
by: Pei, Muleilan, et al.
Published: (2025)
RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
by: Wu, Jiahe, et al.
Published: (2026)
by: Wu, Jiahe, et al.
Published: (2026)
SFOOD: A Multimodal Benchmark for Comprehensive Food Attribute Analysis Beyond RGB with Spectral Insights
by: Xu, Zhenbo, et al.
Published: (2025)
by: Xu, Zhenbo, et al.
Published: (2025)
Accelerating Diffusion Transformer via Gradient-Optimized Cache
by: Qiu, Junxiang, et al.
Published: (2025)
by: Qiu, Junxiang, et al.
Published: (2025)
Look-Ahead and Look-Back Flows: Training-Free Image Generation with Trajectory Smoothing
by: Luo, Yan, et al.
Published: (2026)
by: Luo, Yan, et al.
Published: (2026)
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling
by: Wang, Yuan, et al.
Published: (2026)
by: Wang, Yuan, et al.
Published: (2026)
RED: Robust Environmental Design
by: Yang, Jinghan
Published: (2024)
by: Yang, Jinghan
Published: (2024)
Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
by: Hu, Ruina, et al.
Published: (2026)
by: Hu, Ruina, et al.
Published: (2026)
Densely Distilling Cumulative Knowledge for Continual Learning
by: Shi, Zenglin, et al.
Published: (2024)
by: Shi, Zenglin, et al.
Published: (2024)
CAST-LUT: Tokenizer-Guided HSV Look-Up Tables for Purple Flare Removal
by: Wang, Pu, et al.
Published: (2025)
by: Wang, Pu, et al.
Published: (2025)
R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
by: Zhang, Zirui, et al.
Published: (2026)
by: Zhang, Zirui, et al.
Published: (2026)
Accelerating Diffusion Transformer via Error-Optimized Cache
by: Qiu, Junxiang, et al.
Published: (2025)
by: Qiu, Junxiang, et al.
Published: (2025)
Reinforcement Learning meets Masked Video Modeling : Trajectory-Guided Adaptive Token Selection
by: Rai, Ayush K., et al.
Published: (2025)
by: Rai, Ayush K., et al.
Published: (2025)
Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search
by: Aljehrai, Faisal, et al.
Published: (2026)
by: Aljehrai, Faisal, et al.
Published: (2026)
Learning Where to Edit Vision Transformers
by: Yang, Yunqiao, et al.
Published: (2024)
by: Yang, Yunqiao, et al.
Published: (2024)
Multimodal Label Relevance Ranking via Reinforcement Learning
by: Guo, Taian, et al.
Published: (2024)
by: Guo, Taian, et al.
Published: (2024)
Where do Large Vision-Language Models Look at when Answering Questions?
by: Xing, Xiaoying, et al.
Published: (2025)
by: Xing, Xiaoying, et al.
Published: (2025)
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
by: Yang, Siqi, et al.
Published: (2025)
by: Yang, Siqi, et al.
Published: (2025)
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
by: Salamatian, Ali, et al.
Published: (2026)
by: Salamatian, Ali, et al.
Published: (2026)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
by: Huang, Kexin, et al.
Published: (2026)
by: Huang, Kexin, et al.
Published: (2026)
ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe
by: Bai, Yifan, et al.
Published: (2023)
by: Bai, Yifan, et al.
Published: (2023)
Similar Items
-
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
by: Lu, Jinda, et al.
Published: (2026) -
AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
by: Lu, Jinda, et al.
Published: (2025) -
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
by: Lu, Jinda, et al.
Published: (2025) -
Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model
by: Wang, Yuan, et al.
Published: (2026) -
Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era
by: Hong, Qiuhe, et al.
Published: (2026)