Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Jiaer, Zang, Yuhang, Gao, Peng, Li, Sharon, Zhou, Kaiyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
by: Tong, Bingkui, et al.
Published: (2025)
by: Tong, Bingkui, et al.
Published: (2025)
Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
by: Xia, Jiaer, et al.
Published: (2025)
by: Xia, Jiaer, et al.
Published: (2025)
Measuring Epistemic Humility in Multimodal Large Language Models
by: Tong, Bingkui, et al.
Published: (2025)
by: Tong, Bingkui, et al.
Published: (2025)
Streaming Video Instruction Tuning
by: Xia, Jiaer, et al.
Published: (2025)
by: Xia, Jiaer, et al.
Published: (2025)
Learning to Think Fast and Slow for Visual Language Models
by: Lin, Chenyu, et al.
Published: (2025)
by: Lin, Chenyu, et al.
Published: (2025)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
Visual-RFT: Visual Reinforcement Fine-Tuning
by: Liu, Ziyu, et al.
Published: (2025)
by: Liu, Ziyu, et al.
Published: (2025)
Contextual Object Detection with Multimodal Large Language Models
by: Zang, Yuhang, et al.
Published: (2023)
by: Zang, Yuhang, et al.
Published: (2023)
MedReason-R1: Learning to Reason for CT Diagnosis with Reinforcement Learning and Local Zoom
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
by: Ding, Shengyuan, et al.
Published: (2025)
by: Ding, Shengyuan, et al.
Published: (2025)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
by: Bai, Sule, et al.
Published: (2025)
by: Bai, Sule, et al.
Published: (2025)
Gaze-directed Vision GNN for Mitigating Shortcut Learning in Medical Image
by: Wu, Shaoxuan, et al.
Published: (2024)
by: Wu, Shaoxuan, et al.
Published: (2024)
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
by: Cheng, Zixu, et al.
Published: (2026)
by: Cheng, Zixu, et al.
Published: (2026)
STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning
by: Zhang, Xiaowen, et al.
Published: (2026)
by: Zhang, Xiaowen, et al.
Published: (2026)
Take A Shortcut Back: Mitigating the Gradient Vanishing for Training Spiking Neural Networks
by: Guo, Yufei, et al.
Published: (2024)
by: Guo, Yufei, et al.
Published: (2024)
Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
by: Wang, Peiyu, et al.
Published: (2025)
by: Wang, Peiyu, et al.
Published: (2025)
Visual Agentic Reinforcement Fine-Tuning
by: Liu, Ziyu, et al.
Published: (2025)
by: Liu, Ziyu, et al.
Published: (2025)
HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
by: Zhang, Zhenhao, et al.
Published: (2025)
by: Zhang, Zhenhao, et al.
Published: (2025)
Visual Self-Refine: A Pixel-Guided Paradigm for Accurate Chart Parsing
by: Li, Jinsong, et al.
Published: (2026)
by: Li, Jinsong, et al.
Published: (2026)
Video-R1: Reinforcing Video Reasoning in MLLMs
by: Feng, Kaituo, et al.
Published: (2025)
by: Feng, Kaituo, et al.
Published: (2025)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
by: Liu, Yuqi, et al.
Published: (2025)
by: Liu, Yuqi, et al.
Published: (2025)
Efficient Unsupervised Shortcut Learning Detection and Mitigation in Transformers
by: Kuhn, Lukas, et al.
Published: (2025)
by: Kuhn, Lukas, et al.
Published: (2025)
Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank
by: Wu, Tianhe, et al.
Published: (2025)
by: Wu, Tianhe, et al.
Published: (2025)
Grounded Reinforcement Learning for Visual Reasoning
by: Sarch, Gabriel, et al.
Published: (2025)
by: Sarch, Gabriel, et al.
Published: (2025)
Touch-R1: Reinforcing Touch Reasoning in MLLMs
by: Lai, Yingxin, et al.
Published: (2026)
by: Lai, Yingxin, et al.
Published: (2026)
The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space
by: Hu, Xia, et al.
Published: (2026)
by: Hu, Xia, et al.
Published: (2026)
BackMix: Mitigating Shortcut Learning in Echocardiography with Minimal Supervision
by: Bransby, Kit Mills, et al.
Published: (2024)
by: Bransby, Kit Mills, et al.
Published: (2024)
Wan-R1: Verifiable-Reinforcement Learning for Video Reasoning
by: Liu, Ming, et al.
Published: (2026)
by: Liu, Ming, et al.
Published: (2026)
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models
by: Lai, Yuxiang, et al.
Published: (2025)
by: Lai, Yuxiang, et al.
Published: (2025)
PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks
by: Wu, Jianyu, et al.
Published: (2025)
by: Wu, Jianyu, et al.
Published: (2025)
Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
by: Zhao, Shijie, et al.
Published: (2025)
by: Zhao, Shijie, et al.
Published: (2025)
Attention Disturbance and Dual-Path Constraint Network for Occluded Person Re-identification
by: Xia, Jiaer, et al.
Published: (2023)
by: Xia, Jiaer, et al.
Published: (2023)
VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning
by: Xu, Zishan, et al.
Published: (2025)
by: Xu, Zishan, et al.
Published: (2025)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
by: Lee, Dosung, et al.
Published: (2025)
by: Lee, Dosung, et al.
Published: (2025)
ClearGCD: Mitigating Shortcut Learning For Robust Generalized Category Discovery
by: Lyu, Kailin, et al.
Published: (2025)
by: Lyu, Kailin, et al.
Published: (2025)
GMAI-VL-R1: Harnessing Reinforcement Learning for Multimodal Medical Reasoning
by: Su, Yanzhou, et al.
Published: (2025)
by: Su, Yanzhou, et al.
Published: (2025)
Similar Items
-
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
by: Tong, Bingkui, et al.
Published: (2025) -
Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
by: Xia, Jiaer, et al.
Published: (2025) -
Measuring Epistemic Humility in Multimodal Large Language Models
by: Tong, Bingkui, et al.
Published: (2025) -
Streaming Video Instruction Tuning
by: Xia, Jiaer, et al.
Published: (2025) -
Learning to Think Fast and Slow for Visual Language Models
by: Lin, Chenyu, et al.
Published: (2025)