Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Minghe, Zhi, Zhuo, Liu, Chonghan, Xing, Shuo, Tu, Zhengzhong, Liu, Che |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
by: Pan, Chenbin, et al.
Published: (2025)
by: Pan, Chenbin, et al.
Published: (2025)
The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
by: Yao, Xinhao, et al.
Published: (2025)
by: Yao, Xinhao, et al.
Published: (2025)
How Does Controllability Emerge In Language Models During Pretraining?
by: She, Jianshu, et al.
Published: (2025)
by: She, Jianshu, et al.
Published: (2025)
Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving
by: Luo, Xuewen, et al.
Published: (2025)
by: Luo, Xuewen, et al.
Published: (2025)
Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance
by: Yang, Fengze, et al.
Published: (2025)
by: Yang, Fengze, et al.
Published: (2025)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
by: Tian, Kexin, et al.
Published: (2025)
by: Tian, Kexin, et al.
Published: (2025)
Towards Long-window Anchoring in Vision-Language Model Distillation
by: Zhou, Haoyi, et al.
Published: (2025)
by: Zhou, Haoyi, et al.
Published: (2025)
Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models
by: McGinness, Lachlan, et al.
Published: (2025)
by: McGinness, Lachlan, et al.
Published: (2025)
CAPS: Cascaded Adaptive Pairwise Selection for Efficient Parallel Reasoning
by: Lin, Fangzhou, et al.
Published: (2026)
by: Lin, Fangzhou, et al.
Published: (2026)
PathCal: State-Aware Reflection-Marker Calibration for Efficient Reasoning
by: Jiang, Lingyu, et al.
Published: (2026)
by: Jiang, Lingyu, et al.
Published: (2026)
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
by: Pan, Jiazhen, et al.
Published: (2025)
by: Pan, Jiazhen, et al.
Published: (2025)
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
by: Zhang, Zhaowei, et al.
Published: (2026)
by: Zhang, Zhaowei, et al.
Published: (2026)
Generalization of RLVR Using Causal Reasoning as a Testbed
by: Lu, Brian, et al.
Published: (2025)
by: Lu, Brian, et al.
Published: (2025)
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
by: Xing, Shuo, et al.
Published: (2025)
by: Xing, Shuo, et al.
Published: (2025)
RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
by: Chen, Qiguang, et al.
Published: (2025)
by: Chen, Qiguang, et al.
Published: (2025)
Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
by: Yin, Cheng, et al.
Published: (2025)
by: Yin, Cheng, et al.
Published: (2025)
Think in Sentences: Explicit Sentence Boundaries Enhance Language Model's Capabilities
by: Liu, Zhichen, et al.
Published: (2026)
by: Liu, Zhichen, et al.
Published: (2026)
T2T-VICL: Unlocking the Boundaries of Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
by: Xia, Shao-Jun, et al.
Published: (2025)
by: Xia, Shao-Jun, et al.
Published: (2025)
DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization
by: Li, Gang, et al.
Published: (2025)
by: Li, Gang, et al.
Published: (2025)
SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning
by: Xiang, Kun, et al.
Published: (2026)
by: Xiang, Kun, et al.
Published: (2026)
DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
by: Godbole, Mihir, et al.
Published: (2025)
by: Godbole, Mihir, et al.
Published: (2025)
Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
by: Zhang, Qingjie, et al.
Published: (2025)
by: Zhang, Qingjie, et al.
Published: (2025)
ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
by: Yang, Junyao, et al.
Published: (2026)
by: Yang, Junyao, et al.
Published: (2026)
Distilling Mathematical Reasoning Capabilities into Small Language Models
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Extending RLVR to Open-Ended Tasks via Verifiable Multiple-Choice Reformulation
by: Zhang, Mengyu, et al.
Published: (2025)
by: Zhang, Mengyu, et al.
Published: (2025)
Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models
by: Chen, Qiguang, et al.
Published: (2025)
by: Chen, Qiguang, et al.
Published: (2025)
Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
by: Bai, Qianqian, et al.
Published: (2025)
by: Bai, Qianqian, et al.
Published: (2025)
Zero-shot Object Navigation with Vision-Language Models Reasoning
by: Wen, Congcong, et al.
Published: (2024)
by: Wen, Congcong, et al.
Published: (2024)
ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models
by: Qin, Chonghan, et al.
Published: (2026)
by: Qin, Chonghan, et al.
Published: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
by: Huang, Kexin, et al.
Published: (2026)
by: Huang, Kexin, et al.
Published: (2026)
Geometry of Knowledge Allows Extending Diversity Boundaries of Large Language Models
by: Bystroński, Mateusz, et al.
Published: (2025)
by: Bystroński, Mateusz, et al.
Published: (2025)
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
by: Xie, Can, et al.
Published: (2025)
by: Xie, Can, et al.
Published: (2025)
CAPF: Guiding Search-Agent Rollouts with Credit-Attenuated Privileged Feedback
by: Chen, Bin, et al.
Published: (2026)
by: Chen, Bin, et al.
Published: (2026)
Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks
by: Zhu, Yihua, et al.
Published: (2026)
by: Zhu, Yihua, et al.
Published: (2026)
GraphInstruct: Empowering Large Language Models with Graph Understanding and Reasoning Capability
by: Luo, Zihan, et al.
Published: (2024)
by: Luo, Zihan, et al.
Published: (2024)
AdaRing: Towards Ultra-Light Vision-Language Adaptation via Cross-Layer Tensor Ring Decomposition
by: Huang, Ying, et al.
Published: (2025)
by: Huang, Ying, et al.
Published: (2025)
Similar Items
-
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
by: Pan, Chenbin, et al.
Published: (2025) -
The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
by: Yao, Xinhao, et al.
Published: (2025) -
How Does Controllability Emerge In Language Models During Pretraining?
by: She, Jianshu, et al.
Published: (2025) -
Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models
by: Li, Yuan, et al.
Published: (2025) -
V2X-UniPool: Unifying Multimodal Perception and Knowledge Reasoning for Autonomous Driving
by: Luo, Xuewen, et al.
Published: (2025)