RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Kaiyuan, Pang, Jing-Cheng, Yu, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
von: Wei, Zhepei, et al.
Veröffentlicht: (2026)
von: Wei, Zhepei, et al.
Veröffentlicht: (2026)
How Far Can Unsupervised RLVR Scale LLM Training?
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
Linear Dynamics in the RLVR Training of Large Language Models
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
Self-Distilled RLVR
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
von: Cui, Sijia, et al.
Veröffentlicht: (2026)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
von: Liu, Yijun, et al.
Veröffentlicht: (2024)
von: Liu, Yijun, et al.
Veröffentlicht: (2024)
Efficient RLVR Training via Weighted Mutual Information Data Selection
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
von: Yakushev, George, et al.
Veröffentlicht: (2025)
von: Yakushev, George, et al.
Veröffentlicht: (2025)
Agentic Adversarial QA for Improving Domain-Specific LLMs
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
Does Refusal Training in LLMs Generalize to the Past Tense?
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
von: Burgess, James, et al.
Veröffentlicht: (2026)
von: Burgess, James, et al.
Veröffentlicht: (2026)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2026)
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2026)
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
von: Zhao, Jiale, et al.
Veröffentlicht: (2026)
von: Zhao, Jiale, et al.
Veröffentlicht: (2026)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghye, et al.
Veröffentlicht: (2026)
RLPR: Extrapolating RLVR to General Domains without Verifiers
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
The Unlearnability Phenomenon in RLVR for Language Models
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing
von: Xie, Pei-Xi, et al.
Veröffentlicht: (2026)
von: Xie, Pei-Xi, et al.
Veröffentlicht: (2026)
Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity
von: Lochab, Anamika, et al.
Veröffentlicht: (2026)
von: Lochab, Anamika, et al.
Veröffentlicht: (2026)
Learning to Correct for QA Reasoning with Black-box LLMs
von: Kim, Jaehyung, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyung, et al.
Veröffentlicht: (2024)
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
von: Meng, Haoming, et al.
Veröffentlicht: (2026)
von: Meng, Haoming, et al.
Veröffentlicht: (2026)
Generalization of RLVR Using Causal Reasoning as a Testbed
von: Lu, Brian, et al.
Veröffentlicht: (2025)
von: Lu, Brian, et al.
Veröffentlicht: (2025)
SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
von: Xu, Haozhou, et al.
Veröffentlicht: (2025)
von: Xu, Haozhou, et al.
Veröffentlicht: (2025)
PeruMedQA: Benchmarking Large Language Models (LLMs) on Peruvian Medical Exams -- Dataset Construction and Evaluation
von: Carrillo-Larco, Rodrigo M., et al.
Veröffentlicht: (2025)
von: Carrillo-Larco, Rodrigo M., et al.
Veröffentlicht: (2025)
Pre-trained Language Models Improve the Few-shot Prompt Ability of Decision Transformer
von: Yang, Yu, et al.
Veröffentlicht: (2024)
von: Yang, Yu, et al.
Veröffentlicht: (2024)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
Compressing LLMs: The Truth is Rarely Pure and Never Simple
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2023)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2023)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
von: Wang, Shouren, et al.
Veröffentlicht: (2025)
von: Wang, Shouren, et al.
Veröffentlicht: (2025)
Continuous Approximations for Improving Quantization Aware Training of LLMs
von: Li, He, et al.
Veröffentlicht: (2024)
von: Li, He, et al.
Veröffentlicht: (2024)
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
von: Yang, Wang, et al.
Veröffentlicht: (2025)
von: Yang, Wang, et al.
Veröffentlicht: (2025)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
von: Chen, Feng, et al.
Veröffentlicht: (2024)
von: Chen, Feng, et al.
Veröffentlicht: (2024)
CodeMirage: A Multi-Lingual Benchmark for Detecting AI-Generated and Paraphrased Source Code from Production-Level LLMs
von: Guo, Hanxi, et al.
Veröffentlicht: (2025)
von: Guo, Hanxi, et al.
Veröffentlicht: (2025)
Rewards as Labels: Revisiting RLVR from a Classification Perspective
von: Zhai, Zepeng, et al.
Veröffentlicht: (2026)
von: Zhai, Zepeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
von: Wei, Zhepei, et al.
Veröffentlicht: (2026) -
How Far Can Unsupervised RLVR Scale LLM Training?
von: He, Bingxiang, et al.
Veröffentlicht: (2026) -
Linear Dynamics in the RLVR Training of Large Language Models
von: Wang, Tianle, et al.
Veröffentlicht: (2026) -
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
von: Yan, Lecheng, et al.
Veröffentlicht: (2026) -
Self-Distilled RLVR
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)