How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Yifan, Chen, Junren, Chen, Yifan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
von: Xiong, Zidi, et al.
Veröffentlicht: (2026)
von: Xiong, Zidi, et al.
Veröffentlicht: (2026)
What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity
von: Li, Haoxi, et al.
Veröffentlicht: (2026)
von: Li, Haoxi, et al.
Veröffentlicht: (2026)
PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
von: Wu, Zongqian, et al.
Veröffentlicht: (2025)
von: Wu, Zongqian, et al.
Veröffentlicht: (2025)
Think Before You Lie: How Reasoning Leads to Honesty
von: Yuan, Ann, et al.
Veröffentlicht: (2026)
von: Yuan, Ann, et al.
Veröffentlicht: (2026)
How Much You Ate? Food Portion Estimation on Spoons
von: Sharma, Aaryam, et al.
Veröffentlicht: (2024)
von: Sharma, Aaryam, et al.
Veröffentlicht: (2024)
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
von: Xie, Can, et al.
Veröffentlicht: (2025)
von: Xie, Can, et al.
Veröffentlicht: (2025)
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2026)
Seeing with You: Perception-Reasoning Coevolution for Multimodal Reasoning
von: Miao, Ziqi, et al.
Veröffentlicht: (2026)
von: Miao, Ziqi, et al.
Veröffentlicht: (2026)
Exploitation Is All You Need... for Exploration
von: Rentschler, Micah, et al.
Veröffentlicht: (2025)
von: Rentschler, Micah, et al.
Veröffentlicht: (2025)
Look as You Think: Unifying Reasoning and Visual Evidence Attribution for Verifiable Document RAG via Reinforcement Learning
von: Liu, Shuochen, et al.
Veröffentlicht: (2025)
von: Liu, Shuochen, et al.
Veröffentlicht: (2025)
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
von: Ren, Yanwei, et al.
Veröffentlicht: (2026)
von: Ren, Yanwei, et al.
Veröffentlicht: (2026)
Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing
von: Yuan, Wenhao, et al.
Veröffentlicht: (2026)
von: Yuan, Wenhao, et al.
Veröffentlicht: (2026)
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
von: Kim, Soeun, et al.
Veröffentlicht: (2026)
von: Kim, Soeun, et al.
Veröffentlicht: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
Is Exploration All You Need? Effective Exploration Characteristics for Transfer in Reinforcement Learning
von: Balloch, Jonathan C., et al.
Veröffentlicht: (2024)
von: Balloch, Jonathan C., et al.
Veröffentlicht: (2024)
Detecting RLVR Training Data via Structural Convergence of Reasoning
von: Zhang, Hongbo, et al.
Veröffentlicht: (2026)
von: Zhang, Hongbo, et al.
Veröffentlicht: (2026)
Tensor Product Attention Is All You Need
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Driving with Regulation: Trustworthy and Interpretable Decision-Making for Autonomous Driving with Retrieval-Augmented Reasoning
von: Cai, Tianhui, et al.
Veröffentlicht: (2024)
von: Cai, Tianhui, et al.
Veröffentlicht: (2024)
Know What You Know: Metacognitive Entropy Calibration for Verifiable RL Reasoning
von: Zhao, Qiannian, et al.
Veröffentlicht: (2026)
von: Zhao, Qiannian, et al.
Veröffentlicht: (2026)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
Reasoning Is All You Need for Urban Planning AI
von: Yang, Sijie, et al.
Veröffentlicht: (2025)
von: Yang, Sijie, et al.
Veröffentlicht: (2025)
Look Before You Leap: Autonomous Exploration for LLM Agents
von: Ye, Ziang, et al.
Veröffentlicht: (2026)
von: Ye, Ziang, et al.
Veröffentlicht: (2026)
Open-Medical-R1: How to Choose Data for RLVR Training at Medicine Domain
von: Qiu, Zhongxi, et al.
Veröffentlicht: (2025)
von: Qiu, Zhongxi, et al.
Veröffentlicht: (2025)
LensWalk: Agentic Video Understanding by Planning How You See in Videos
von: Li, Keliang, et al.
Veröffentlicht: (2026)
von: Li, Keliang, et al.
Veröffentlicht: (2026)
All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks
von: Takemoto, Kazuhiro
Veröffentlicht: (2024)
von: Takemoto, Kazuhiro
Veröffentlicht: (2024)
How Many Bytes Can You Take Out Of Brain-To-Text Decoding?
von: Antonello, Richard, et al.
Veröffentlicht: (2024)
von: Antonello, Richard, et al.
Veröffentlicht: (2024)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
Does Your Optimizer Care How You Normalize? Normalization-Optimizer Coupling in LLM Training
von: Abouzeid, Abdelrahman
Veröffentlicht: (2026)
von: Abouzeid, Abdelrahman
Veröffentlicht: (2026)
Game of Trust: How Trustworthy Does Your Blockchain Think You Are?
von: Drineas, Petros, et al.
Veröffentlicht: (2025)
von: Drineas, Petros, et al.
Veröffentlicht: (2025)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
von: Liao, Haicheng, et al.
Veröffentlicht: (2025)
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
You Only Forward Once: An Efficient Compositional Judging Paradigm
von: Zhang, Tianlong, et al.
Veröffentlicht: (2025)
von: Zhang, Tianlong, et al.
Veröffentlicht: (2025)
Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning
von: Ma, Hao, et al.
Veröffentlicht: (2024)
von: Ma, Hao, et al.
Veröffentlicht: (2024)
LoRA is All You Need for Safety Alignment of Reasoning LLMs
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
What Do You Mean? Exploring How Humans and AI Interact with Symbols and Meanings in Their Interactions
von: Habibi, Reza, et al.
Veröffentlicht: (2025)
von: Habibi, Reza, et al.
Veröffentlicht: (2025)
Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning
von: Sun, Yiliu, et al.
Veröffentlicht: (2025)
von: Sun, Yiliu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
von: Xiong, Zidi, et al.
Veröffentlicht: (2026) -
What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity
von: Li, Haoxi, et al.
Veröffentlicht: (2026) -
PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention
von: Wang, Haonan, et al.
Veröffentlicht: (2025) -
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
von: Wu, Zongqian, et al.
Veröffentlicht: (2025) -
Think Before You Lie: How Reasoning Leads to Honesty
von: Yuan, Ann, et al.
Veröffentlicht: (2026)