Diversity-Incentivized Exploration for Versatile Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Zican, Zhang, Shilin, Li, Yafu, Yan, Jianhao, Hu, Xuyang, Cui, Leyang, Qu, Xiaoye, Chen, Chunlin, Cheng, Yu, Wang, Zhi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Reason under Off-Policy Guidance
by: Yan, Jianhao, et al.
Published: (2025)
by: Yan, Jianhao, et al.
Published: (2025)
Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning
by: Hu, Zican, et al.
Published: (2025)
by: Hu, Zican, et al.
Published: (2025)
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
by: Fu, Tingchen, et al.
Published: (2025)
by: Fu, Tingchen, et al.
Published: (2025)
Text-to-Decision Agent: Offline Meta-Reinforcement Learning from Natural Language Supervision
by: Zhang, Shilin, et al.
Published: (2025)
by: Zhang, Shilin, et al.
Published: (2025)
SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
by: Shen, Chuming, et al.
Published: (2025)
by: Shen, Chuming, et al.
Published: (2025)
ExGRPO: Learning to Reason from Experience
by: Zhan, Runzhe, et al.
Published: (2025)
by: Zhan, Runzhe, et al.
Published: (2025)
Potential and Challenges of Model Editing for Social Debiasing
by: Yan, Jianhao, et al.
Published: (2024)
by: Yan, Jianhao, et al.
Published: (2024)
Mixture-of-Experts Meets In-Context Reinforcement Learning
by: Wu, Wenhao, et al.
Published: (2025)
by: Wu, Wenhao, et al.
Published: (2025)
CoTEvol: Self-Evolving Chain-of-Thoughts for Data Synthesis in Mathematical Reasoning
by: Wang, Zhuo, et al.
Published: (2026)
by: Wang, Zhuo, et al.
Published: (2026)
Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
by: Li, Yafu, et al.
Published: (2025)
by: Li, Yafu, et al.
Published: (2025)
Spotting AI's Touch: Identifying LLM-Paraphrased Spans in Text
by: Li, Yafu, et al.
Published: (2024)
by: Li, Yafu, et al.
Published: (2024)
From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG
by: Wu, Wenhao, et al.
Published: (2026)
by: Wu, Wenhao, et al.
Published: (2026)
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
by: Chen, Shuang, et al.
Published: (2025)
by: Chen, Shuang, et al.
Published: (2025)
Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation
by: Liu, Jinmei, et al.
Published: (2026)
by: Liu, Jinmei, et al.
Published: (2026)
Timo: Towards Better Temporal Reasoning for Language Models
by: Su, Zhaochen, et al.
Published: (2024)
by: Su, Zhaochen, et al.
Published: (2024)
Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning
by: Qu, Xiaoye, et al.
Published: (2024)
by: Qu, Xiaoye, et al.
Published: (2024)
Continual Offline Reinforcement Learning via Diffusion-based Dual Generative Replay
by: Liu, Jinmei, et al.
Published: (2024)
by: Liu, Jinmei, et al.
Published: (2024)
DROJ: A Prompt-Driven Attack against Large Language Models
by: Hu, Leyang, et al.
Published: (2024)
by: Hu, Leyang, et al.
Published: (2024)
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
by: Li, Yafu, et al.
Published: (2026)
by: Li, Yafu, et al.
Published: (2026)
Next Token Perception Score: Analytical Assessment of your LLM Perception Skills
by: Cheng, Yu-Ang, et al.
Published: (2025)
by: Cheng, Yu-Ang, et al.
Published: (2025)
Rethinking Entropy Regularization in Large Reasoning Models
by: Jiang, Yuxian, et al.
Published: (2025)
by: Jiang, Yuxian, et al.
Published: (2025)
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
by: Huang, Siyuan, et al.
Published: (2026)
by: Huang, Siyuan, et al.
Published: (2026)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
by: Zhang, Jihai, et al.
Published: (2024)
by: Zhang, Jihai, et al.
Published: (2024)
Neuro-Symbolic Integration Brings Causal and Reliable Reasoning Proofs
by: Yang, Sen, et al.
Published: (2023)
by: Yang, Sen, et al.
Published: (2023)
Multi-LLM Collaborative Search for Complex Problem Solving
by: Yang, Sen, et al.
Published: (2025)
by: Yang, Sen, et al.
Published: (2025)
Incentivized Exploration of Non-Stationary Stochastic Bandits
by: Chakraborty, Sourav, et al.
Published: (2024)
by: Chakraborty, Sourav, et al.
Published: (2024)
Learning Compact Representations of LLM Abilities via Item Response Theory
by: Chen, Jianhao, et al.
Published: (2025)
by: Chen, Jianhao, et al.
Published: (2025)
Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted Concepts
by: Li, Leyang, et al.
Published: (2025)
by: Li, Leyang, et al.
Published: (2025)
From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning
by: Li, Yafu, et al.
Published: (2025)
by: Li, Yafu, et al.
Published: (2025)
$π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
by: Zhang, Haoran, et al.
Published: (2026)
by: Zhang, Haoran, et al.
Published: (2026)
GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
by: Kang, Weitai, et al.
Published: (2025)
by: Kang, Weitai, et al.
Published: (2025)
Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scaling
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
by: Chen, Guanxu, et al.
Published: (2025)
by: Chen, Guanxu, et al.
Published: (2025)
Incentivizing In-depth Reasoning over Long Contexts with Process Advantage Shaping
by: Peng, Miao, et al.
Published: (2026)
by: Peng, Miao, et al.
Published: (2026)
MCRanker: Generating Diverse Criteria On-the-Fly to Improve Point-wise LLM Rankers
by: Guo, Fang, et al.
Published: (2024)
by: Guo, Fang, et al.
Published: (2024)
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
by: Qiao, Runqi, et al.
Published: (2025)
by: Qiao, Runqi, et al.
Published: (2025)
RAP: Runtime Adaptive Pruning for LLM Inference
by: Liu, Huanrong, et al.
Published: (2025)
by: Liu, Huanrong, et al.
Published: (2025)
Less is More: Resource-Efficient Low-Rank Adaptation
by: Tian, Chunlin, et al.
Published: (2025)
by: Tian, Chunlin, et al.
Published: (2025)
Similar Items
-
Learning to Reason under Off-Policy Guidance
by: Yan, Jianhao, et al.
Published: (2025) -
Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning
by: Hu, Zican, et al.
Published: (2025) -
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
by: Fu, Tingchen, et al.
Published: (2025) -
Text-to-Decision Agent: Offline Meta-Reinforcement Learning from Natural Language Supervision
by: Zhang, Shilin, et al.
Published: (2025) -
SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
by: Shen, Chuming, et al.
Published: (2025)