Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Zishang, Han, Jinyi, Li, Tingyun, Wang, Xinyi, Jiang, Sihang, Liang, Jiaqing, Dai, Zhaoqian, Ma, Shuguang, Yu, Fei, Xiao, Yanghua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HINT: Helping Ineffective Rollouts Navigate Towards Effectiveness
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
A Stitch in Time Saves Nine: Proactive Self-Refinement for Language Models
by: Han, Jinyi, et al.
Published: (2025)
by: Han, Jinyi, et al.
Published: (2025)
Structured Reasoning for Large Language Models
by: Han, Jinyi, et al.
Published: (2026)
by: Han, Jinyi, et al.
Published: (2026)
Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking
by: Han, Jinyi, et al.
Published: (2025)
by: Han, Jinyi, et al.
Published: (2025)
GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)
by: Liang, Jiaqing, et al.
Published: (2026)
by: Liang, Jiaqing, et al.
Published: (2026)
LogReasoner: Empowering LLMs with Expert-like Coarse-to-Fine Reasoning for Automated Log Analysis
by: Ma, Lipeng, et al.
Published: (2025)
by: Ma, Lipeng, et al.
Published: (2025)
Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension Perception
by: Huang, Yuncheng, et al.
Published: (2023)
by: Huang, Yuncheng, et al.
Published: (2023)
Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation
by: Han, Jinyi, et al.
Published: (2025)
by: Han, Jinyi, et al.
Published: (2025)
SED: Self-Evaluation Decoding Enhances Large Language Models for Better Generation
by: Luo, Ziqin, et al.
Published: (2024)
by: Luo, Ziqin, et al.
Published: (2024)
CultureScope: A Dimensional Lens for Probing Cultural Understanding in LLMs
by: Zhang, Jinghao, et al.
Published: (2025)
by: Zhang, Jinghao, et al.
Published: (2025)
CDS: Knowledge Component-Driven Data Synthesis Guided by Cognitive Diagnosis Theory
by: Zhao, Haokun, et al.
Published: (2025)
by: Zhao, Haokun, et al.
Published: (2025)
Enhancing Confidence Expression in Large Language Models Through Learning from Past Experience
by: Han, Haixia, et al.
Published: (2024)
by: Han, Haixia, et al.
Published: (2024)
LUK: Empowering Log Understanding with Expert Knowledge from Large Language Models
by: Ma, Lipeng, et al.
Published: (2024)
by: Ma, Lipeng, et al.
Published: (2024)
Toward Automated Robustness Evaluation of Mathematical Reasoning
by: Hou, Yutao, et al.
Published: (2025)
by: Hou, Yutao, et al.
Published: (2025)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
by: He, Qianxi, et al.
Published: (2025)
by: He, Qianxi, et al.
Published: (2025)
Adaptive Ordered Information Extraction with Deep Reinforcement Learning
by: Huang, Wenhao, et al.
Published: (2023)
by: Huang, Wenhao, et al.
Published: (2023)
Do Large Language Models Truly Understand Cross-cultural Differences?
by: Guo, Shiwei, et al.
Published: (2025)
by: Guo, Shiwei, et al.
Published: (2025)
Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation
by: Xia, Sirui, et al.
Published: (2024)
by: Xia, Sirui, et al.
Published: (2024)
Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
by: Ren, Qingyu, et al.
Published: (2025)
by: Ren, Qingyu, et al.
Published: (2025)
Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance
by: Fang, Yue, et al.
Published: (2025)
by: Fang, Yue, et al.
Published: (2025)
Small Language Model Can Self-correct
by: Han, Haixia, et al.
Published: (2024)
by: Han, Haixia, et al.
Published: (2024)
SEIF: Self-Evolving Reinforcement Learning for Instruction Following
by: Ren, Qingyu, et al.
Published: (2026)
by: Ren, Qingyu, et al.
Published: (2026)
EDGE: Enhanced Grounded GUI Understanding with Enriched Multi-Granularity Synthetic Data
by: Chen, Xuetian, et al.
Published: (2024)
by: Chen, Xuetian, et al.
Published: (2024)
Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following
by: Ren, Qingyu, et al.
Published: (2025)
by: Ren, Qingyu, et al.
Published: (2025)
VCEval: Rethinking What is a Good Educational Video and How to Automatically Evaluate It
by: Zhu, Xiaoxuan, et al.
Published: (2024)
by: Zhu, Xiaoxuan, et al.
Published: (2024)
SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
by: Jiang, Sihang, et al.
Published: (2026)
by: Jiang, Sihang, et al.
Published: (2026)
Reason from Fallacy: Enhancing Large Language Models' Logical Reasoning through Logical Fallacy Understanding
by: Li, Yanda, et al.
Published: (2024)
by: Li, Yanda, et al.
Published: (2024)
CEM: A Data-Efficient Method for Large Language Models to Continue Evolving From Mistakes
by: Zhao, Haokun, et al.
Published: (2024)
by: Zhao, Haokun, et al.
Published: (2024)
AdaptiveLog: An Adaptive Log Analysis Framework with the Collaboration of Large and Small Language Model
by: Ma, Lipeng, et al.
Published: (2025)
by: Ma, Lipeng, et al.
Published: (2025)
Laying the Foundation First? Investigating the Generalization from Atomic Skills to Complex Reasoning Tasks
by: Huang, Yuncheng, et al.
Published: (2024)
by: Huang, Yuncheng, et al.
Published: (2024)
From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models
by: He, Qianyu, et al.
Published: (2024)
by: He, Qianyu, et al.
Published: (2024)
Is There a One-Model-Fits-All Approach to Information Extraction? Revisiting Task Definition Biases
by: Huang, Wenhao, et al.
Published: (2024)
by: Huang, Wenhao, et al.
Published: (2024)
Improving Recall of Large Language Models: A Model Collaboration Approach for Relational Triple Extraction
by: Ding, Zepeng, et al.
Published: (2024)
by: Ding, Zepeng, et al.
Published: (2024)
MultiLingPoT: Enhancing Mathematical Reasoning with Multilingual Program Fine-tuning
by: Li, Nianqi, et al.
Published: (2024)
by: Li, Nianqi, et al.
Published: (2024)
Order Matters: Investigate the Position Bias in Multi-constraint Instruction Following
by: Zeng, Jie, et al.
Published: (2025)
by: Zeng, Jie, et al.
Published: (2025)
Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models
by: Ren, Qingyu, et al.
Published: (2025)
by: Ren, Qingyu, et al.
Published: (2025)
Why Did Apple Fall: Evaluating Curiosity in Large Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Mixture of Autoencoder Experts Guidance using Unlabeled and Incomplete Data for Exploration in Reinforcement Learning
by: Malomgré, Elias, et al.
Published: (2025)
by: Malomgré, Elias, et al.
Published: (2025)
Local Hamiltonian Problem with succinct ground state is MA-Complete
by: Jiang, Jiaqing
Published: (2023)
by: Jiang, Jiaqing
Published: (2023)
Can Pre-trained Language Models Understand Chinese Humor?
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Similar Items
-
HINT: Helping Ineffective Rollouts Navigate Towards Effectiveness
by: Wang, Xinyi, et al.
Published: (2025) -
A Stitch in Time Saves Nine: Proactive Self-Refinement for Language Models
by: Han, Jinyi, et al.
Published: (2025) -
Structured Reasoning for Large Language Models
by: Han, Jinyi, et al.
Published: (2026) -
Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking
by: Han, Jinyi, et al.
Published: (2025) -
GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)
by: Liang, Jiaqing, et al.
Published: (2026)