Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tian, Ye, Peng, Baolin, Song, Linfeng, Jin, Lifeng, Yu, Dian, Mi, Haitao, Yu, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
LiteSearch: Efficacious Tree Search for LLM
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
von: Yu, Dian, et al.
Veröffentlicht: (2024)
von: Yu, Dian, et al.
Veröffentlicht: (2024)
Collaborative decoding of critical tokens for boosting factuality of large language models
von: Jin, Lifeng, et al.
Veröffentlicht: (2024)
von: Jin, Lifeng, et al.
Veröffentlicht: (2024)
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
von: Yue, Murong, et al.
Veröffentlicht: (2024)
von: Yue, Murong, et al.
Veröffentlicht: (2024)
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
von: Yu, Dian, et al.
Veröffentlicht: (2025)
von: Yu, Dian, et al.
Veröffentlicht: (2025)
Self-Consistency Boosts Calibration for Math Reasoning
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
Fine-Grained Self-Endorsement Improves Factuality and Reasoning
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
Entropy Guided Extrapolative Decoding to Improve Factuality in Large Language Models
von: Das, Souvik, et al.
Veröffentlicht: (2024)
von: Das, Souvik, et al.
Veröffentlicht: (2024)
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
von: Zhou, Yujun, et al.
Veröffentlicht: (2025)
Scaling Synthetic Data Creation with 1,000,000,000 Personas
von: Ge, Tao, et al.
Veröffentlicht: (2024)
von: Ge, Tao, et al.
Veröffentlicht: (2024)
Guided Self-Evolving LLMs with Minimal Human Supervision
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
von: Yu, Wenhao, et al.
Veröffentlicht: (2025)
One Token to Fool LLM-as-a-Judge
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
Inconsistent dialogue responses and how to recover from them
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
von: Zhang, Mian, et al.
Veröffentlicht: (2024)
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
Teaching LLMs to Refine with Tools
von: Yu, Dian, et al.
Veröffentlicht: (2024)
von: Yu, Dian, et al.
Veröffentlicht: (2024)
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
von: Wang, Ante, et al.
Veröffentlicht: (2025)
von: Wang, Ante, et al.
Veröffentlicht: (2025)
A Knowledge Plug-and-Play Test Bed for Open-domain Dialogue Generation
von: Li, Xiangci, et al.
Veröffentlicht: (2024)
von: Li, Xiangci, et al.
Veröffentlicht: (2024)
Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition
von: Ye, Guanghao, et al.
Veröffentlicht: (2025)
von: Ye, Guanghao, et al.
Veröffentlicht: (2025)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
Evolving LLMs' Self-Refinement Capability via Synergistic Training-Inference Optimization
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2025)
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2025)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
von: Lee, Yu-Ting, et al.
Veröffentlicht: (2025)
von: Lee, Yu-Ting, et al.
Veröffentlicht: (2025)
R-Zero: Self-Evolving Reasoning LLM from Zero Data
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
von: Huang, Chengsong, et al.
Veröffentlicht: (2025)
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
von: Shi, Yucheng, et al.
Veröffentlicht: (2025)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Is Parameter Collision Hindering Continual Learning in LLMs?
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
Self-Imagine: Effective Unimodal Reasoning with Multimodal Models using Self-Imagination
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
von: Yang, An, et al.
Veröffentlicht: (2024)
von: Yang, An, et al.
Veröffentlicht: (2024)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
von: Li, Moxin, et al.
Veröffentlicht: (2025)
von: Li, Moxin, et al.
Veröffentlicht: (2025)
Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
von: Wang, Xiyao, et al.
Veröffentlicht: (2024) -
LiteSearch: Efficacious Tree Search for LLM
von: Wang, Ante, et al.
Veröffentlicht: (2024) -
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024) -
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
von: Yu, Dian, et al.
Veröffentlicht: (2024) -
Collaborative decoding of critical tokens for boosting factuality of large language models
von: Jin, Lifeng, et al.
Veröffentlicht: (2024)