SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Dengjia, Liu, Xiaoou, Cheng, Lu, Wang, Yaqing, Murray, Kenton, Wei, Hua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
by: Liu, Xiaoou, et al.
Published: (2026)
by: Liu, Xiaoou, et al.
Published: (2026)
DOTResize: Reducing LLM Width via Discrete Optimal Transport-based Neuron Merging
by: Verma, Neha, et al.
Published: (2025)
by: Verma, Neha, et al.
Published: (2025)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
by: Karlekar, Sweta, et al.
Published: (2026)
by: Karlekar, Sweta, et al.
Published: (2026)
Merging Feed-Forward Sublayers for Compressed Transformers
by: Verma, Neha, et al.
Published: (2025)
by: Verma, Neha, et al.
Published: (2025)
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
SAUP: Situation Awareness Uncertainty Propagation on LLM Agent
by: Zhao, Qiwei, et al.
Published: (2024)
by: Zhao, Qiwei, et al.
Published: (2024)
Reward-Based Online LLM Routing via NeuralUCB
by: Tsai, Ming-Hua, et al.
Published: (2026)
by: Tsai, Ming-Hua, et al.
Published: (2026)
RewardHarness: Self-Evolving Agentic Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
by: Zhang, Haozhen, et al.
Published: (2026)
by: Zhang, Haozhen, et al.
Published: (2026)
AgentEvolver: Towards Efficient Self-Evolving Agent System
by: Zhai, Yunpeng, et al.
Published: (2025)
by: Zhai, Yunpeng, et al.
Published: (2025)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Symbolic Learning Enables Self-Evolving Agents
by: Zhou, Wangchunshu, et al.
Published: (2024)
by: Zhou, Wangchunshu, et al.
Published: (2024)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
by: Chen, Jiaqi, et al.
Published: (2025)
by: Chen, Jiaqi, et al.
Published: (2025)
Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
by: Wang, Jiawei, et al.
Published: (2025)
by: Wang, Jiawei, et al.
Published: (2025)
MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
by: Wen, Bingbing, et al.
Published: (2026)
by: Wen, Bingbing, et al.
Published: (2026)
DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization
by: She, Shuaijie, et al.
Published: (2025)
by: She, Shuaijie, et al.
Published: (2025)
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
Unified Multimodal Uncertain Inference
by: Zhang, Dengjia, et al.
Published: (2026)
by: Zhang, Dengjia, et al.
Published: (2026)
CREAM: Consistency Regularized Self-Rewarding Language Models
by: Wang, Zhaoyang, et al.
Published: (2024)
by: Wang, Zhaoyang, et al.
Published: (2024)
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
by: Wang, Xiaobo, et al.
Published: (2025)
by: Wang, Xiaobo, et al.
Published: (2025)
Evolving LLMs' Self-Refinement Capability via Synergistic Training-Inference Optimization
by: Zeng, Yongcheng, et al.
Published: (2025)
by: Zeng, Yongcheng, et al.
Published: (2025)
Upsample or Upweight? Balanced Training on Heavily Imbalanced Datasets
by: Li, Tianjian, et al.
Published: (2024)
by: Li, Tianjian, et al.
Published: (2024)
Every Response Counts: Quantifying Uncertainty of LLM-based Multi-Agent Systems through Tensor Decomposition
by: Chen, Tiejin, et al.
Published: (2026)
by: Chen, Tiejin, et al.
Published: (2026)
Diving into Self-Evolving Training for Multimodal Reasoning
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
R-Zero: Self-Evolving Reasoning LLM from Zero Data
by: Huang, Chengsong, et al.
Published: (2025)
by: Huang, Chengsong, et al.
Published: (2025)
Universe Routing: Why Self-Evolving Agents Need Epistemic Control
by: Wang, Zhaohui Geoffrey
Published: (2026)
by: Wang, Zhaohui Geoffrey
Published: (2026)
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
by: Xia, Chunqiu Steven, et al.
Published: (2025)
by: Xia, Chunqiu Steven, et al.
Published: (2025)
Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory
by: Wang, Zhenting, et al.
Published: (2026)
by: Wang, Zhenting, et al.
Published: (2026)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
by: Xia, Fanzeng, et al.
Published: (2024)
by: Xia, Fanzeng, et al.
Published: (2024)
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
by: Li, Sijia, et al.
Published: (2026)
by: Li, Sijia, et al.
Published: (2026)
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Scaling Long-Horizon LLM Agent via Context-Folding
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
AgentRM: Enhancing Agent Generalization with Reward Modeling
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective
by: Wang, Yuyao, et al.
Published: (2026)
by: Wang, Yuyao, et al.
Published: (2026)
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
by: Shao, Shuai, et al.
Published: (2025)
by: Shao, Shuai, et al.
Published: (2025)
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
by: Chen, Tiejin, et al.
Published: (2026)
by: Chen, Tiejin, et al.
Published: (2026)
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
by: Sun, Shengyang, et al.
Published: (2025)
by: Sun, Shengyang, et al.
Published: (2025)
Similar Items
-
Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
by: Liu, Xiaoou, et al.
Published: (2026) -
DOTResize: Reducing LLM Width via Discrete Optimal Transport-based Neuron Merging
by: Verma, Neha, et al.
Published: (2025) -
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
by: Karlekar, Sweta, et al.
Published: (2026) -
Merging Feed-Forward Sublayers for Compressed Transformers
by: Verma, Neha, et al.
Published: (2025) -
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
by: Da, Longchao, et al.
Published: (2025)