EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Wujiang, Zhao, Wentian, Wang, Zhenting, Li, Yu-Jhe, Jin, Can, Jin, Mingyu, Mei, Kai, Wan, Kun, Metaxas, Dimitris N. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
by: Wang, Zhenting, et al.
Published: (2025)
by: Wang, Zhenting, et al.
Published: (2025)
Pooling and Semantic Shift: The Fundamental Challenges in Long Text Embedding and Retrieval
by: Gao, Hang, et al.
Published: (2026)
by: Gao, Hang, et al.
Published: (2026)
LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
APEER: Automatic Prompt Engineering Enhances Large Language Model Reranking
by: Jin, Can, et al.
Published: (2024)
by: Jin, Can, et al.
Published: (2024)
AEL: Agent Evolving Learning for Open-Ended Environments
by: Xu, Wujiang, et al.
Published: (2026)
by: Xu, Wujiang, et al.
Published: (2026)
M^3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Token-Controlled Re-ranking for Sequential Recommendation via LLMs
by: Dai, Wenxi, et al.
Published: (2025)
by: Dai, Wenxi, et al.
Published: (2025)
LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
AIOS: LLM Agent Operating System
by: Mei, Kai, et al.
Published: (2024)
by: Mei, Kai, et al.
Published: (2024)
Individual Turing Test: A Case Study of LLM-based Simulation Using Longitudinal Personal Data
by: Guo, Minghao, et al.
Published: (2026)
by: Guo, Minghao, et al.
Published: (2026)
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight
by: Jin, Can, et al.
Published: (2026)
by: Jin, Can, et al.
Published: (2026)
EPO: Hierarchical LLM Agents with Environment Preference Optimization
by: Zhao, Qi, et al.
Published: (2024)
by: Zhao, Qi, et al.
Published: (2024)
AgentSelect: Benchmark for Narrative Query-to-Agent Recommendation
by: Shi, Yunxiao, et al.
Published: (2026)
by: Shi, Yunxiao, et al.
Published: (2026)
Trust or Abstain? A Self-Aware RAG Approach
by: Zhu, Xi, et al.
Published: (2026)
by: Zhu, Xi, et al.
Published: (2026)
MHB: Multimodal Handshape-aware Boundary Detection for Continuous Sign Language Recognition
by: Zhao, Mingyu, et al.
Published: (2025)
by: Zhao, Mingyu, et al.
Published: (2025)
Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs
by: Jin, Mingyu, et al.
Published: (2026)
by: Jin, Mingyu, et al.
Published: (2026)
SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models
by: Han, Jiaojiao, et al.
Published: (2025)
by: Han, Jiaojiao, et al.
Published: (2025)
A-MEM: Agentic Memory for LLM Agents
by: Xu, Wujiang, et al.
Published: (2025)
by: Xu, Wujiang, et al.
Published: (2025)
iAgent: LLM Agent as a Shield between User and Recommender Systems
by: Xu, Wujiang, et al.
Published: (2025)
by: Xu, Wujiang, et al.
Published: (2025)
DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models
by: Wang, Zhenting, et al.
Published: (2023)
by: Wang, Zhenting, et al.
Published: (2023)
Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
by: Deng, Shijian, et al.
Published: (2024)
by: Deng, Shijian, et al.
Published: (2024)
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
by: Zeng, Yifan, et al.
Published: (2026)
by: Zeng, Yifan, et al.
Published: (2026)
Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety
by: Jin, Can, et al.
Published: (2026)
by: Jin, Can, et al.
Published: (2026)
Evidence Over Plans: Online Trajectory Verification for Skill Distillation
by: Zhou, Yang, et al.
Published: (2026)
by: Zhou, Yang, et al.
Published: (2026)
Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference
by: Yuan, Jiayi, et al.
Published: (2025)
by: Yuan, Jiayi, et al.
Published: (2025)
EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning
by: Liu, Xiaoqian, et al.
Published: (2025)
by: Liu, Xiaoqian, et al.
Published: (2025)
OmniRouter: Budget and Performance Controllable Multi-LLM Routing
by: Mei, Kai, et al.
Published: (2025)
by: Mei, Kai, et al.
Published: (2025)
Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding
by: Jin, Mingyu, et al.
Published: (2025)
by: Jin, Mingyu, et al.
Published: (2025)
How to Trace Latent Generative Model Generated Images without Artificial Watermark?
by: Wang, Zhenting, et al.
Published: (2024)
by: Wang, Zhenting, et al.
Published: (2024)
Learning from Teaching Regularization: Generalizable Correlations Should be Easy to Imitate
by: Jin, Can, et al.
Published: (2024)
by: Jin, Can, et al.
Published: (2024)
Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Seeing Farther and Smarter: Value-Guided Multi-Path Reflection for VLM Policy Optimization
by: Yang, Yanting, et al.
Published: (2026)
by: Yang, Yanting, et al.
Published: (2026)
Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
Improving Visual Reasoning with Iterative Evidence Refinement
by: Shi, Zeru, et al.
Published: (2026)
by: Shi, Zeru, et al.
Published: (2026)
Beyond Explicit Edges: Robust Reasoning over Noisy and Sparse Knowledge Graphs
by: Gao, Hang, et al.
Published: (2026)
by: Gao, Hang, et al.
Published: (2026)
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation
by: Zhou, Yang, et al.
Published: (2026)
by: Zhou, Yang, et al.
Published: (2026)
From Commands to Prompts: LLM-based Semantic File System for AIOS
by: Shi, Zeru, et al.
Published: (2024)
by: Shi, Zeru, et al.
Published: (2024)
MoralBench: Moral Evaluation of LLMs
by: Ji, Jianchao, et al.
Published: (2024)
by: Ji, Jianchao, et al.
Published: (2024)
Agents in the Wild: Safety, Society, and the Illusion of Sociality on Moltbook
by: Zhang, Yunbei, et al.
Published: (2026)
by: Zhang, Yunbei, et al.
Published: (2026)
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
by: Dao, Quan, et al.
Published: (2026)
by: Dao, Quan, et al.
Published: (2026)
Similar Items
-
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
by: Wang, Zhenting, et al.
Published: (2025) -
Pooling and Semantic Shift: The Fundamental Challenges in Long Text Embedding and Retrieval
by: Gao, Hang, et al.
Published: (2026) -
LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
by: Zhou, Yang, et al.
Published: (2025) -
APEER: Automatic Prompt Engineering Enhances Large Language Model Reranking
by: Jin, Can, et al.
Published: (2024) -
AEL: Agent Evolving Learning for Open-Ended Environments
by: Xu, Wujiang, et al.
Published: (2026)