Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample Efficiency
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zhihan, Hu, Hao, Zhang, Shenao, Guo, Hongyi, Ke, Shuqi, Liu, Boyi, Wang, Zhaoran |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
How Can LLM Guide RL? A Value-Based Approach
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
by: Guo, Hongyi, et al.
Published: (2024)
by: Guo, Hongyi, et al.
Published: (2024)
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning
by: Zhong, Han, et al.
Published: (2025)
by: Zhong, Han, et al.
Published: (2025)
Hindsight Planner: A Closed-Loop Few-Shot Planner for Embodied Instruction Following
by: Yang, Yuxiao, et al.
Published: (2024)
by: Yang, Yuxiao, et al.
Published: (2024)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
Embed to Control Partially Observed Systems: Representation Learning with Provable Sample Efficiency
by: Wang, Lingxiao, et al.
Published: (2022)
by: Wang, Lingxiao, et al.
Published: (2022)
Human-Instruction-Free LLM Self-Alignment with Limited Samples
by: Guo, Hongyi, et al.
Published: (2024)
by: Guo, Hongyi, et al.
Published: (2024)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
by: Liu, Zhihan, et al.
Published: (2026)
by: Liu, Zhihan, et al.
Published: (2026)
PRACT: Optimizing Principled Reasoning and Acting of LLM Agent
by: Liu, Zhiwei, et al.
Published: (2024)
by: Liu, Zhiwei, et al.
Published: (2024)
Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
by: Zhang, Shenao, et al.
Published: (2024)
by: Zhang, Shenao, et al.
Published: (2024)
ClimAgent: LLM as Agents for Autonomous Open-ended Climate Science Analysis
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving
by: Sun, Zhengqi, et al.
Published: (2026)
by: Sun, Zhengqi, et al.
Published: (2026)
Toward Optimal LLM Alignments Using Two-Player Games
by: Zheng, Rui, et al.
Published: (2024)
by: Zheng, Rui, et al.
Published: (2024)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Distilling the Thought, Watermarking the Answer: A Principle Semantic Guided Watermark for Large Reasoning Models
by: Liu, Shuliang, et al.
Published: (2026)
by: Liu, Shuliang, et al.
Published: (2026)
Pre-Act: Multi-Step Planning and Reasoning Improves Acting in LLM Agents
by: Rawat, Mrinal, et al.
Published: (2025)
by: Rawat, Mrinal, et al.
Published: (2025)
ODAR: Principled Adaptive Routing for LLM Reasoning via Active Inference
by: Ma, Siyuan, et al.
Published: (2026)
by: Ma, Siyuan, et al.
Published: (2026)
TurboAgent: An LLM-Driven Autonomous Multi-Agent Framework for Turbomachinery Aerodynamic Design
by: Du, Juan, et al.
Published: (2026)
by: Du, Juan, et al.
Published: (2026)
ActMem: Bridging the Gap Between Memory Retrieval and Reasoning in LLM Agents
by: Zhang, Xiaohui, et al.
Published: (2026)
by: Zhang, Xiaohui, et al.
Published: (2026)
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
by: Xu, Ruijie, et al.
Published: (2024)
by: Xu, Ruijie, et al.
Published: (2024)
R&D-Agent: An LLM-Agent Framework Towards Autonomous Data Science
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
Verifiability-First Agents: Provable Observability and Lightweight Audit Agents for Controlling Autonomous LLM Systems
by: Gupta, Abhivansh
Published: (2025)
by: Gupta, Abhivansh
Published: (2025)
A Concurrent Modular Agent: Framework for Autonomous LLM Agents
by: Maruyama, Norihiro, et al.
Published: (2025)
by: Maruyama, Norihiro, et al.
Published: (2025)
The Reasons that Agents Act: Intention and Instrumental Goals
by: Ward, Francis Rhys, et al.
Published: (2024)
by: Ward, Francis Rhys, et al.
Published: (2024)
Time-Scaling Is What Agents Need Now
by: Liu, Zhi, et al.
Published: (2026)
by: Liu, Zhi, et al.
Published: (2026)
ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
by: Tawosi, Vali, et al.
Published: (2025)
by: Tawosi, Vali, et al.
Published: (2025)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
by: Wang, Huaijie, et al.
Published: (2024)
by: Wang, Huaijie, et al.
Published: (2024)
Scalable and Accurate Graph Reasoning with LLM-based Multi-Agents
by: Hu, Yuwei, et al.
Published: (2024)
by: Hu, Yuwei, et al.
Published: (2024)
AgentFactory: A Self-Evolving Framework Through Executable Subagent Accumulation and Reuse
by: Zhang, Zhang, et al.
Published: (2026)
by: Zhang, Zhang, et al.
Published: (2026)
MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning
by: Liu, Yuxin, et al.
Published: (2026)
by: Liu, Yuxin, et al.
Published: (2026)
Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
by: Wei, Xiaolong, et al.
Published: (2025)
by: Wei, Xiaolong, et al.
Published: (2025)
AgriWorld:A World Tools Protocol Framework for Verifiable Agricultural Reasoning with Code-Executing LLM Agents
by: Zhang, Zhixing, et al.
Published: (2026)
by: Zhang, Zhixing, et al.
Published: (2026)
When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems
by: Huang, Donghao, et al.
Published: (2026)
by: Huang, Donghao, et al.
Published: (2026)
Ten Principles of AI Agent Economics
by: Yang, Ke, et al.
Published: (2025)
by: Yang, Ke, et al.
Published: (2025)
SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
by: Lin, Jiaye, et al.
Published: (2025)
by: Lin, Jiaye, et al.
Published: (2025)
TRACE: A Multi-Agent System for Autonomous Physical Reasoning for Seismology
by: Liu, Feng, et al.
Published: (2026)
by: Liu, Feng, et al.
Published: (2026)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
by: Fang, Yi, et al.
Published: (2026)
by: Fang, Yi, et al.
Published: (2026)
Similar Items
-
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
by: Liu, Zhihan, et al.
Published: (2024) -
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs
by: Liu, Zhihan, et al.
Published: (2024) -
How Can LLM Guide RL? A Value-Based Approach
by: Zhang, Shenao, et al.
Published: (2024) -
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
by: Guo, Hongyi, et al.
Published: (2024) -
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning
by: Zhong, Han, et al.
Published: (2025)