Reward Bound for Behavioral Guarantee of Model-based Planning Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | An, Zhiyu, Ding, Xianzhong, Du, Wan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Go Beyond Black-box Policies: Rethinking the Design of Learning Agent for Interpretable and Verifiable HVAC Control
di: An, Zhiyu, et al.
Pubblicazione: (2024)
di: An, Zhiyu, et al.
Pubblicazione: (2024)
DIML: Differentiable Inverse Mechanism Learning from Behaviors of Multi-Agent Learning Trajectories
di: An, Zhiyu, et al.
Pubblicazione: (2026)
di: An, Zhiyu, et al.
Pubblicazione: (2026)
MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning
di: An, Zhiyu, et al.
Pubblicazione: (2025)
di: An, Zhiyu, et al.
Pubblicazione: (2025)
Representational Homomorphism Predicts and Improves Compositional Generalization In Transformer Language Model
di: An, Zhiyu, et al.
Pubblicazione: (2026)
di: An, Zhiyu, et al.
Pubblicazione: (2026)
Golden-Retriever: High-Fidelity Agentic Retrieval Augmented Generation for Industrial Knowledge Base
di: An, Zhiyu, et al.
Pubblicazione: (2024)
di: An, Zhiyu, et al.
Pubblicazione: (2024)
Open Problems in Differentiable Social Choice: Learning Mechanisms, Decisions, and Alignment
di: An, Zhiyu, et al.
Pubblicazione: (2026)
di: An, Zhiyu, et al.
Pubblicazione: (2026)
Scaling Autonomous Agents via Automatic Reward Modeling And Planning
di: Chen, Zhenfang, et al.
Pubblicazione: (2025)
di: Chen, Zhenfang, et al.
Pubblicazione: (2025)
Debiasing Reward Models by Representation Learning with Guarantees
di: Ng, Ignavier, et al.
Pubblicazione: (2025)
di: Ng, Ignavier, et al.
Pubblicazione: (2025)
A Safe and Data-efficient Model-based Reinforcement Learning System for HVAC Control
di: Ding, Xianzhong, et al.
Pubblicazione: (2024)
di: Ding, Xianzhong, et al.
Pubblicazione: (2024)
Beyond Noisy-TVs: Noise-Robust Exploration Via Learning Progress Monitoring
di: Hou, Zhibo, et al.
Pubblicazione: (2025)
di: Hou, Zhibo, et al.
Pubblicazione: (2025)
Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling
di: Wang, Jiaxuan, et al.
Pubblicazione: (2026)
di: Wang, Jiaxuan, et al.
Pubblicazione: (2026)
Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models
di: Fu, Shi, et al.
Pubblicazione: (2026)
di: Fu, Shi, et al.
Pubblicazione: (2026)
Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences
di: An, Zhiyu, et al.
Pubblicazione: (2026)
di: An, Zhiyu, et al.
Pubblicazione: (2026)
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
di: Men, Tianyi, et al.
Pubblicazione: (2025)
di: Men, Tianyi, et al.
Pubblicazione: (2025)
HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
di: Wu, Peilin, et al.
Pubblicazione: (2025)
di: Wu, Peilin, et al.
Pubblicazione: (2025)
Confidence as a Reward: Transforming LLMs into Reward Models
di: Du, He, et al.
Pubblicazione: (2025)
di: Du, He, et al.
Pubblicazione: (2025)
TPTU: Large Language Model-based AI Agents for Task Planning and Tool Usage
di: Ruan, Jingqing, et al.
Pubblicazione: (2023)
di: Ruan, Jingqing, et al.
Pubblicazione: (2023)
Rectifying Shortcut Behaviors in Preference-based Reward Learning
di: Ye, Wenqian, et al.
Pubblicazione: (2025)
di: Ye, Wenqian, et al.
Pubblicazione: (2025)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
di: Ding, Liang
Pubblicazione: (2026)
di: Ding, Liang
Pubblicazione: (2026)
Exploring Reasoning Reward Model for Agents
di: Fan, Kaixuan, et al.
Pubblicazione: (2026)
di: Fan, Kaixuan, et al.
Pubblicazione: (2026)
Adapting the Behavior of Reinforcement Learning Agents to Changing Action Spaces and Reward Functions
di: de la Rosa, Raul, et al.
Pubblicazione: (2026)
di: de la Rosa, Raul, et al.
Pubblicazione: (2026)
TRIP-PAL: Travel Planning with Guarantees by Combining Large Language Models and Automated Planners
di: de la Rosa, Tomas, et al.
Pubblicazione: (2024)
di: de la Rosa, Tomas, et al.
Pubblicazione: (2024)
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
Simplifying Complex Observation Models in Continuous POMDP Planning with Probabilistic Guarantees and Practice
di: Lev-Yehudi, Idan, et al.
Pubblicazione: (2023)
di: Lev-Yehudi, Idan, et al.
Pubblicazione: (2023)
Decomposability-Guaranteed Cooperative Coevolution for Large-Scale Itinerary Planning
di: Zhang, Ziyu, et al.
Pubblicazione: (2025)
di: Zhang, Ziyu, et al.
Pubblicazione: (2025)
Online POMDP Planning with Anytime Deterministic Optimality Guarantees
di: Barenboim, Moran, et al.
Pubblicazione: (2023)
di: Barenboim, Moran, et al.
Pubblicazione: (2023)
Depth-Bounded Epistemic Planning
di: Bolander, Thomas, et al.
Pubblicazione: (2024)
di: Bolander, Thomas, et al.
Pubblicazione: (2024)
Agentified Assessment of Logical Reasoning Agents
di: Ni, Zhiyu, et al.
Pubblicazione: (2026)
di: Ni, Zhiyu, et al.
Pubblicazione: (2026)
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
di: Lu, Qingyu, et al.
Pubblicazione: (2025)
di: Lu, Qingyu, et al.
Pubblicazione: (2025)
Path Planning based on 2D Object Bounding-box
di: Huang, Yanliang, et al.
Pubblicazione: (2024)
di: Huang, Yanliang, et al.
Pubblicazione: (2024)
Principal-Agent Reward Shaping in MDPs
di: Ben-Porat, Omer, et al.
Pubblicazione: (2023)
di: Ben-Porat, Omer, et al.
Pubblicazione: (2023)
Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing
di: Jiao, Fangkai, et al.
Pubblicazione: (2024)
di: Jiao, Fangkai, et al.
Pubblicazione: (2024)
MASP: Scalable GNN-based Planning for Multi-Agent Navigation
di: Yang, Xinyi, et al.
Pubblicazione: (2023)
di: Yang, Xinyi, et al.
Pubblicazione: (2023)
Tiered Reward: Designing Rewards for Specification and Fast Learning of Desired Behavior
di: Zhou, Zhiyuan, et al.
Pubblicazione: (2022)
di: Zhou, Zhiyuan, et al.
Pubblicazione: (2022)
User Behavior Simulation with Large Language Model based Agents
di: Wang, Lei, et al.
Pubblicazione: (2023)
di: Wang, Lei, et al.
Pubblicazione: (2023)
Explaining an Agent's Future Beliefs through Temporally Decomposing Future Reward Estimators
di: Towers, Mark, et al.
Pubblicazione: (2024)
di: Towers, Mark, et al.
Pubblicazione: (2024)
Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System
di: Li, Yanming, et al.
Pubblicazione: (2026)
di: Li, Yanming, et al.
Pubblicazione: (2026)
AgentRM: Enhancing Agent Generalization with Reward Modeling
di: Xia, Yu, et al.
Pubblicazione: (2025)
di: Xia, Yu, et al.
Pubblicazione: (2025)
Reward Centering
di: Naik, Abhishek, et al.
Pubblicazione: (2024)
di: Naik, Abhishek, et al.
Pubblicazione: (2024)
MASPRM: Multi-Agent System Process Reward Model
di: Yazdani, Milad, et al.
Pubblicazione: (2025)
di: Yazdani, Milad, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Go Beyond Black-box Policies: Rethinking the Design of Learning Agent for Interpretable and Verifiable HVAC Control
di: An, Zhiyu, et al.
Pubblicazione: (2024) -
DIML: Differentiable Inverse Mechanism Learning from Behaviors of Multi-Agent Learning Trajectories
di: An, Zhiyu, et al.
Pubblicazione: (2026) -
MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning
di: An, Zhiyu, et al.
Pubblicazione: (2025) -
Representational Homomorphism Predicts and Improves Compositional Generalization In Transformer Language Model
di: An, Zhiyu, et al.
Pubblicazione: (2026) -
Golden-Retriever: High-Fidelity Agentic Retrieval Augmented Generation for Industrial Knowledge Base
di: An, Zhiyu, et al.
Pubblicazione: (2024)