Improve Value Estimation of Q Function and Reshape Reward with Monte Carlo Tree Search
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Li, Jiamian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Value Internalization: Learning and Generalizing from Social Reward
von: Rong, Frieda, et al.
Veröffentlicht: (2024)
von: Rong, Frieda, et al.
Veröffentlicht: (2024)
RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value Factorization
von: Shen, Siqi, et al.
Veröffentlicht: (2023)
von: Shen, Siqi, et al.
Veröffentlicht: (2023)
SMCEvolve: Principled Scientific Discovery via Sequential Monte Carlo Evolution
von: Jiang, Jiachen, et al.
Veröffentlicht: (2026)
von: Jiang, Jiachen, et al.
Veröffentlicht: (2026)
Mathematical Framework for Custom Reward Functions in Job Application Evaluation using Reinforcement Learning
von: Jain, Shreyansh, et al.
Veröffentlicht: (2025)
von: Jain, Shreyansh, et al.
Veröffentlicht: (2025)
Lookahead Pathology in Monte-Carlo Tree Search
von: Nguyen, Khoi P. N., et al.
Veröffentlicht: (2022)
von: Nguyen, Khoi P. N., et al.
Veröffentlicht: (2022)
Value-Based Rationales Improve Social Experience: A Multiagent Simulation Study
von: Tzeng, Sz-Ting, et al.
Veröffentlicht: (2024)
von: Tzeng, Sz-Ting, et al.
Veröffentlicht: (2024)
Distributed Multi-Agent Reinforcement Learning Based on Graph-Induced Local Value Functions
von: Jing, Gangshan, et al.
Veröffentlicht: (2022)
von: Jing, Gangshan, et al.
Veröffentlicht: (2022)
SrSv: Integrating Sequential Rollouts with Sequential Value Estimation for Multi-agent Reinforcement Learning
von: Wan, Xu, et al.
Veröffentlicht: (2025)
von: Wan, Xu, et al.
Veröffentlicht: (2025)
LERO: LLM-driven Evolutionary framework with Hybrid Rewards and Enhanced Observation for Multi-Agent Reinforcement Learning
von: Wei, Yuan, et al.
Veröffentlicht: (2025)
von: Wei, Yuan, et al.
Veröffentlicht: (2025)
When Is Diversity Rewarded in Cooperative Multi-Agent Learning?
von: Amir, Michael, et al.
Veröffentlicht: (2025)
von: Amir, Michael, et al.
Veröffentlicht: (2025)
Multi-Agent Reinforcement Learning with a Hierarchy of Reward Machines
von: Zheng, Xuejing, et al.
Veröffentlicht: (2024)
von: Zheng, Xuejing, et al.
Veröffentlicht: (2024)
QSIM: Mitigating Overestimation in Multi-Agent Reinforcement Learning via Action Similarity Weighted Q-Learning
von: Li, Yuanjun, et al.
Veröffentlicht: (2026)
von: Li, Yuanjun, et al.
Veröffentlicht: (2026)
Reward-Independent Messaging for Decentralized Multi-Agent Reinforcement Learning
von: Yoshida, Naoto, et al.
Veröffentlicht: (2025)
von: Yoshida, Naoto, et al.
Veröffentlicht: (2025)
A finite time analysis of distributed Q-learning
von: Lim, Han-Dong, et al.
Veröffentlicht: (2024)
von: Lim, Han-Dong, et al.
Veröffentlicht: (2024)
Balancing Act: Prioritization Strategies for LLM-Designed Restless Bandit Rewards
von: Verma, Shresth, et al.
Veröffentlicht: (2024)
von: Verma, Shresth, et al.
Veröffentlicht: (2024)
Incorporating Human Flexibility through Reward Preferences in Human-AI Teaming
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2023)
von: Bhambri, Siddhant, et al.
Veröffentlicht: (2023)
Achieving Optimal Tissue Repair Through MARL with Reward Shaping and Curriculum Learning
von: Khan, Muhammad Al-Zafar, et al.
Veröffentlicht: (2025)
von: Khan, Muhammad Al-Zafar, et al.
Veröffentlicht: (2025)
Embodiment-Induced Coordination Regimes in Tabular Multi-Agent Q-Learning
von: Atif, Muhammad Ahmed, et al.
Veröffentlicht: (2026)
von: Atif, Muhammad Ahmed, et al.
Veröffentlicht: (2026)
Beyond Shallow Behavior: Task-Efficient Value-Based Multi-Task Offline MARL via Skill Discovery
von: Wang, Xun, et al.
Veröffentlicht: (2025)
von: Wang, Xun, et al.
Veröffentlicht: (2025)
Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning via Incorporating Generalized Human Expertise
von: Wu, Xuefei, et al.
Veröffentlicht: (2025)
von: Wu, Xuefei, et al.
Veröffentlicht: (2025)
Large Language Model-Based Reward Design for Deep Reinforcement Learning-Driven Autonomous Cyber Defense
von: Mukherjee, Sayak, et al.
Veröffentlicht: (2025)
von: Mukherjee, Sayak, et al.
Veröffentlicht: (2025)
GHQ: Grouped Hybrid Q Learning for Heterogeneous Cooperative Multi-agent Reinforcement Learning
von: Yu, Xiaoyang, et al.
Veröffentlicht: (2023)
von: Yu, Xiaoyang, et al.
Veröffentlicht: (2023)
Using Deep Q-Learning to Dynamically Toggle between Push/Pull Actions in Computational Trust Mechanisms
von: Lygizou, Zoi, et al.
Veröffentlicht: (2024)
von: Lygizou, Zoi, et al.
Veröffentlicht: (2024)
VDFD: Multi-Agent Value Decomposition Framework with Disentangled World Model
von: Wang, Zhizun, et al.
Veröffentlicht: (2023)
von: Wang, Zhizun, et al.
Veröffentlicht: (2023)
Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness
von: Parthasarathy, Ambreesh, et al.
Veröffentlicht: (2025)
von: Parthasarathy, Ambreesh, et al.
Veröffentlicht: (2025)
Discovering Sensorimotor Agency in Cellular Automata using Diversity Search
von: Hamon, Gautier, et al.
Veröffentlicht: (2024)
von: Hamon, Gautier, et al.
Veröffentlicht: (2024)
Meritocratic Fairness in Budgeted Combinatorial Multi-armed Bandits via Shapley Values
von: Sharma, Shradha, et al.
Veröffentlicht: (2026)
von: Sharma, Shradha, et al.
Veröffentlicht: (2026)
SPIRAL: Symbolic LLM Planning via Grounded and Reflective Search
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
BMG-Q: Localized Bipartite Match Graph Attention Q-Learning for Ride-Pooling Order Dispatch
von: Hu, Yulong, et al.
Veröffentlicht: (2025)
von: Hu, Yulong, et al.
Veröffentlicht: (2025)
Multi-Agent Inverse Q-Learning from Demonstrations
von: Haynam, Nathaniel, et al.
Veröffentlicht: (2025)
von: Haynam, Nathaniel, et al.
Veröffentlicht: (2025)
Redistributing Rewards Across Time and Agents for Multi-Agent Reinforcement Learning
von: Kapoor, Aditya, et al.
Veröffentlicht: (2025)
von: Kapoor, Aditya, et al.
Veröffentlicht: (2025)
Assigning Credit with Partial Reward Decoupling in Multi-Agent Proximal Policy Optimization
von: Kapoor, Aditya, et al.
Veröffentlicht: (2024)
von: Kapoor, Aditya, et al.
Veröffentlicht: (2024)
EvoMem: Improving Multi-Agent Planning with Dual-Evolving Memory
von: Fan, Wenzhe, et al.
Veröffentlicht: (2025)
von: Fan, Wenzhe, et al.
Veröffentlicht: (2025)
A Subgoal-driven Framework for Improving Long-Horizon LLM Agents
von: Wang, Taiyi, et al.
Veröffentlicht: (2026)
von: Wang, Taiyi, et al.
Veröffentlicht: (2026)
Real-time Adapting Routing (RAR): Improving Efficiency Through Continuous Learning in Software Powered by Layered Foundation Models
von: Vasilevski, Kirill, et al.
Veröffentlicht: (2024)
von: Vasilevski, Kirill, et al.
Veröffentlicht: (2024)
MFC-EQ: Mean-Field Control with Envelope Q-Learning for Moving Decentralized Agents in Formation
von: Lin, Qiushi, et al.
Veröffentlicht: (2024)
von: Lin, Qiushi, et al.
Veröffentlicht: (2024)
Improving Mixed-Criticality Scheduling with Reinforcement Learning
von: El-Mahdy, Muhammad, et al.
Veröffentlicht: (2025)
von: El-Mahdy, Muhammad, et al.
Veröffentlicht: (2025)
Graph Attention-Guided Search for Dense Multi-Agent Pathfinding
von: Jain, Rishabh, et al.
Veröffentlicht: (2025)
von: Jain, Rishabh, et al.
Veröffentlicht: (2025)
Innate-Values-driven Reinforcement Learning based Cooperative Multi-Agent Cognitive Modeling
von: Yang, Qin
Veröffentlicht: (2024)
von: Yang, Qin
Veröffentlicht: (2024)
Multi-Agent Deep Q-Network with Layer-based Communication Channel for Autonomous Internal Logistics Vehicle Scheduling in Smart Manufacturing
von: Feizabadi, Mohammad, et al.
Veröffentlicht: (2024)
von: Feizabadi, Mohammad, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Value Internalization: Learning and Generalizing from Social Reward
von: Rong, Frieda, et al.
Veröffentlicht: (2024) -
RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value Factorization
von: Shen, Siqi, et al.
Veröffentlicht: (2023) -
SMCEvolve: Principled Scientific Discovery via Sequential Monte Carlo Evolution
von: Jiang, Jiachen, et al.
Veröffentlicht: (2026) -
Mathematical Framework for Custom Reward Functions in Job Application Evaluation using Reinforcement Learning
von: Jain, Shreyansh, et al.
Veröffentlicht: (2025) -
Lookahead Pathology in Monte-Carlo Tree Search
von: Nguyen, Khoi P. N., et al.
Veröffentlicht: (2022)