Reward Shaping and Action Masking for Compositional Tasks using Behavior Trees and LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Potteiger, Nicholas, Samaddar, Ankita, Johnson, Taylor T., Koutsoukos, Xenofon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Out-of-Distribution Detection for Neurosymbolic Autonomous Cyber Agents
por: Samaddar, Ankita, et al.
Publicado: (2024)
por: Samaddar, Ankita, et al.
Publicado: (2024)
Designing Robust Cyber-Defense Agents with Evolving Behavior Trees
por: Potteiger, Nicholas, et al.
Publicado: (2024)
por: Potteiger, Nicholas, et al.
Publicado: (2024)
Resilient Peer-to-peer Learning based on Adaptive Aggregation
por: Bhowmick, Chandreyee, et al.
Publicado: (2025)
por: Bhowmick, Chandreyee, et al.
Publicado: (2025)
Robust Anomaly Detection with Graph Neural Networks using Controllability
por: Wei, Yifan, et al.
Publicado: (2025)
por: Wei, Yifan, et al.
Publicado: (2025)
PropEnc: A Property Encoder for Graph Neural Networks
por: Said, Anwar, et al.
Publicado: (2024)
por: Said, Anwar, et al.
Publicado: (2024)
Improving Graph Machine Learning Performance Through Feature Augmentation Based on Network Control Theory
por: Said, Anwar, et al.
Publicado: (2024)
por: Said, Anwar, et al.
Publicado: (2024)
Feature Construction Using Network Control Theory and Rank Encoding for Graph Machine Learning
por: Said, Anwar, et al.
Publicado: (2025)
por: Said, Anwar, et al.
Publicado: (2025)
Quantifying the Generalization Gap: A New Benchmark for Out-of-Distribution Graph-Based Android Malware Classification
por: Tran, Ngoc N., et al.
Publicado: (2025)
por: Tran, Ngoc N., et al.
Publicado: (2025)
Learning Backbones: Sparsifying Graphs through Zero Forcing for Effective Graph-Based Learning
por: Ahmad, Obaid Ullah, et al.
Publicado: (2025)
por: Ahmad, Obaid Ullah, et al.
Publicado: (2025)
Control-based Graph Embeddings with Data Augmentation for Contrastive Learning
por: Ahmad, Obaid Ullah, et al.
Publicado: (2024)
por: Ahmad, Obaid Ullah, et al.
Publicado: (2024)
Verification of Behavior Trees with Contingency Monitors
por: Serbinowska, Serena S., et al.
Publicado: (2024)
por: Serbinowska, Serena S., et al.
Publicado: (2024)
A Survey of Graph Unlearning
por: Said, Anwar, et al.
Publicado: (2023)
por: Said, Anwar, et al.
Publicado: (2023)
Zero Shot Coordination for Sparse Reward Tasks with Diverse Reward Shapings
por: Powell, Keenan, et al.
Publicado: (2026)
por: Powell, Keenan, et al.
Publicado: (2026)
Defining and Benchmarking a Data-Centric Design Space for Brain Graph Construction
por: Ge, Qinwen, et al.
Publicado: (2025)
por: Ge, Qinwen, et al.
Publicado: (2025)
Action-Dependent Optimality-Preserving Reward Shaping
por: Forbes, Grant C., et al.
Publicado: (2025)
por: Forbes, Grant C., et al.
Publicado: (2025)
NeuroGraph: Benchmarks for Graph Machine Learning in Brain Connectomics
por: Said, Anwar, et al.
Publicado: (2023)
por: Said, Anwar, et al.
Publicado: (2023)
From Diet to Free Lunch: Estimating Auxiliary Signal Properties using Dynamic Pruning Masks in Speech Enhancement Networks
por: Miccini, Riccardo, et al.
Publicado: (2026)
por: Miccini, Riccardo, et al.
Publicado: (2026)
Training RL Agents for Multi-Objective Network Defense Tasks
por: Molina-Markham, Andres, et al.
Publicado: (2025)
por: Molina-Markham, Andres, et al.
Publicado: (2025)
Reward Hacking Mitigation using Verifiable Composite Rewards
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025)
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025)
Cooperative Task Offloading through Asynchronous Deep Reinforcement Learning in Mobile Edge Computing for Future Networks
por: Liu, Yuelin, et al.
Publicado: (2025)
por: Liu, Yuelin, et al.
Publicado: (2025)
MAESTRO: Multi-Agent Environment Shaping through Task and Reward Optimization
por: Wu, Boyuan
Publicado: (2025)
por: Wu, Boyuan
Publicado: (2025)
Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling
por: Li, Derek, et al.
Publicado: (2025)
por: Li, Derek, et al.
Publicado: (2025)
Adapting the Behavior of Reinforcement Learning Agents to Changing Action Spaces and Reward Functions
por: de la Rosa, Raul, et al.
Publicado: (2026)
por: de la Rosa, Raul, et al.
Publicado: (2026)
Guided Star-Shaped Masked Diffusion
por: Meshchaninov, Viacheslav, et al.
Publicado: (2025)
por: Meshchaninov, Viacheslav, et al.
Publicado: (2025)
Bootstrapped Reward Shaping
por: Adamczyk, Jacob, et al.
Publicado: (2025)
por: Adamczyk, Jacob, et al.
Publicado: (2025)
Decoupling Task and Behavior: A Two-Stage Reward Curriculum in Reinforcement Learning for Robotics
por: Freitag, Kilian, et al.
Publicado: (2026)
por: Freitag, Kilian, et al.
Publicado: (2026)
Multi-Task Reward Learning from Human Ratings
por: Wu, Mingkang, et al.
Publicado: (2025)
por: Wu, Mingkang, et al.
Publicado: (2025)
Zero-Shot LLMs in Human-in-the-Loop RL: Replacing Human Feedback for Reward Shaping
por: Nazir, Mohammad Saif, et al.
Publicado: (2025)
por: Nazir, Mohammad Saif, et al.
Publicado: (2025)
Efficient Flow Matching using Latent Variables
por: Samaddar, Anirban, et al.
Publicado: (2025)
por: Samaddar, Anirban, et al.
Publicado: (2025)
Attention-Based Reward Shaping for Sparse and Delayed Rewards
por: Holmes, Ian, et al.
Publicado: (2025)
por: Holmes, Ian, et al.
Publicado: (2025)
Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
por: Yu, Zichao, et al.
Publicado: (2025)
por: Yu, Zichao, et al.
Publicado: (2025)
TIPS: Turn-Level Information-Potential Reward Shaping for Search-Augmented LLMs
por: Xie, Yutao, et al.
Publicado: (2026)
por: Xie, Yutao, et al.
Publicado: (2026)
TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
por: Naim, Omar, et al.
Publicado: (2025)
por: Naim, Omar, et al.
Publicado: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models
por: Luo, Ziwei, et al.
Publicado: (2026)
por: Luo, Ziwei, et al.
Publicado: (2026)
Physics-Aware Combinatorial Assembly Sequence Planning using Data-free Action Masking
por: Liu, Ruixuan, et al.
Publicado: (2024)
por: Liu, Ruixuan, et al.
Publicado: (2024)
Mask Is What DLLM Needs: A Masked Data Training Paradigm for Diffusion LLMs
por: Ma, Linrui, et al.
Publicado: (2026)
por: Ma, Linrui, et al.
Publicado: (2026)
BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping
por: Lidayan, Aly, et al.
Publicado: (2024)
por: Lidayan, Aly, et al.
Publicado: (2024)
Aligning LLMs with Domain Invariant Reward Models
por: Wu, David, et al.
Publicado: (2025)
por: Wu, David, et al.
Publicado: (2025)
Learning Explainable Dense Reward Shapes via Bayesian Optimization
por: Koo, Ryan, et al.
Publicado: (2025)
por: Koo, Ryan, et al.
Publicado: (2025)
Ejemplares similares
-
Out-of-Distribution Detection for Neurosymbolic Autonomous Cyber Agents
por: Samaddar, Ankita, et al.
Publicado: (2024) -
Designing Robust Cyber-Defense Agents with Evolving Behavior Trees
por: Potteiger, Nicholas, et al.
Publicado: (2024) -
Resilient Peer-to-peer Learning based on Adaptive Aggregation
por: Bhowmick, Chandreyee, et al.
Publicado: (2025) -
Robust Anomaly Detection with Graph Neural Networks using Controllability
por: Wei, Yifan, et al.
Publicado: (2025) -
PropEnc: A Property Encoder for Graph Neural Networks
por: Said, Anwar, et al.
Publicado: (2024)