Overcoming Valid Action Suppression in Unmasked Policy Gradient Algorithms
Fuente:
arXiv
Guardado en:
| Autores principales: | Zabounidis, Renos, Siegelmann, Roy, Qadri, Mohamad, Kim, Woojun, Stepputtis, Simon, Sycara, Katia P. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SCALAR: Learning and Composing Skills through LLM Guided Symbolic Planning and Deep RL Grounding
por: Zabounidis, Renos, et al.
Publicado: (2026)
por: Zabounidis, Renos, et al.
Publicado: (2026)
B3C: A Minimalist Approach to Offline Multi-Agent Reinforcement Learning
por: Kim, Woojun, et al.
Publicado: (2025)
por: Kim, Woojun, et al.
Publicado: (2025)
Dual Prototype Evolving for Test-Time Generalization of Vision-Language Models
por: Zhang, Ce, et al.
Publicado: (2024)
por: Zhang, Ce, et al.
Publicado: (2024)
Fair Cooperation in Mixed-Motive Games via Conflict-Aware Gradient Adjustment
por: Kim, Woojun, et al.
Publicado: (2025)
por: Kim, Woojun, et al.
Publicado: (2025)
Model-Agnostic Policy Explanations with Large Language Models
por: Xi-Jia, Zhang, et al.
Publicado: (2025)
por: Xi-Jia, Zhang, et al.
Publicado: (2025)
Adaptively Coordinating with Novel Partners via Learned Latent Strategies
por: Li, Benjamin, et al.
Publicado: (2025)
por: Li, Benjamin, et al.
Publicado: (2025)
Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
por: Zabounidis, Renos, et al.
Publicado: (2025)
por: Zabounidis, Renos, et al.
Publicado: (2025)
HiMemFormer: Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation
por: Wang, Zirui, et al.
Publicado: (2024)
por: Wang, Zirui, et al.
Publicado: (2024)
Overcoming Slow Decision Frequencies in Continuous Control: Model-Based Sequence Reinforcement Learning for Model-Free Control
por: Patel, Devdhar, et al.
Publicado: (2024)
por: Patel, Devdhar, et al.
Publicado: (2024)
Theory of Mind Guided Strategy Adaptation for Zero-Shot Coordination
por: Ni, Andrew, et al.
Publicado: (2026)
por: Ni, Andrew, et al.
Publicado: (2026)
ShapeGrasp: Zero-Shot Task-Oriented Grasping with Large Language Models through Geometric Decomposition
por: Li, Samuel, et al.
Publicado: (2024)
por: Li, Samuel, et al.
Publicado: (2024)
LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework
por: Kim, Woojun, et al.
Publicado: (2023)
por: Kim, Woojun, et al.
Publicado: (2023)
Multi-Agent Transfer Learning via Temporal Contrastive Learning
por: Zeng, Weihao, et al.
Publicado: (2024)
por: Zeng, Weihao, et al.
Publicado: (2024)
Enhancing Vision-Language Few-Shot Adaptation with Negative Learning
por: Zhang, Ce, et al.
Publicado: (2024)
por: Zhang, Ce, et al.
Publicado: (2024)
GL-NeRF: Gauss-Laguerre Quadrature Enables Training-Free NeRF Acceleration
por: Yong, Silong, et al.
Publicado: (2024)
por: Yong, Silong, et al.
Publicado: (2024)
Segment-driven Structural Induction and Semantic Alignment for Heterogeneous Tabular Representation
por: Jung, Woojun, et al.
Publicado: (2026)
por: Jung, Woojun, et al.
Publicado: (2026)
Decision ConvFormer: Local Filtering in MetaFormer is Sufficient for Decision Making
por: Kim, Jeonghye, et al.
Publicado: (2023)
por: Kim, Jeonghye, et al.
Publicado: (2023)
Adaptive $Q$-Aid for Conditional Supervised Learning in Offline Reinforcement Learning
por: Kim, Jeonghye, et al.
Publicado: (2024)
por: Kim, Jeonghye, et al.
Publicado: (2024)
Learning Unmasking Policies for Diffusion Language Models
por: Jazbec, Metod, et al.
Publicado: (2025)
por: Jazbec, Metod, et al.
Publicado: (2025)
CBGT-Net: A Neuromimetic Architecture for Robust Classification of Streaming Data
por: Sharma, Shreya, et al.
Publicado: (2024)
por: Sharma, Shreya, et al.
Publicado: (2024)
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient
por: Zhou, Zhongzhu, et al.
Publicado: (2025)
por: Zhou, Zhongzhu, et al.
Publicado: (2025)
A Comparison of Imitation Learning Algorithms for Bimanual Manipulation
por: Drolet, Michael, et al.
Publicado: (2024)
por: Drolet, Michael, et al.
Publicado: (2024)
CARE: Enhancing Safety of Visual Navigation through Collision Avoidance via Repulsive Estimation
por: Kim, Joonkyung, et al.
Publicado: (2025)
por: Kim, Joonkyung, et al.
Publicado: (2025)
Spectral-Aware Global Fusion for RGB-Thermal Semantic Segmentation
por: Zhang, Ce, et al.
Publicado: (2025)
por: Zhang, Ce, et al.
Publicado: (2025)
HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph Generation
por: Zhang, Ce, et al.
Publicado: (2024)
por: Zhang, Ce, et al.
Publicado: (2024)
On the Second-Order Convergence of Biased Policy Gradient Algorithms
por: Mu, Siqiao, et al.
Publicado: (2023)
por: Mu, Siqiao, et al.
Publicado: (2023)
DyPNIPP: Predicting Environment Dynamics for RL-based Robust Informative Path Planning
por: Deolasee, Srujan, et al.
Publicado: (2024)
por: Deolasee, Srujan, et al.
Publicado: (2024)
Learning General Policies with Policy Gradient Methods
por: Ståhlberg, Simon, et al.
Publicado: (2025)
por: Ståhlberg, Simon, et al.
Publicado: (2025)
Your Learned Constraint is Secretly a Backward Reachable Tube
por: Qadri, Mohamad, et al.
Publicado: (2025)
por: Qadri, Mohamad, et al.
Publicado: (2025)
Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
por: Hong, Chunsan, et al.
Publicado: (2025)
por: Hong, Chunsan, et al.
Publicado: (2025)
Modeling Latent Partner Strategies for Adaptive Zero-Shot Human-Agent Collaboration
por: Li, Benjamin, et al.
Publicado: (2025)
por: Li, Benjamin, et al.
Publicado: (2025)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
por: Armandpour, Mohammadreza, et al.
Publicado: (2026)
por: Armandpour, Mohammadreza, et al.
Publicado: (2026)
Energy-Based Transfer for Reinforcement Learning
por: Deng, Zeyun, et al.
Publicado: (2025)
por: Deng, Zeyun, et al.
Publicado: (2025)
FlowPG: Action-constrained Policy Gradient with Normalizing Flows
por: Brahmanage, Janaka Chathuranga, et al.
Publicado: (2024)
por: Brahmanage, Janaka Chathuranga, et al.
Publicado: (2024)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
por: Corrado, Nicholas E., et al.
Publicado: (2023)
por: Corrado, Nicholas E., et al.
Publicado: (2023)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
por: Meng, Wenjia, et al.
Publicado: (2024)
por: Meng, Wenjia, et al.
Publicado: (2024)
Efficient and Optimal Policy Gradient Algorithm for Corrupted Multi-armed Bandits
por: Liu, Jiayuan, et al.
Publicado: (2025)
por: Liu, Jiayuan, et al.
Publicado: (2025)
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
por: Lee, Jongmin, et al.
Publicado: (2025)
por: Lee, Jongmin, et al.
Publicado: (2025)
Algorithm-Relative Trajectory Valuation in Policy Gradient Control
por: Li, Shihao, et al.
Publicado: (2025)
por: Li, Shihao, et al.
Publicado: (2025)
Speaking the Language of Teamwork: LLM-Guided Credit Assignment in Multi-Agent Reinforcement Learning
por: Lin, Muhan, et al.
Publicado: (2025)
por: Lin, Muhan, et al.
Publicado: (2025)
Ejemplares similares
-
SCALAR: Learning and Composing Skills through LLM Guided Symbolic Planning and Deep RL Grounding
por: Zabounidis, Renos, et al.
Publicado: (2026) -
B3C: A Minimalist Approach to Offline Multi-Agent Reinforcement Learning
por: Kim, Woojun, et al.
Publicado: (2025) -
Dual Prototype Evolving for Test-Time Generalization of Vision-Language Models
por: Zhang, Ce, et al.
Publicado: (2024) -
Fair Cooperation in Mixed-Motive Games via Conflict-Aware Gradient Adjustment
por: Kim, Woojun, et al.
Publicado: (2025) -
Model-Agnostic Policy Explanations with Large Language Models
por: Xi-Jia, Zhang, et al.
Publicado: (2025)