Learning in complex action spaces without policy gradients
Fuente:
arXiv
Guardado en:
| Autores principales: | Tavakoli, Arash, Ghiassian, Sina, Rakićević, Nemanja |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
In-context Exploration-Exploitation for Reinforcement Learning
por: Dai, Zhenwen, et al.
Publicado: (2024)
por: Dai, Zhenwen, et al.
Publicado: (2024)
Soft Preference Optimization: Aligning Language Models to Expert Distributions
por: Sharifnassab, Arsalan, et al.
Publicado: (2024)
por: Sharifnassab, Arsalan, et al.
Publicado: (2024)
Auxiliary task discovery through generate-and-test
por: Rafiee, Banafsheh, et al.
Publicado: (2022)
por: Rafiee, Banafsheh, et al.
Publicado: (2022)
Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
por: Ding, Dongsheng, et al.
Publicado: (2022)
por: Ding, Dongsheng, et al.
Publicado: (2022)
Generalizing soft actor-critic algorithms to discrete action spaces
por: Zhang, Le, et al.
Publicado: (2024)
por: Zhang, Le, et al.
Publicado: (2024)
Combining LLM decision and RL action selection to improve RL policy for adaptive interventions
por: Karine, Karine, et al.
Publicado: (2025)
por: Karine, Karine, et al.
Publicado: (2025)
Active Inference and Reinforcement Learning: A unified inference on continuous state and action spaces under partial observability
por: Malekzadeh, Parvin, et al.
Publicado: (2022)
por: Malekzadeh, Parvin, et al.
Publicado: (2022)
Hallucination Detection on a Budget: Efficient Bayesian Estimation of Semantic Entropy
por: Ciosek, Kamil, et al.
Publicado: (2025)
por: Ciosek, Kamil, et al.
Publicado: (2025)
Boosting Hierarchical Reinforcement Learning with Meta-Learning for Complex Task Adaptation
por: Khajooeinejad, Arash, et al.
Publicado: (2024)
por: Khajooeinejad, Arash, et al.
Publicado: (2024)
Decoding Funded Research: Comparative Analysis of Topic Models and Uncovering the Effect of Gender and Geographic Location
por: Kafiabad, Shirin Tavakoli, et al.
Publicado: (2025)
por: Kafiabad, Shirin Tavakoli, et al.
Publicado: (2025)
Finding Optimal Trading History in Reinforcement Learning for Stock Market Trading
por: Montazeri, Sina, et al.
Publicado: (2025)
por: Montazeri, Sina, et al.
Publicado: (2025)
Large Language Models for Sequential Decision-Making: Improving In-Context Learning via Supervised Fine-Tuning
por: Zhang, Minmin, et al.
Publicado: (2026)
por: Zhang, Minmin, et al.
Publicado: (2026)
Learning to Act without Actions
por: Schmidt, Dominik, et al.
Publicado: (2023)
por: Schmidt, Dominik, et al.
Publicado: (2023)
Exploiting Expertise of Non-Expert and Diverse Agents in Social Bandit Learning: A Free Energy Approach
por: Mirzaei, Erfan, et al.
Publicado: (2026)
por: Mirzaei, Erfan, et al.
Publicado: (2026)
Learning with Conflicts of Interest
por: Aryal, Nischal, et al.
Publicado: (2026)
por: Aryal, Nischal, et al.
Publicado: (2026)
Deep deterministic policy gradient with symmetric data augmentation for lateral attitude tracking control of a fixed-wing aircraft
por: Li, Yifei, et al.
Publicado: (2024)
por: Li, Yifei, et al.
Publicado: (2024)
Edge-DIRECT: A Deep Reinforcement Learning-based Method for Solving Heterogeneous Electric Vehicle Routing Problem with Time Window Constraints
por: Mozhdehi, Arash, et al.
Publicado: (2024)
por: Mozhdehi, Arash, et al.
Publicado: (2024)
Softmax gradient policy for variance minimization and risk-averse multi armed bandits
por: Turinici, Gabriel
Publicado: (2026)
por: Turinici, Gabriel
Publicado: (2026)
Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity
por: Yang, Yan, et al.
Publicado: (2024)
por: Yang, Yan, et al.
Publicado: (2024)
Using the Path of Least Resistance to Explain Deep Networks
por: Salek, Sina, et al.
Publicado: (2025)
por: Salek, Sina, et al.
Publicado: (2025)
Streaming Flow Policy: Simplifying diffusion/flow-matching policies by treating action trajectories as flow trajectories
por: Jiang, Sunshine, et al.
Publicado: (2025)
por: Jiang, Sunshine, et al.
Publicado: (2025)
Learning with Logical Constraints but without Shortcut Satisfaction
por: Li, Zenan, et al.
Publicado: (2024)
por: Li, Zenan, et al.
Publicado: (2024)
Safe-Support Q-Learning: Learning without Unsafe Exploration
por: Lim, Yeeun, et al.
Publicado: (2026)
por: Lim, Yeeun, et al.
Publicado: (2026)
SED2AM: Solving Multi-Trip Time-Dependent Vehicle Routing Problem using Deep Reinforcement Learning
por: Mozhdehi, Arash, et al.
Publicado: (2025)
por: Mozhdehi, Arash, et al.
Publicado: (2025)
Contrastive Preference Learning: Learning from Human Feedback without RL
por: Hejna, Joey, et al.
Publicado: (2023)
por: Hejna, Joey, et al.
Publicado: (2023)
Class-Imbalanced Graph Learning without Class Rebalancing
por: Liu, Zhining, et al.
Publicado: (2023)
por: Liu, Zhining, et al.
Publicado: (2023)
Learning Relational Tabular Data without Shared Features
por: Wu, Zhaomin, et al.
Publicado: (2025)
por: Wu, Zhaomin, et al.
Publicado: (2025)
Learning Over Dirty Data with Minimal Repairs
por: Zhen, Cheng, et al.
Publicado: (2025)
por: Zhen, Cheng, et al.
Publicado: (2025)
Augmenting the action space with conventions to improve multi-agent cooperation in Hanabi
por: Bredell, F., et al.
Publicado: (2024)
por: Bredell, F., et al.
Publicado: (2024)
RL-GPT: Integrating Reinforcement Learning and Code-as-policy
por: Liu, Shaoteng, et al.
Publicado: (2024)
por: Liu, Shaoteng, et al.
Publicado: (2024)
Learning without Global Backpropagation via Synergistic Information Distillation
por: Ye, Chenhao, et al.
Publicado: (2025)
por: Ye, Chenhao, et al.
Publicado: (2025)
Online Reinforcement Learning in Non-Stationary Context-Driven Environments
por: Hamadanian, Pouya, et al.
Publicado: (2023)
por: Hamadanian, Pouya, et al.
Publicado: (2023)
Graph Neural Thompson Sampling
por: Wu, Shuang, et al.
Publicado: (2024)
por: Wu, Shuang, et al.
Publicado: (2024)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
por: Rezazadeh, Navid, et al.
Publicado: (2026)
por: Rezazadeh, Navid, et al.
Publicado: (2026)
SimMerge: Learning to Select Merge Operators from Similarity Signals
por: Bolton, Oliver, et al.
Publicado: (2026)
por: Bolton, Oliver, et al.
Publicado: (2026)
State-space models can learn in-context by gradient descent
por: Sushma, Neeraj Mohan, et al.
Publicado: (2024)
por: Sushma, Neeraj Mohan, et al.
Publicado: (2024)
Complex behavior from intrinsic motivation to occupy action-state path space
por: Ramírez-Ruiz, Jorge, et al.
Publicado: (2022)
por: Ramírez-Ruiz, Jorge, et al.
Publicado: (2022)
Residual Q-Learning: Offline and Online Policy Customization without Value
por: Li, Chenran, et al.
Publicado: (2023)
por: Li, Chenran, et al.
Publicado: (2023)
Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting
por: Nam, Yunhun, et al.
Publicado: (2025)
por: Nam, Yunhun, et al.
Publicado: (2025)
Automatic generation of insights from workers' actions in industrial workflows with explainable Machine Learning
por: de Arriba-Pérez, Francisco, et al.
Publicado: (2024)
por: de Arriba-Pérez, Francisco, et al.
Publicado: (2024)
Ejemplares similares
-
In-context Exploration-Exploitation for Reinforcement Learning
por: Dai, Zhenwen, et al.
Publicado: (2024) -
Soft Preference Optimization: Aligning Language Models to Expert Distributions
por: Sharifnassab, Arsalan, et al.
Publicado: (2024) -
Auxiliary task discovery through generate-and-test
por: Rafiee, Banafsheh, et al.
Publicado: (2022) -
Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
por: Ding, Dongsheng, et al.
Publicado: (2022) -
Generalizing soft actor-critic algorithms to discrete action spaces
por: Zhang, Le, et al.
Publicado: (2024)