Discrete Variational Autoencoding via Policy Search
Fuente:
arXiv
Saved in:
| Main Authors: | Drolet, Michael, Al-Hafez, Firas, Bhatt, Aditya, Peters, Jan, Arenz, Oleg |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Role of Domain Randomization in Training Diffusion Policies for Whole-Body Humanoid Control
by: Kaidanov, Oleg, et al.
Published: (2024)
by: Kaidanov, Oleg, et al.
Published: (2024)
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
by: Diwan, Anish, et al.
Published: (2026)
by: Diwan, Anish, et al.
Published: (2026)
MuJoCo MPC for Humanoid Control: Evaluation on HumanoidBench
by: Meser, Moritz, et al.
Published: (2024)
by: Meser, Moritz, et al.
Published: (2024)
Learning and Blending Robot Hugging Behaviors in Time and Space
by: Drolet, Michael, et al.
Published: (2022)
by: Drolet, Michael, et al.
Published: (2022)
Mutual Information Tracks Policy Coherence in Reinforcement Learning
by: Reid, Cameron, et al.
Published: (2025)
by: Reid, Cameron, et al.
Published: (2025)
Context-Aware Deep Lagrangian Networks for Model Predictive Control
by: Schulze, Lucas, et al.
Published: (2025)
by: Schulze, Lucas, et al.
Published: (2025)
Exciting Action: Investigating Efficient Exploration for Learning Musculoskeletal Humanoid Locomotion
by: Geiß, Henri-Jacques, et al.
Published: (2024)
by: Geiß, Henri-Jacques, et al.
Published: (2024)
Learning Hierarchical Domain Models Through Environment-Grounded Interaction
by: Kienle, Claudius, et al.
Published: (2025)
by: Kienle, Claudius, et al.
Published: (2025)
The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Policy Learning
by: von Hartz, Jan Ole, et al.
Published: (2025)
by: von Hartz, Jan Ole, et al.
Published: (2025)
Variational Distillation of Diffusion Policies into Mixture of Experts
by: Zhou, Hongyi, et al.
Published: (2024)
by: Zhou, Hongyi, et al.
Published: (2024)
Q-Guided Stein Variational Model Predictive Control via RL-informed Policy Prior
by: Cai, Shizhe, et al.
Published: (2025)
by: Cai, Shizhe, et al.
Published: (2025)
Reliable and Efficient Multi-Agent Coordination via Graph Neural Network Variational Autoencoders
by: Meng, Yue, et al.
Published: (2025)
by: Meng, Yue, et al.
Published: (2025)
Neuro-Symbolic Imitation Learning: Discovering Symbolic Abstractions for Skill Learning
by: Keller, Leon, et al.
Published: (2025)
by: Keller, Leon, et al.
Published: (2025)
A Multi-Fidelity Control Variate Approach for Policy Gradient Estimation
by: Liu, Xinjie, et al.
Published: (2025)
by: Liu, Xinjie, et al.
Published: (2025)
Unsupervised Skill Discovery for Robotic Manipulation through Automatic Task Generation
by: Jansonnie, Paul, et al.
Published: (2024)
by: Jansonnie, Paul, et al.
Published: (2024)
Motion Planning Diffusion: Learning and Planning of Robot Motions with Diffusion Models
by: Carvalho, Joao, et al.
Published: (2023)
by: Carvalho, Joao, et al.
Published: (2023)
Safe Exploration via Policy Priors
by: Wendl, Manuel, et al.
Published: (2026)
by: Wendl, Manuel, et al.
Published: (2026)
Drifting Field Policy: A One-Step Generative Policy via Wasserstein Gradient Flow
by: Koo, Juil, et al.
Published: (2026)
by: Koo, Juil, et al.
Published: (2026)
Redistributing Rewards Across Time and Agents for Multi-Agent Reinforcement Learning
by: Kapoor, Aditya, et al.
Published: (2025)
by: Kapoor, Aditya, et al.
Published: (2025)
IMLE Policy: Fast and Sample Efficient Visuomotor Policy Learning via Implicit Maximum Likelihood Estimation
by: Rana, Krishan, et al.
Published: (2025)
by: Rana, Krishan, et al.
Published: (2025)
Robust Policy Learning via Offline Skill Diffusion
by: Kim, Woo Kyung, et al.
Published: (2024)
by: Kim, Woo Kyung, et al.
Published: (2024)
Policy-Guided Diffusion
by: Jackson, Matthew Thomas, et al.
Published: (2024)
by: Jackson, Matthew Thomas, et al.
Published: (2024)
Learning from Observation: A Survey of Recent Advances
by: Burnwal, Returaj, et al.
Published: (2025)
by: Burnwal, Returaj, et al.
Published: (2025)
Low-Rank Agent-Specific Adaptation (LoRASA) for Multi-Agent Policy Learning
by: Zhang, Beining, et al.
Published: (2025)
by: Zhang, Beining, et al.
Published: (2025)
MSG: Multi-Stream Generative Policies for Sample-Efficient Robotic Manipulation
by: von Hartz, Jan Ole, et al.
Published: (2025)
by: von Hartz, Jan Ole, et al.
Published: (2025)
Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation
by: Nakanishi, Kosuke, et al.
Published: (2025)
by: Nakanishi, Kosuke, et al.
Published: (2025)
Tactile MNIST: Benchmarking Active Tactile Perception
by: Schneider, Tim, et al.
Published: (2025)
by: Schneider, Tim, et al.
Published: (2025)
Scaling Policy Gradient Quality-Diversity with Massive Parallelization via Behavioral Variations
by: Mitsides, Konstantinos, et al.
Published: (2025)
by: Mitsides, Konstantinos, et al.
Published: (2025)
Data Augmentation for Instruction Following Policies via Trajectory Segmentation
by: Höpner, Niklas, et al.
Published: (2025)
by: Höpner, Niklas, et al.
Published: (2025)
In-Context Policy Adaptation via Cross-Domain Skill Diffusion
by: Yoo, Minjong, et al.
Published: (2025)
by: Yoo, Minjong, et al.
Published: (2025)
Multi-Modal Manipulation via Multi-Modal Policy Consensus
by: Chen, Haonan, et al.
Published: (2025)
by: Chen, Haonan, et al.
Published: (2025)
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
by: Xu, Charles, et al.
Published: (2024)
by: Xu, Charles, et al.
Published: (2024)
FFHFlow: Diverse and Uncertainty-Aware Dexterous Grasp Generation via Flow Variational Inference
by: Feng, Qian, et al.
Published: (2024)
by: Feng, Qian, et al.
Published: (2024)
Learning Long-Context Diffusion Policies via Past-Token Prediction
by: Torne, Marcel, et al.
Published: (2025)
by: Torne, Marcel, et al.
Published: (2025)
SPRINT: Scalable Policy Pre-Training via Language Instruction Relabeling
by: Zhang, Jesse, et al.
Published: (2023)
by: Zhang, Jesse, et al.
Published: (2023)
Assigning Credit with Partial Reward Decoupling in Multi-Agent Proximal Policy Optimization
by: Kapoor, Aditya, et al.
Published: (2024)
by: Kapoor, Aditya, et al.
Published: (2024)
Offline Reinforcement Learning with Discrete Diffusion Skills
by: Qiao, RuiXi, et al.
Published: (2025)
by: Qiao, RuiXi, et al.
Published: (2025)
SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
by: Jin, Yang, et al.
Published: (2025)
by: Jin, Yang, et al.
Published: (2025)
DISC: Decoupling Instruction from State-Conditioned Control via Policy Generation
by: Ren, Hanxiang, et al.
Published: (2026)
by: Ren, Hanxiang, et al.
Published: (2026)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
Similar Items
-
The Role of Domain Randomization in Training Diffusion Policies for Whole-Body Humanoid Control
by: Kaidanov, Oleg, et al.
Published: (2024) -
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
by: Diwan, Anish, et al.
Published: (2026) -
MuJoCo MPC for Humanoid Control: Evaluation on HumanoidBench
by: Meser, Moritz, et al.
Published: (2024) -
Learning and Blending Robot Hugging Behaviors in Time and Space
by: Drolet, Michael, et al.
Published: (2022) -
Mutual Information Tracks Policy Coherence in Reinforcement Learning
by: Reid, Cameron, et al.
Published: (2025)