Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Sungyoung, Kim, Dohyeong, Balachandar, Eshan, Mustafaoglu, Zelal Su, Pingali, Keshav |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
por: Su, Zelal, et al.
Publicado: (2026)
por: Su, Zelal, et al.
Publicado: (2026)
Evolutionary Policy Optimization
por: Mustafaoglu, Zelal Su "Lain", et al.
Publicado: (2025)
por: Mustafaoglu, Zelal Su "Lain", et al.
Publicado: (2025)
ReFORM: Reflected Flows for On-support Offline RL via Noise Manipulation
por: Zhang, Songyuan, et al.
Publicado: (2026)
por: Zhang, Songyuan, et al.
Publicado: (2026)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
por: You, Bozhi, et al.
Publicado: (2025)
por: You, Bozhi, et al.
Publicado: (2025)
Causal Flow Q-Learning for Robust Offline Reinforcement Learning
por: Li, Mingxuan, et al.
Publicado: (2026)
por: Li, Mingxuan, et al.
Publicado: (2026)
FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning
por: Alles, Marvin, et al.
Publicado: (2025)
por: Alles, Marvin, et al.
Publicado: (2025)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
por: Kim, Changyeon, et al.
Publicado: (2025)
por: Kim, Changyeon, et al.
Publicado: (2025)
Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning
por: Koirala, Prajwal, et al.
Publicado: (2025)
por: Koirala, Prajwal, et al.
Publicado: (2025)
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
por: Zhang, Hongyin, et al.
Publicado: (2025)
por: Zhang, Hongyin, et al.
Publicado: (2025)
FLAG: Flow Policy MaxEnt-RL by Latent Augmented Guidance
por: Kim, Sungha, et al.
Publicado: (2026)
por: Kim, Sungha, et al.
Publicado: (2026)
Adaptive Q-Chunking for Offline-to-Online Reinforcement Learning
por: Gireesh, Nandiraju, et al.
Publicado: (2026)
por: Gireesh, Nandiraju, et al.
Publicado: (2026)
Uncertainty-Aware Rank-One MIMO Q Network Framework for Accelerated Offline Reinforcement Learning
por: Nguyen, Thanh, et al.
Publicado: (2026)
por: Nguyen, Thanh, et al.
Publicado: (2026)
Q-value Regularized Decision ConvFormer for Offline Reinforcement Learning
por: Yan, Teng, et al.
Publicado: (2024)
por: Yan, Teng, et al.
Publicado: (2024)
Learning Generalizable Visuomotor Policy through Dynamics-Alignment
por: Lee, Dohyeok, et al.
Publicado: (2025)
por: Lee, Dohyeok, et al.
Publicado: (2025)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
por: Xiao, Wei, et al.
Publicado: (2025)
por: Xiao, Wei, et al.
Publicado: (2025)
Diffusion Models as Optimizers for Efficient Planning in Offline RL
por: Huang, Renming, et al.
Publicado: (2024)
por: Huang, Renming, et al.
Publicado: (2024)
ENOTO: Improving Offline-to-Online Reinforcement Learning with Q-Ensembles
por: Zhao, Kai, et al.
Publicado: (2023)
por: Zhao, Kai, et al.
Publicado: (2023)
CtRL-Sim: Reactive and Controllable Driving Agents with Offline Reinforcement Learning
por: Rowe, Luke, et al.
Publicado: (2024)
por: Rowe, Luke, et al.
Publicado: (2024)
Compositional Conservatism: A Transductive Approach in Offline Reinforcement Learning
por: Song, Yeda, et al.
Publicado: (2024)
por: Song, Yeda, et al.
Publicado: (2024)
Robust Policy Learning via Offline Skill Diffusion
por: Kim, Woo Kyung, et al.
Publicado: (2024)
por: Kim, Woo Kyung, et al.
Publicado: (2024)
Language-Conditioned Offline RL for Multi-Robot Navigation
por: Morad, Steven, et al.
Publicado: (2024)
por: Morad, Steven, et al.
Publicado: (2024)
ACL-QL: Adaptive Conservative Level in Q-Learning for Offline Reinforcement Learning
por: Wu, Kun, et al.
Publicado: (2024)
por: Wu, Kun, et al.
Publicado: (2024)
Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking
por: Bagaria, Vaidehi, et al.
Publicado: (2026)
por: Bagaria, Vaidehi, et al.
Publicado: (2026)
HIQL: Offline Goal-Conditioned RL with Latent States as Actions
por: Park, Seohong, et al.
Publicado: (2023)
por: Park, Seohong, et al.
Publicado: (2023)
SutureFormer: Learning Surgical Trajectories via Goal-conditioned Offline RL in Pixel Space
por: Liu, Huanrong, et al.
Publicado: (2026)
por: Liu, Huanrong, et al.
Publicado: (2026)
Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation
por: Li, Huanyu, et al.
Publicado: (2026)
por: Li, Huanyu, et al.
Publicado: (2026)
Scaling Offline RL via Efficient and Expressive Shortcut Models
por: Espinosa-Dice, Nicolas, et al.
Publicado: (2025)
por: Espinosa-Dice, Nicolas, et al.
Publicado: (2025)
Sim-Anchored Learning for On-the-Fly Adaptation
por: Mabsout, Bassel El, et al.
Publicado: (2023)
por: Mabsout, Bassel El, et al.
Publicado: (2023)
A Recipe for Stable Offline Multi-agent Reinforcement Learning
por: Lee, Dongsu, et al.
Publicado: (2026)
por: Lee, Dongsu, et al.
Publicado: (2026)
Equivariant Offline Reinforcement Learning
por: Tangri, Arsh, et al.
Publicado: (2024)
por: Tangri, Arsh, et al.
Publicado: (2024)
Q-Guided Stein Variational Model Predictive Control via RL-informed Policy Prior
por: Cai, Shizhe, et al.
Publicado: (2025)
por: Cai, Shizhe, et al.
Publicado: (2025)
H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics Gaps
por: Niu, Haoyi, et al.
Publicado: (2023)
por: Niu, Haoyi, et al.
Publicado: (2023)
From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning
por: Sun, Zhanyi, et al.
Publicado: (2026)
por: Sun, Zhanyi, et al.
Publicado: (2026)
Dual-Granularity Contrastive Reward via Generated Episodic Guidance for Efficient Embodied RL
por: Liu, Xin, et al.
Publicado: (2026)
por: Liu, Xin, et al.
Publicado: (2026)
SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling
por: Zhang, Yixian, et al.
Publicado: (2025)
por: Zhang, Yixian, et al.
Publicado: (2025)
Boundary-to-Region Supervision for Offline Safe Reinforcement Learning
por: Su, Huikang, et al.
Publicado: (2025)
por: Su, Huikang, et al.
Publicado: (2025)
Graph-Assisted Stitching for Offline Hierarchical Reinforcement Learning
por: Baek, Seungho, et al.
Publicado: (2025)
por: Baek, Seungho, et al.
Publicado: (2025)
FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy
por: He, Qian, et al.
Publicado: (2026)
por: He, Qian, et al.
Publicado: (2026)
Towards Robust Policy: Enhancing Offline Reinforcement Learning with Adversarial Attacks and Defenses
por: Nguyen, Thanh, et al.
Publicado: (2024)
por: Nguyen, Thanh, et al.
Publicado: (2024)
Dataset Clustering for Improved Offline Policy Learning
por: Wang, Qiang, et al.
Publicado: (2024)
por: Wang, Qiang, et al.
Publicado: (2024)
Ejemplares similares
-
Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
por: Su, Zelal, et al.
Publicado: (2026) -
Evolutionary Policy Optimization
por: Mustafaoglu, Zelal Su "Lain", et al.
Publicado: (2025) -
ReFORM: Reflected Flows for On-support Offline RL via Noise Manipulation
por: Zhang, Songyuan, et al.
Publicado: (2026) -
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
por: You, Bozhi, et al.
Publicado: (2025) -
Causal Flow Q-Learning for Robust Offline Reinforcement Learning
por: Li, Mingxuan, et al.
Publicado: (2026)