Decision Flow Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Jifeng, Huang, Sili, Guo, Siyuan, Liu, Zhaogeng, Shen, Li, Sun, Lichao, Chen, Hechang, Chang, Yi, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning
by: Hu, Jifeng, et al.
Published: (2025)
by: Hu, Jifeng, et al.
Published: (2025)
In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought
by: Huang, Sili, et al.
Published: (2024)
by: Huang, Sili, et al.
Published: (2024)
Continual Diffuser (CoD): Mastering Continual Offline Reinforcement Learning with Experience Rehearsal
by: Hu, Jifeng, et al.
Published: (2024)
by: Hu, Jifeng, et al.
Published: (2024)
A Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement Learning
by: Guo, Siyuan, et al.
Published: (2023)
by: Guo, Siyuan, et al.
Published: (2023)
Solving Continual Offline RL through Selective Weights Activation on Aligned Spaces
by: Hu, Jifeng, et al.
Published: (2024)
by: Hu, Jifeng, et al.
Published: (2024)
Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling
by: Huang, Sili, et al.
Published: (2024)
by: Huang, Sili, et al.
Published: (2024)
Continual Task Learning through Adaptive Policy Self-Composition
by: Hu, Shengchao, et al.
Published: (2024)
by: Hu, Shengchao, et al.
Published: (2024)
CASCADE: Case-Based Continual Adaptation for Large Language Models During Deployment
by: Guo, Siyuan, et al.
Published: (2026)
by: Guo, Siyuan, et al.
Published: (2026)
Flow marching for a generative PDE foundation model
by: Chen, Zituo, et al.
Published: (2025)
by: Chen, Zituo, et al.
Published: (2025)
Solving Continual Offline Reinforcement Learning with Decision Transformer
by: Huang, Kaixin, et al.
Published: (2024)
by: Huang, Kaixin, et al.
Published: (2024)
A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization
by: Lei, Shiye, et al.
Published: (2026)
by: Lei, Shiye, et al.
Published: (2026)
Policy Dispersion in Non-Markovian Environment
by: Qu, Bohao, et al.
Published: (2023)
by: Qu, Bohao, et al.
Published: (2023)
Prompt Tuning with Diffusion for Few-Shot Pre-trained Policy Generalization
by: Hu, Shengchao, et al.
Published: (2024)
by: Hu, Shengchao, et al.
Published: (2024)
Task-Aware Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning
by: Fan, Ziqing, et al.
Published: (2024)
by: Fan, Ziqing, et al.
Published: (2024)
Latent Generative Solvers for Generalizable Long-Term Physics Simulation
by: Chen, Zituo, et al.
Published: (2026)
by: Chen, Zituo, et al.
Published: (2026)
Reparameterization Flow Policy Optimization
by: Zhong, Hai, et al.
Published: (2026)
by: Zhong, Hai, et al.
Published: (2026)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
by: Ren, Zhiyao, et al.
Published: (2025)
by: Ren, Zhiyao, et al.
Published: (2025)
Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision Transformer
by: Kong, Yilun, et al.
Published: (2025)
by: Kong, Yilun, et al.
Published: (2025)
Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data Scheduler
by: Hu, Zixuan, et al.
Published: (2025)
by: Hu, Zixuan, et al.
Published: (2025)
AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies
by: Hu, Xixi, et al.
Published: (2024)
by: Hu, Xixi, et al.
Published: (2024)
Flow-Based Policy for Online Reinforcement Learning
by: Lv, Lei, et al.
Published: (2025)
by: Lv, Lei, et al.
Published: (2025)
Discrete Flow Matching Policy Optimization
by: Su, Maojiang, et al.
Published: (2026)
by: Su, Maojiang, et al.
Published: (2026)
Poisson Process for Bayesian Optimization
by: Wang, Xiaoxing, et al.
Published: (2024)
by: Wang, Xiaoxing, et al.
Published: (2024)
LLM Data Selection and Utilization via Dynamic Bi-level Optimization
by: Yu, Yang, et al.
Published: (2025)
by: Yu, Yang, et al.
Published: (2025)
SPOT: Scalable Policy Optimization with Trees for Markov Decision Processes
by: Xiong, Xuyuan, et al.
Published: (2025)
by: Xiong, Xuyuan, et al.
Published: (2025)
Absolute Policy Optimization
by: Zhao, Weiye, et al.
Published: (2023)
by: Zhao, Weiye, et al.
Published: (2023)
Online Policy Distillation with Decision-Attention
by: Yu, Xinqiang, et al.
Published: (2024)
by: Yu, Xinqiang, et al.
Published: (2024)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
Neuron-level Balance between Stability and Plasticity in Deep Reinforcement Learning
by: Lan, Jiahua, et al.
Published: (2025)
by: Lan, Jiahua, et al.
Published: (2025)
Fractal Landscapes in Policy Optimization
by: Wang, Tao, et al.
Published: (2023)
by: Wang, Tao, et al.
Published: (2023)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
CoVeR: Conformal Calibration for Versatile and Reliable Autoregressive Next-Token Prediction
by: Chen, Yuzhu, et al.
Published: (2025)
by: Chen, Yuzhu, et al.
Published: (2025)
Task-Distributionally Robust Data-Free Meta-Learning
by: Hu, Zixuan, et al.
Published: (2023)
by: Hu, Zixuan, et al.
Published: (2023)
A Theoretical Survey on Foundation Models
by: Fu, Shi, et al.
Published: (2024)
by: Fu, Shi, et al.
Published: (2024)
Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
by: Hao, Zhiwei, et al.
Published: (2025)
by: Hao, Zhiwei, et al.
Published: (2025)
Offline Behavior Distillation
by: Lei, Shiye, et al.
Published: (2024)
by: Lei, Shiye, et al.
Published: (2024)
Continual Learning on Graphs: Challenges, Solutions, and Opportunities
by: Zhang, Xikun, et al.
Published: (2024)
by: Zhang, Xikun, et al.
Published: (2024)
Offline Behavioral Data Selection
by: Lei, Shiye, et al.
Published: (2025)
by: Lei, Shiye, et al.
Published: (2025)
TimeGuard: Channel-wise Pool Training for Backdoor Defense in Time Series Forecasting
by: Nguyen, Quang Duc, et al.
Published: (2026)
by: Nguyen, Quang Duc, et al.
Published: (2026)
TreeRPO: Tree Relative Policy Optimization
by: Yang, Zhicheng, et al.
Published: (2025)
by: Yang, Zhicheng, et al.
Published: (2025)
Similar Items
-
Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning
by: Hu, Jifeng, et al.
Published: (2025) -
In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought
by: Huang, Sili, et al.
Published: (2024) -
Continual Diffuser (CoD): Mastering Continual Offline Reinforcement Learning with Experience Rehearsal
by: Hu, Jifeng, et al.
Published: (2024) -
A Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement Learning
by: Guo, Siyuan, et al.
Published: (2023) -
Solving Continual Offline RL through Selective Weights Activation on Aligned Spaces
by: Hu, Jifeng, et al.
Published: (2024)