Diffusing States and Matching Scores: A New Framework for Imitation Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Runzhe, Chen, Yiding, Swamy, Gokul, Brantley, Kianté, Sun, Wen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Offline RL via Efficient and Expressive Shortcut Models
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
Adversarial Imitation Learning via Boosting
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
Expressive Value Learning for Scalable Offline Reinforcement Learning
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
Efficient Imitation under Misspecification
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
Multi-Agent Imitation Learning: Value is Easy, Regret is Hard
von: Tang, Jingwu, et al.
Veröffentlicht: (2024)
von: Tang, Jingwu, et al.
Veröffentlicht: (2024)
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
von: Wu, Anne, et al.
Veröffentlicht: (2024)
von: Wu, Anne, et al.
Veröffentlicht: (2024)
LLMs Can Learn to Reason Via Off-Policy RL
von: Ritter, Daniel, et al.
Veröffentlicht: (2026)
von: Ritter, Daniel, et al.
Veröffentlicht: (2026)
REBEL: Reinforcement Learning via Regressing Relative Rewards
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
$p1$: Better Prompt Optimization with Fewer Prompts
von: Gao, Zhaolin, et al.
Veröffentlicht: (2026)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2026)
EvIL: Evolution Strategies for Generalisable Imitation Learning
von: Sapora, Silvia, et al.
Veröffentlicht: (2024)
von: Sapora, Silvia, et al.
Veröffentlicht: (2024)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
von: Oertell, Owen, et al.
Veröffentlicht: (2024)
von: Oertell, Owen, et al.
Veröffentlicht: (2024)
LLMs Are In-Context Bandit Reinforcement Learners
von: Monea, Giovanni, et al.
Veröffentlicht: (2024)
von: Monea, Giovanni, et al.
Veröffentlicht: (2024)
All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning
von: Swamy, Gokul, et al.
Veröffentlicht: (2025)
von: Swamy, Gokul, et al.
Veröffentlicht: (2025)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
von: Brantley, Kianté, et al.
Veröffentlicht: (2025)
von: Brantley, Kianté, et al.
Veröffentlicht: (2025)
A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to Search
von: Jain, Arnav Kumar, et al.
Veröffentlicht: (2025)
von: Jain, Arnav Kumar, et al.
Veröffentlicht: (2025)
Inverse Reinforcement Learning without Reinforcement Learning
von: Swamy, Gokul, et al.
Veröffentlicht: (2023)
von: Swamy, Gokul, et al.
Veröffentlicht: (2023)
Making RL with Preference-based Feedback Efficient via Randomization
von: Wu, Runzhe, et al.
Veröffentlicht: (2023)
von: Wu, Runzhe, et al.
Veröffentlicht: (2023)
Distributional Offline Policy Evaluation with Predictive Error Guarantees
von: Wu, Runzhe, et al.
Veröffentlicht: (2023)
von: Wu, Runzhe, et al.
Veröffentlicht: (2023)
Imitation Learning by State-Only Distribution Matching
von: Boborzi, Damian, et al.
Veröffentlicht: (2022)
von: Boborzi, Damian, et al.
Veröffentlicht: (2022)
A Minimaximalist Approach to Reinforcement Learning from Human Feedback
von: Swamy, Gokul, et al.
Veröffentlicht: (2024)
von: Swamy, Gokul, et al.
Veröffentlicht: (2024)
Value-Guided Search for Efficient Chain-of-Thought Reasoning
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
The Virtues of Pessimism in Inverse Reinforcement Learning
von: Wu, David, et al.
Veröffentlicht: (2024)
von: Wu, David, et al.
Veröffentlicht: (2024)
From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment
von: Wu, Yilin, et al.
Veröffentlicht: (2025)
von: Wu, Yilin, et al.
Veröffentlicht: (2025)
Dataset Reset Policy Optimization for RLHF
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024)
Policy-Gradient Training of Language Models for Ranking
von: Gao, Ge, et al.
Veröffentlicht: (2023)
von: Gao, Ge, et al.
Veröffentlicht: (2023)
Scaling Reward Modeling without Human Supervision
von: Fan, Jingxuan, et al.
Veröffentlicht: (2026)
von: Fan, Jingxuan, et al.
Veröffentlicht: (2026)
Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion
von: Zhang, Ruixiang, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixiang, et al.
Veröffentlicht: (2025)
The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
von: Song, Yuda, et al.
Veröffentlicht: (2024)
von: Song, Yuda, et al.
Veröffentlicht: (2024)
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
von: Sutawika, Lintang, et al.
Veröffentlicht: (2026)
von: Sutawika, Lintang, et al.
Veröffentlicht: (2026)
Hybrid Inverse Reinforcement Learning
von: Ren, Juntao, et al.
Veröffentlicht: (2024)
von: Ren, Juntao, et al.
Veröffentlicht: (2024)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2025)
Physics-informed Imitative Reinforcement Learning for Real-world Driving
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
Your Learned Constraint is Secretly a Backward Reachable Tube
von: Qadri, Mohamad, et al.
Veröffentlicht: (2025)
von: Qadri, Mohamad, et al.
Veröffentlicht: (2025)
Diffusion Imitation from Observation
von: Huang, Bo-Ruei, et al.
Veröffentlicht: (2024)
von: Huang, Bo-Ruei, et al.
Veröffentlicht: (2024)
Computationally Efficient RL under Linear Bellman Completeness for Deterministic Dynamics
von: Wu, Runzhe, et al.
Veröffentlicht: (2024)
von: Wu, Runzhe, et al.
Veröffentlicht: (2024)
Simulating Diffusion Bridges with Score Matching
von: Heng, Jeremy, et al.
Veröffentlicht: (2021)
von: Heng, Jeremy, et al.
Veröffentlicht: (2021)
Imitation Learning as Return Distribution Matching
von: Lazzati, Filippo, et al.
Veröffentlicht: (2025)
von: Lazzati, Filippo, et al.
Veröffentlicht: (2025)
Diffusion-Reward Adversarial Imitation Learning
von: Lai, Chun-Mao, et al.
Veröffentlicht: (2024)
von: Lai, Chun-Mao, et al.
Veröffentlicht: (2024)
Control Variate Score Matching for Diffusion Models
von: Kahouli, Khaled, et al.
Veröffentlicht: (2025)
von: Kahouli, Khaled, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling Offline RL via Efficient and Expressive Shortcut Models
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025) -
Adversarial Imitation Learning via Boosting
von: Chang, Jonathan D., et al.
Veröffentlicht: (2024) -
Expressive Value Learning for Scalable Offline Reinforcement Learning
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025) -
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024) -
Efficient Imitation under Misspecification
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)