Transitive RL: Value Learning via Divide and Conquer
Fuente:
arXiv
Salvato in:
| Autori principali: | Park, Seohong, Oberai, Aditya, Atreya, Pranav, Levine, Sergey |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Is Value Learning Really the Main Bottleneck in Offline RL?
di: Park, Seohong, et al.
Pubblicazione: (2024)
di: Park, Seohong, et al.
Pubblicazione: (2024)
METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
di: Park, Seohong, et al.
Pubblicazione: (2023)
di: Park, Seohong, et al.
Pubblicazione: (2023)
Scalable Offline Model-Based RL with Action Chunks
di: Park, Kwanyoung, et al.
Pubblicazione: (2025)
di: Park, Kwanyoung, et al.
Pubblicazione: (2025)
OGBench: Benchmarking Offline Goal-Conditioned RL
di: Park, Seohong, et al.
Pubblicazione: (2024)
di: Park, Seohong, et al.
Pubblicazione: (2024)
Flow Q-Learning
di: Park, Seohong, et al.
Pubblicazione: (2025)
di: Park, Seohong, et al.
Pubblicazione: (2025)
HIQL: Offline Goal-Conditioned RL with Latent States as Actions
di: Park, Seohong, et al.
Pubblicazione: (2023)
di: Park, Seohong, et al.
Pubblicazione: (2023)
Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings
di: Frans, Kevin, et al.
Pubblicazione: (2024)
di: Frans, Kevin, et al.
Pubblicazione: (2024)
Dual Goal Representations
di: Park, Seohong, et al.
Pubblicazione: (2025)
di: Park, Seohong, et al.
Pubblicazione: (2025)
Horizon Reduction Makes RL Scalable
di: Park, Seohong, et al.
Pubblicazione: (2025)
di: Park, Seohong, et al.
Pubblicazione: (2025)
Decoupled Q-Chunking
di: Li, Qiyang, et al.
Pubblicazione: (2025)
di: Li, Qiyang, et al.
Pubblicazione: (2025)
Foundation Policies with Hilbert Representations
di: Park, Seohong, et al.
Pubblicazione: (2024)
di: Park, Seohong, et al.
Pubblicazione: (2024)
Intention-Conditioned Flow Occupancy Models
di: Zheng, Chongyi, et al.
Pubblicazione: (2025)
di: Zheng, Chongyi, et al.
Pubblicazione: (2025)
ViVa: Video-Trained Value Functions for Guiding Online RL from Diverse Data
di: Dashora, Nitish, et al.
Pubblicazione: (2025)
di: Dashora, Nitish, et al.
Pubblicazione: (2025)
Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning
di: Zhang, Weiliang, et al.
Pubblicazione: (2025)
di: Zhang, Weiliang, et al.
Pubblicazione: (2025)
Recursive Decomposition with Dependencies for Generic Divide-and-Conquer Reasoning
di: Hernández-Gutiérrez, Sergio, et al.
Pubblicazione: (2025)
di: Hernández-Gutiérrez, Sergio, et al.
Pubblicazione: (2025)
Unsupervised-to-Online Reinforcement Learning
di: Kim, Junsu, et al.
Pubblicazione: (2024)
di: Kim, Junsu, et al.
Pubblicazione: (2024)
RACER: Epistemic Risk-Sensitive RL Enables Fast Driving with Fewer Crashes
di: Stachowicz, Kyle, et al.
Pubblicazione: (2024)
di: Stachowicz, Kyle, et al.
Pubblicazione: (2024)
FPGA Divide-and-Conquer Placement using Deep Reinforcement Learning
di: Wang, Shang, et al.
Pubblicazione: (2024)
di: Wang, Shang, et al.
Pubblicazione: (2024)
Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
di: Farebrother, Jesse, et al.
Pubblicazione: (2024)
di: Farebrother, Jesse, et al.
Pubblicazione: (2024)
Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning
di: Huang, Chengsong, et al.
Pubblicazione: (2024)
di: Huang, Chengsong, et al.
Pubblicazione: (2024)
Divide And Conquer: Learning Chaotic Dynamical Systems With Multistep Penalty Neural Ordinary Differential Equations
di: Chakraborty, Dibyajyoti, et al.
Pubblicazione: (2024)
di: Chakraborty, Dibyajyoti, et al.
Pubblicazione: (2024)
PLM-eXplain: Divide and Conquer the Protein Embedding Space
di: van Eck, Jan, et al.
Pubblicazione: (2025)
di: van Eck, Jan, et al.
Pubblicazione: (2025)
An Examination on the Effectiveness of Divide-and-Conquer Prompting in Large Language Models
di: Zhang, Yizhou, et al.
Pubblicazione: (2024)
di: Zhang, Yizhou, et al.
Pubblicazione: (2024)
Divide-and-Conquer Predictive Coding: a structured Bayesian inference algorithm
di: Sennesh, Eli, et al.
Pubblicazione: (2024)
di: Sennesh, Eli, et al.
Pubblicazione: (2024)
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
di: Zhou, Yifei, et al.
Pubblicazione: (2024)
di: Zhou, Yifei, et al.
Pubblicazione: (2024)
Divide and Conquer Self-Supervised Learning for High-Content Imaging
di: Farndale, Lucas, et al.
Pubblicazione: (2025)
di: Farndale, Lucas, et al.
Pubblicazione: (2025)
Divide (Text) and Conquer (Sentiment): Improved Sentiment Classification by Constituent Conflict Resolution
di: Kościałkowski, Jan, et al.
Pubblicazione: (2025)
di: Kościałkowski, Jan, et al.
Pubblicazione: (2025)
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
di: Wagenmaker, Andrew, et al.
Pubblicazione: (2025)
di: Wagenmaker, Andrew, et al.
Pubblicazione: (2025)
Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing
di: Wahréus, Johan, et al.
Pubblicazione: (2025)
di: Wahréus, Johan, et al.
Pubblicazione: (2025)
Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data
di: Zheng, Chongyi, et al.
Pubblicazione: (2023)
di: Zheng, Chongyi, et al.
Pubblicazione: (2023)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
di: Hong, Joey, et al.
Pubblicazione: (2024)
di: Hong, Joey, et al.
Pubblicazione: (2024)
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
di: Xu, Charles, et al.
Pubblicazione: (2024)
di: Xu, Charles, et al.
Pubblicazione: (2024)
Toward Effective Multimodal Graph Foundation Model: A Divide-and-Conquer Based Approach
di: Liu, Sicheng, et al.
Pubblicazione: (2026)
di: Liu, Sicheng, et al.
Pubblicazione: (2026)
Reinforcement Learning with Action Chunking
di: Li, Qiyang, et al.
Pubblicazione: (2025)
di: Li, Qiyang, et al.
Pubblicazione: (2025)
Diffusion Guidance Is a Controllable Policy Improvement Operator
di: Frans, Kevin, et al.
Pubblicazione: (2025)
di: Frans, Kevin, et al.
Pubblicazione: (2025)
Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning
di: Nakamoto, Mitsuhiko, et al.
Pubblicazione: (2023)
di: Nakamoto, Mitsuhiko, et al.
Pubblicazione: (2023)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
di: Hong, Joey, et al.
Pubblicazione: (2024)
di: Hong, Joey, et al.
Pubblicazione: (2024)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
di: Abdulhai, Marwa, et al.
Pubblicazione: (2025)
di: Abdulhai, Marwa, et al.
Pubblicazione: (2025)
RL$^3$: Boosting Meta Reinforcement Learning via RL inside RL$^2$
di: Bhatia, Abhinav, et al.
Pubblicazione: (2023)
di: Bhatia, Abhinav, et al.
Pubblicazione: (2023)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
di: Banerjee, Debangshu, et al.
Pubblicazione: (2023)
di: Banerjee, Debangshu, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Is Value Learning Really the Main Bottleneck in Offline RL?
di: Park, Seohong, et al.
Pubblicazione: (2024) -
METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
di: Park, Seohong, et al.
Pubblicazione: (2023) -
Scalable Offline Model-Based RL with Action Chunks
di: Park, Kwanyoung, et al.
Pubblicazione: (2025) -
OGBench: Benchmarking Offline Goal-Conditioned RL
di: Park, Seohong, et al.
Pubblicazione: (2024) -
Flow Q-Learning
di: Park, Seohong, et al.
Pubblicazione: (2025)