DynaMITE-RL: A Dynamic Model for Improved Temporal Meta-Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Anthony, Tennenholtz, Guy, Hsu, Chih-wei, Chow, Yinlam, Bıyık, Erdem, Boutilier, Craig |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Embedding-Aligned Language Models
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2024)
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2024)
Descriptive History Representations: Learning Representations by Answering Questions
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2025)
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2025)
Preference Adaptive and Sequential Text-to-Image Generation
von: Nabati, Ofir, et al.
Veröffentlicht: (2024)
von: Nabati, Ofir, et al.
Veröffentlicht: (2024)
Demystifying Embedding Spaces using Large Language Models
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2023)
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2023)
ViSaRL: Visual Reinforcement Learning Guided by Human Saliency
von: Liang, Anthony, et al.
Veröffentlicht: (2024)
von: Liang, Anthony, et al.
Veröffentlicht: (2024)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
von: Chow, Yinlam, et al.
Veröffentlicht: (2024)
Synthetic Dialogue Generation for Interactive Conversational Elicitation & Recommendation (ICER)
von: Ryu, Moonkyung, et al.
Veröffentlicht: (2025)
von: Ryu, Moonkyung, et al.
Veröffentlicht: (2025)
Diffusion Controller: Framework, Algorithms and Parameterization
von: Yang, Tong, et al.
Veröffentlicht: (2026)
von: Yang, Tong, et al.
Veröffentlicht: (2026)
Spectral Souping: A Unified Framework for Online Preference Alignment
von: Chow, Yinlam, et al.
Veröffentlicht: (2026)
von: Chow, Yinlam, et al.
Veröffentlicht: (2026)
MILE: Model-based Intervention Learning
von: Korkmaz, Yigit, et al.
Veröffentlicht: (2025)
von: Korkmaz, Yigit, et al.
Veröffentlicht: (2025)
Asking Clarifying Questions for Preference Elicitation With Large Language Models
von: Montazeralghaem, Ali, et al.
Veröffentlicht: (2025)
von: Montazeralghaem, Ali, et al.
Veröffentlicht: (2025)
Controllable User Simulation
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2026)
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2026)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
von: Wang, Yufei, et al.
Veröffentlicht: (2024)
von: Wang, Yufei, et al.
Veröffentlicht: (2024)
Representation-Driven Reinforcement Learning
von: Nabati, Ofir, et al.
Veröffentlicht: (2023)
von: Nabati, Ofir, et al.
Veröffentlicht: (2023)
Benchmarks for Reinforcement Learning with Biased Offline Data and Imperfect Simulators
von: Linial, Ori, et al.
Veröffentlicht: (2024)
von: Linial, Ori, et al.
Veröffentlicht: (2024)
CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations
von: Liang, Anthony, et al.
Veröffentlicht: (2025)
von: Liang, Anthony, et al.
Veröffentlicht: (2025)
Spectral Bellman Method: Unifying Representation and Exploration in RL
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
von: Nabati, Ofir, et al.
Veröffentlicht: (2025)
Batch Active Learning of Reward Functions from Human Preferences
von: Bıyık, Erdem, et al.
Veröffentlicht: (2024)
von: Bıyık, Erdem, et al.
Veröffentlicht: (2024)
Coprocessor Actor Critic: A Model-Based Reinforcement Learning Approach For Adaptive Brain Stimulation
von: Pan, Michelle, et al.
Veröffentlicht: (2024)
von: Pan, Michelle, et al.
Veröffentlicht: (2024)
Training robots with natural and lightweight human feedback
von: Erdem Bıyık
Veröffentlicht: (2026)
von: Erdem Bıyık
Veröffentlicht: (2026)
GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2025)
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2025)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
von: Hwang, Minjune, et al.
Veröffentlicht: (2026)
von: Hwang, Minjune, et al.
Veröffentlicht: (2026)
Dyna-Style Reinforcement Learning Modeling and Control of Non-linear Dynamics
von: Abdelsalam, Karim, et al.
Veröffentlicht: (2025)
von: Abdelsalam, Karim, et al.
Veröffentlicht: (2025)
Value Explicit Pretraining for Learning Transferable Representations
von: Lekkala, Kiran, et al.
Veröffentlicht: (2023)
von: Lekkala, Kiran, et al.
Veröffentlicht: (2023)
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
von: Korkmaz, Yigit, et al.
Veröffentlicht: (2025)
von: Korkmaz, Yigit, et al.
Veröffentlicht: (2025)
Bayesian Regret Minimization in Offline Bandits
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
von: Li, Xinhu, et al.
Veröffentlicht: (2025)
von: Li, Xinhu, et al.
Veröffentlicht: (2025)
Multi-Agent Inverse Q-Learning from Demonstrations
von: Haynam, Nathaniel, et al.
Veröffentlicht: (2025)
von: Haynam, Nathaniel, et al.
Veröffentlicht: (2025)
IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models
von: Ling, Yiyang, et al.
Veröffentlicht: (2025)
von: Ling, Yiyang, et al.
Veröffentlicht: (2025)
RL$^3$: Boosting Meta Reinforcement Learning via RL inside RL$^2$
von: Bhatia, Abhinav, et al.
Veröffentlicht: (2023)
von: Bhatia, Abhinav, et al.
Veröffentlicht: (2023)
A Generalized Acquisition Function for Preference-based Reward Learning
von: Ellis, Evan, et al.
Veröffentlicht: (2024)
von: Ellis, Evan, et al.
Veröffentlicht: (2024)
DynaGen: Unifying Temporal Knowledge Graph Reasoning with Dynamic Subgraphs and Generative Regularization
von: Shen, Jiawei, et al.
Veröffentlicht: (2025)
von: Shen, Jiawei, et al.
Veröffentlicht: (2025)
Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
von: Zhang, Rongtao, et al.
Veröffentlicht: (2026)
von: Zhang, Rongtao, et al.
Veröffentlicht: (2026)
ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
von: Li, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Li, Zhaoyang, et al.
Veröffentlicht: (2025)
Knowledge Transfer in Deep Reinforcement Learning via an RL-Specific GAN-Based Correspondence Function
von: Ruman, Marko, et al.
Veröffentlicht: (2022)
von: Ruman, Marko, et al.
Veröffentlicht: (2022)
Meta-Gradient Search Control: A Method for Improving the Efficiency of Dyna-style Planning
von: Burega, Bradley, et al.
Veröffentlicht: (2024)
von: Burega, Bradley, et al.
Veröffentlicht: (2024)
Physics-Informed Model and Hybrid Planning for Efficient Dyna-Style Reinforcement Learning
von: Asri, Zakariae El, et al.
Veröffentlicht: (2024)
von: Asri, Zakariae El, et al.
Veröffentlicht: (2024)
Minimizing Live Experiments in Recommender Systems: User Simulation to Evaluate Preference Elicitation Policies
von: Hsu, Chih-Wei, et al.
Veröffentlicht: (2024)
von: Hsu, Chih-Wei, et al.
Veröffentlicht: (2024)
ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders
von: Meshi, Ofer, et al.
Veröffentlicht: (2026)
von: Meshi, Ofer, et al.
Veröffentlicht: (2026)
DynaSTy: A Framework for SpatioTemporal Node Attribute Prediction in Dynamic Graphs
von: Banerji, Namrata, et al.
Veröffentlicht: (2026)
von: Banerji, Namrata, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Embedding-Aligned Language Models
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2024) -
Descriptive History Representations: Learning Representations by Answering Questions
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2025) -
Preference Adaptive and Sequential Text-to-Image Generation
von: Nabati, Ofir, et al.
Veröffentlicht: (2024) -
Demystifying Embedding Spaces using Large Language Models
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2023) -
ViSaRL: Visual Reinforcement Learning Guided by Human Saliency
von: Liang, Anthony, et al.
Veröffentlicht: (2024)