Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Hao, van der Schaar, Mihaela |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
von: Sun, Hao, et al.
Veröffentlicht: (2025)
von: Sun, Hao, et al.
Veröffentlicht: (2025)
Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RL
von: Sun, Hao, et al.
Veröffentlicht: (2023)
von: Sun, Hao, et al.
Veröffentlicht: (2023)
Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
von: Fanconi, Claudio, et al.
Veröffentlicht: (2025)
von: Fanconi, Claudio, et al.
Veröffentlicht: (2025)
Preference Learning for AI Alignment: a Causal Perspective
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2025)
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2025)
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
von: Sun, Hao, et al.
Veröffentlicht: (2025)
von: Sun, Hao, et al.
Veröffentlicht: (2025)
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
von: Chiu, Christopher, et al.
Veröffentlicht: (2025)
von: Chiu, Christopher, et al.
Veröffentlicht: (2025)
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2024)
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2024)
Language Bottleneck Models for Qualitative Knowledge State Modeling
von: Berthon, Antonin, et al.
Veröffentlicht: (2025)
von: Berthon, Antonin, et al.
Veröffentlicht: (2025)
Large Language Models to Enhance Bayesian Optimization
von: Liu, Tennison, et al.
Veröffentlicht: (2024)
von: Liu, Tennison, et al.
Veröffentlicht: (2024)
Not All Explanations for Deep Learning Phenomena Are Equally Valuable
von: Jeffares, Alan, et al.
Veröffentlicht: (2025)
von: Jeffares, Alan, et al.
Veröffentlicht: (2025)
Active Timepoint Selection for Learning Measure-Valued Trajectories
von: Huynh, Nicolas, et al.
Veröffentlicht: (2026)
von: Huynh, Nicolas, et al.
Veröffentlicht: (2026)
Hyperparameter Trajectory Inference with Conditional Lagrangian Optimal Transport
von: Amad, Harry, et al.
Veröffentlicht: (2026)
von: Amad, Harry, et al.
Veröffentlicht: (2026)
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
von: Jeffares, Alan, et al.
Veröffentlicht: (2024)
von: Jeffares, Alan, et al.
Veröffentlicht: (2024)
Inverse Reinforcement Learning by Estimating Expertise of Demonstrators
von: Beliaev, Mark, et al.
Veröffentlicht: (2024)
von: Beliaev, Mark, et al.
Veröffentlicht: (2024)
Discovery of Hidden Miscalibration Regimes
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2026)
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2026)
Causal Deep Learning
von: Berrevoets, Jeroen, et al.
Veröffentlicht: (2023)
von: Berrevoets, Jeroen, et al.
Veröffentlicht: (2023)
Eliciting Numerical Predictive Distributions of LLMs Without Autoregression
von: Piskorz, Julianna, et al.
Veröffentlicht: (2026)
von: Piskorz, Julianna, et al.
Veröffentlicht: (2026)
Rethinking Inverse Reinforcement Learning: from Data Alignment to Task Alignment
von: Zhou, Weichao, et al.
Veröffentlicht: (2024)
von: Zhou, Weichao, et al.
Veröffentlicht: (2024)
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
von: Sun, Hao, et al.
Veröffentlicht: (2023)
von: Sun, Hao, et al.
Veröffentlicht: (2023)
Supervised Fine-Tuning as Inverse Reinforcement Learning
von: Sun, Hao
Veröffentlicht: (2024)
von: Sun, Hao
Veröffentlicht: (2024)
Self-Healing Machine Learning: A Framework for Autonomous Adaptation in Real-World Environments
von: Rauba, Paulius, et al.
Veröffentlicht: (2024)
von: Rauba, Paulius, et al.
Veröffentlicht: (2024)
AutoPrognosis 2.0: Democratizing Diagnostic and Prognostic Modeling in Healthcare with Automated Machine Learning
von: Imrie, Fergus, et al.
Veröffentlicht: (2022)
von: Imrie, Fergus, et al.
Veröffentlicht: (2022)
Machine Learning with Requirements: a Manifesto
von: Giunchiglia, Eleonora, et al.
Veröffentlicht: (2023)
von: Giunchiglia, Eleonora, et al.
Veröffentlicht: (2023)
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
von: Seedat, Nabeel, et al.
Veröffentlicht: (2023)
von: Seedat, Nabeel, et al.
Veröffentlicht: (2023)
You can't handle the (dirty) truth: Data-centric insights improve pseudo-labeling
von: Seedat, Nabeel, et al.
Veröffentlicht: (2024)
von: Seedat, Nabeel, et al.
Veröffentlicht: (2024)
Time Series Diffusion in the Frequency Domain
von: Crabbé, Jonathan, et al.
Veröffentlicht: (2024)
von: Crabbé, Jonathan, et al.
Veröffentlicht: (2024)
Meta-Learners for Partially-Identified Treatment Effects Across Multiple Environments
von: Schweisthal, Jonas, et al.
Veröffentlicht: (2024)
von: Schweisthal, Jonas, et al.
Veröffentlicht: (2024)
Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach
von: Kim, Kihyun, et al.
Veröffentlicht: (2026)
von: Kim, Kihyun, et al.
Veröffentlicht: (2026)
Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
von: Rauba, Paulius, et al.
Veröffentlicht: (2025)
von: Rauba, Paulius, et al.
Veröffentlicht: (2025)
Hybrid Inverse Reinforcement Learning
von: Ren, Juntao, et al.
Veröffentlicht: (2024)
von: Ren, Juntao, et al.
Veröffentlicht: (2024)
GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
von: Sapora, Silvia, et al.
Veröffentlicht: (2025)
von: Sapora, Silvia, et al.
Veröffentlicht: (2025)
When Demonstrations Meet Generative World Models: A Maximum Likelihood Framework for Offline Inverse Reinforcement Learning
von: Zeng, Siliang, et al.
Veröffentlicht: (2023)
von: Zeng, Siliang, et al.
Veröffentlicht: (2023)
Defining Expertise: Applications to Treatment Effect Estimation
von: Hüyük, Alihan, et al.
Veröffentlicht: (2024)
von: Hüyük, Alihan, et al.
Veröffentlicht: (2024)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2025)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2025)
OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
von: Sun, Hao, et al.
Veröffentlicht: (2025)
von: Sun, Hao, et al.
Veröffentlicht: (2025)
DC-Check: A Data-Centric AI checklist to guide the development of reliable machine learning systems
von: Seedat, Nabeel, et al.
Veröffentlicht: (2022)
von: Seedat, Nabeel, et al.
Veröffentlicht: (2022)
Fast Rates for Inverse Reinforcement Learning
von: Schlaginhaufen, Andreas, et al.
Veröffentlicht: (2026)
von: Schlaginhaufen, Andreas, et al.
Veröffentlicht: (2026)
On the Effective Horizon of Inverse Reinforcement Learning
von: Xu, Yiqing, et al.
Veröffentlicht: (2023)
von: Xu, Yiqing, et al.
Veröffentlicht: (2023)
Environment Design for Inverse Reinforcement Learning
von: Buening, Thomas Kleine, et al.
Veröffentlicht: (2022)
von: Buening, Thomas Kleine, et al.
Veröffentlicht: (2022)
Recursive Deep Inverse Reinforcement Learning
von: Ghanem, Paul, et al.
Veröffentlicht: (2025)
von: Ghanem, Paul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
von: Sun, Hao, et al.
Veröffentlicht: (2025) -
Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RL
von: Sun, Hao, et al.
Veröffentlicht: (2023) -
Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
von: Fanconi, Claudio, et al.
Veröffentlicht: (2025) -
Preference Learning for AI Alignment: a Causal Perspective
von: Kobalczyk, Katarzyna, et al.
Veröffentlicht: (2025) -
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
von: Sun, Hao, et al.
Veröffentlicht: (2025)