Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Fanconi, Claudio, Astorga, Nicolás, van der Schaar, Mihaela |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Cascaded Language Models for Cost-effective Human-AI Decision-Making
por: Fanconi, Claudio, et al.
Publicado: (2025)
por: Fanconi, Claudio, et al.
Publicado: (2025)
Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
por: Sun, Hao, et al.
Publicado: (2024)
por: Sun, Hao, et al.
Publicado: (2024)
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
por: Kobalczyk, Katarzyna, et al.
Publicado: (2024)
por: Kobalczyk, Katarzyna, et al.
Publicado: (2024)
CauSim: Scaling Causal Reasoning with Increasingly Complex Causal Simulators
por: Astorga, Nicolás, et al.
Publicado: (2026)
por: Astorga, Nicolás, et al.
Publicado: (2026)
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
por: Sun, Hao, et al.
Publicado: (2025)
por: Sun, Hao, et al.
Publicado: (2025)
Timely Clinical Diagnosis through Active Test Selection
por: Estévez, Silas Ruhrberg, et al.
Publicado: (2025)
por: Estévez, Silas Ruhrberg, et al.
Publicado: (2025)
Active Timepoint Selection for Learning Measure-Valued Trajectories
por: Huynh, Nicolas, et al.
Publicado: (2026)
por: Huynh, Nicolas, et al.
Publicado: (2026)
Large Language Models to Enhance Bayesian Optimization
por: Liu, Tennison, et al.
Publicado: (2024)
por: Liu, Tennison, et al.
Publicado: (2024)
Truly Self-Improving Agents Require Intrinsic Metacognitive Learning
por: Liu, Tennison, et al.
Publicado: (2025)
por: Liu, Tennison, et al.
Publicado: (2025)
Active Task Disambiguation with LLMs
por: Kobalczyk, Katarzyna, et al.
Publicado: (2025)
por: Kobalczyk, Katarzyna, et al.
Publicado: (2025)
Not All Explanations for Deep Learning Phenomena Are Equally Valuable
por: Jeffares, Alan, et al.
Publicado: (2025)
por: Jeffares, Alan, et al.
Publicado: (2025)
Preference Learning for AI Alignment: a Causal Perspective
por: Kobalczyk, Katarzyna, et al.
Publicado: (2025)
por: Kobalczyk, Katarzyna, et al.
Publicado: (2025)
Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets
por: Zhang, Simpson, et al.
Publicado: (2025)
por: Zhang, Simpson, et al.
Publicado: (2025)
Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RL
por: Sun, Hao, et al.
Publicado: (2023)
por: Sun, Hao, et al.
Publicado: (2023)
Tiny Autoregressive Recursive Models
por: Rauba, Paulius, et al.
Publicado: (2026)
por: Rauba, Paulius, et al.
Publicado: (2026)
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
por: Chiu, Christopher, et al.
Publicado: (2025)
por: Chiu, Christopher, et al.
Publicado: (2025)
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
por: Jeffares, Alan, et al.
Publicado: (2024)
por: Jeffares, Alan, et al.
Publicado: (2024)
Hyperparameter Trajectory Inference with Conditional Lagrangian Optimal Transport
por: Amad, Harry, et al.
Publicado: (2026)
por: Amad, Harry, et al.
Publicado: (2026)
Causal Deep Learning
por: Berrevoets, Jeroen, et al.
Publicado: (2023)
por: Berrevoets, Jeroen, et al.
Publicado: (2023)
Discovery of Hidden Miscalibration Regimes
por: Kobalczyk, Katarzyna, et al.
Publicado: (2026)
por: Kobalczyk, Katarzyna, et al.
Publicado: (2026)
Language Bottleneck Models for Qualitative Knowledge State Modeling
por: Berthon, Antonin, et al.
Publicado: (2025)
por: Berthon, Antonin, et al.
Publicado: (2025)
Self-Healing Machine Learning: A Framework for Autonomous Adaptation in Real-World Environments
por: Rauba, Paulius, et al.
Publicado: (2024)
por: Rauba, Paulius, et al.
Publicado: (2024)
Machine Learning with Requirements: a Manifesto
por: Giunchiglia, Eleonora, et al.
Publicado: (2023)
por: Giunchiglia, Eleonora, et al.
Publicado: (2023)
The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data
por: Pouplin, Thomas, et al.
Publicado: (2024)
por: Pouplin, Thomas, et al.
Publicado: (2024)
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
por: Sun, Hao, et al.
Publicado: (2025)
por: Sun, Hao, et al.
Publicado: (2025)
You can't handle the (dirty) truth: Data-centric insights improve pseudo-labeling
por: Seedat, Nabeel, et al.
Publicado: (2024)
por: Seedat, Nabeel, et al.
Publicado: (2024)
Time Series Diffusion in the Frequency Domain
por: Crabbé, Jonathan, et al.
Publicado: (2024)
por: Crabbé, Jonathan, et al.
Publicado: (2024)
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
por: Seedat, Nabeel, et al.
Publicado: (2023)
por: Seedat, Nabeel, et al.
Publicado: (2023)
Strategic Self-Improvement for Competitive Agents in AI Labour Markets
por: Chiu, Christopher, et al.
Publicado: (2025)
por: Chiu, Christopher, et al.
Publicado: (2025)
OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
por: Sun, Hao, et al.
Publicado: (2025)
por: Sun, Hao, et al.
Publicado: (2025)
Eliciting Numerical Predictive Distributions of LLMs Without Autoregression
por: Piskorz, Julianna, et al.
Publicado: (2026)
por: Piskorz, Julianna, et al.
Publicado: (2026)
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
por: Sun, Hao, et al.
Publicado: (2023)
por: Sun, Hao, et al.
Publicado: (2023)
Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach
por: Kim, Kihyun, et al.
Publicado: (2026)
por: Kim, Kihyun, et al.
Publicado: (2026)
AutoPrognosis 2.0: Democratizing Diagnostic and Prognostic Modeling in Healthcare with Automated Machine Learning
por: Imrie, Fergus, et al.
Publicado: (2022)
por: Imrie, Fergus, et al.
Publicado: (2022)
Semantic-KG: Using Knowledge Graphs to Construct Benchmarks for Measuring Semantic Similarity
por: Wei, Qiyao, et al.
Publicado: (2025)
por: Wei, Qiyao, et al.
Publicado: (2025)
Inverse Reinforcement Learning by Estimating Expertise of Demonstrators
por: Beliaev, Mark, et al.
Publicado: (2024)
por: Beliaev, Mark, et al.
Publicado: (2024)
DC-Check: A Data-Centric AI checklist to guide the development of reliable machine learning systems
por: Seedat, Nabeel, et al.
Publicado: (2022)
por: Seedat, Nabeel, et al.
Publicado: (2022)
Interpretable DNA Sequence Classification via Dynamic Feature Generation in Decision Trees
por: Huynh, Nicolas, et al.
Publicado: (2026)
por: Huynh, Nicolas, et al.
Publicado: (2026)
Learning Neural Control Barrier Functions from Expert Demonstrations using Inverse Constraint Learning
por: Yang, Yuxuan, et al.
Publicado: (2025)
por: Yang, Yuxuan, et al.
Publicado: (2025)
Meta-Learners for Partially-Identified Treatment Effects Across Multiple Environments
por: Schweisthal, Jonas, et al.
Publicado: (2024)
por: Schweisthal, Jonas, et al.
Publicado: (2024)
Ejemplares similares
-
Cascaded Language Models for Cost-effective Human-AI Decision-Making
por: Fanconi, Claudio, et al.
Publicado: (2025) -
Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
por: Sun, Hao, et al.
Publicado: (2024) -
Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
por: Kobalczyk, Katarzyna, et al.
Publicado: (2024) -
CauSim: Scaling Causal Reasoning with Increasingly Complex Causal Simulators
por: Astorga, Nicolás, et al.
Publicado: (2026) -
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
por: Sun, Hao, et al.
Publicado: (2025)