Insights from the Inverse: Reconstructing LLM Training Goals Through Inverse Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Joselowitz, Jared, Majumdar, Ritam, Jagota, Arjun, Bou, Matthieu, Patel, Nyal, Krishna, Satyapriya, Parbhoo, Sonali |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
by: Patel, Nyal, et al.
Published: (2025)
by: Patel, Nyal, et al.
Published: (2025)
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
by: Bou, Matthieu, et al.
Published: (2025)
by: Bou, Matthieu, et al.
Published: (2025)
Concept-driven Off Policy Evaluation
by: Majumdar, Ritam, et al.
Published: (2024)
by: Majumdar, Ritam, et al.
Published: (2024)
Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
by: Roy, Amartya, et al.
Published: (2026)
by: Roy, Amartya, et al.
Published: (2026)
Improving ARDS Diagnosis Through Context-Aware Concept Bottleneck Models
by: Narain, Anish, et al.
Published: (2025)
by: Narain, Anish, et al.
Published: (2025)
Understanding the Effects of Iterative Prompting on Truthfulness
by: Krishna, Satyapriya, et al.
Published: (2024)
by: Krishna, Satyapriya, et al.
Published: (2024)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
by: Cheng, Ruoxi, et al.
Published: (2025)
by: Cheng, Ruoxi, et al.
Published: (2025)
Imitating Language via Scalable Inverse Reinforcement Learning
by: Wulfmeier, Markus, et al.
Published: (2024)
by: Wulfmeier, Markus, et al.
Published: (2024)
ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems
by: Chowdhury, Mohita, et al.
Published: (2025)
by: Chowdhury, Mohita, et al.
Published: (2025)
Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories
by: Benac, Leo, et al.
Published: (2024)
by: Benac, Leo, et al.
Published: (2024)
Solving the Inverse Alignment Problem for Efficient RLHF
by: Krishna, Shambhavi, et al.
Published: (2024)
by: Krishna, Shambhavi, et al.
Published: (2024)
About Time: Model-free Reinforcement Learning with Timed Reward Machines
by: Roy, Rajarshi, et al.
Published: (2025)
by: Roy, Rajarshi, et al.
Published: (2025)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
by: Li, Aaron J., et al.
Published: (2024)
by: Li, Aaron J., et al.
Published: (2024)
Supervised Fine-Tuning as Inverse Reinforcement Learning
by: Sun, Hao
Published: (2024)
by: Sun, Hao
Published: (2024)
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
by: Krishna, Satyapriya, et al.
Published: (2024)
by: Krishna, Satyapriya, et al.
Published: (2024)
Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data
by: Xia, Han, et al.
Published: (2024)
by: Xia, Han, et al.
Published: (2024)
Pragmatic Instruction Following and Goal Assistance via Cooperative Language-Guided Inverse Planning
by: Zhi-Xuan, Tan, et al.
Published: (2024)
by: Zhi-Xuan, Tan, et al.
Published: (2024)
Reinforcement Learning for LLM Post-Training: A Survey
by: Wang, Zhichao, et al.
Published: (2024)
by: Wang, Zhichao, et al.
Published: (2024)
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage
by: Oh, Byung-Doh, et al.
Published: (2025)
by: Oh, Byung-Doh, et al.
Published: (2025)
Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
by: Zhang, Qinyan, et al.
Published: (2025)
by: Zhang, Qinyan, et al.
Published: (2025)
Self-Correcting Large Language Models: Generation vs. Multiple Choice
by: Rahmani, Hossein A., et al.
Published: (2025)
by: Rahmani, Hossein A., et al.
Published: (2025)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023)
by: Kroeger, Nicholas, et al.
Published: (2023)
Rethinking Weight Tying: Pseudo-Inverse Tying for LM Stable Training and Updates
by: Gu, Jian, et al.
Published: (2026)
by: Gu, Jian, et al.
Published: (2026)
Learning Design Skills as Memory Policies for Agentic Photonic Inverse Design
by: Chen, Shengchao, et al.
Published: (2026)
by: Chen, Shengchao, et al.
Published: (2026)
Many-Shot In-Context Learning for Molecular Inverse Design
by: Moayedpour, Saeed, et al.
Published: (2024)
by: Moayedpour, Saeed, et al.
Published: (2024)
Transferable Post-training via Inverse Value Learning
by: Lu, Xinyu, et al.
Published: (2024)
by: Lu, Xinyu, et al.
Published: (2024)
Byte BPE Tokenization as an Inverse string Homomorphism
by: Geng, Saibo, et al.
Published: (2024)
by: Geng, Saibo, et al.
Published: (2024)
WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue
by: Ellis, Zachary, et al.
Published: (2025)
by: Ellis, Zachary, et al.
Published: (2025)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
by: Krishna, Arjun, et al.
Published: (2025)
by: Krishna, Arjun, et al.
Published: (2025)
Fairness Dynamics During Training
by: Patel, Krishna, et al.
Published: (2025)
by: Patel, Krishna, et al.
Published: (2025)
Seeing Through AI's Lens: Enhancing Human Skepticism Towards LLM-Generated Fake News
by: Ayoobi, Navid, et al.
Published: (2024)
by: Ayoobi, Navid, et al.
Published: (2024)
InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct
by: Wu, Yutong, et al.
Published: (2024)
by: Wu, Yutong, et al.
Published: (2024)
Emergent inabilities? Inverse scaling over the course of pretraining
by: Michaelov, James A., et al.
Published: (2023)
by: Michaelov, James A., et al.
Published: (2023)
Inverse Language Modeling towards Robust and Grounded LLMs
by: Gabrielli, Davide, et al.
Published: (2025)
by: Gabrielli, Davide, et al.
Published: (2025)
Inverse Scaling in Test-Time Compute
by: Gema, Aryo Pradipta, et al.
Published: (2025)
by: Gema, Aryo Pradipta, et al.
Published: (2025)
MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents
by: Anand, Ashwani, et al.
Published: (2026)
by: Anand, Ashwani, et al.
Published: (2026)
Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning
by: Yalcinkaya, Beyazit, et al.
Published: (2024)
by: Yalcinkaya, Beyazit, et al.
Published: (2024)
Tree-Based Leakage Inspection and Control in Concept Bottleneck Models
by: Ragkousis, Angelos, et al.
Published: (2024)
by: Ragkousis, Angelos, et al.
Published: (2024)
Explainability Through Systematicity: The Hard Systematicity Challenge for Artificial Intelligence
by: Queloz, Matthieu
Published: (2025)
by: Queloz, Matthieu
Published: (2025)
Similar Items
-
Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
by: Patel, Nyal, et al.
Published: (2025) -
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
by: Bou, Matthieu, et al.
Published: (2025) -
Concept-driven Off Policy Evaluation
by: Majumdar, Ritam, et al.
Published: (2024) -
Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
by: Roy, Amartya, et al.
Published: (2026) -
Improving ARDS Diagnosis Through Context-Aware Concept Bottleneck Models
by: Narain, Anish, et al.
Published: (2025)