Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Patel, Nyal, Bou, Matthieu, Jagota, Arjun, Krishna, Satyapriya, Parbhoo, Sonali |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
von: Bou, Matthieu, et al.
Veröffentlicht: (2025)
von: Bou, Matthieu, et al.
Veröffentlicht: (2025)
Insights from the Inverse: Reconstructing LLM Training Goals Through Inverse Reinforcement Learning
von: Joselowitz, Jared, et al.
Veröffentlicht: (2024)
von: Joselowitz, Jared, et al.
Veröffentlicht: (2024)
Solving the Inverse Alignment Problem for Efficient RLHF
von: Krishna, Shambhavi, et al.
Veröffentlicht: (2024)
von: Krishna, Shambhavi, et al.
Veröffentlicht: (2024)
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
von: Ji, Xiaotong, et al.
Veröffentlicht: (2025)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2025)
Accurate Failure Prediction in Agents Does Not Imply Effective Failure Prevention
von: Vasudev, Rakshith, et al.
Veröffentlicht: (2026)
von: Vasudev, Rakshith, et al.
Veröffentlicht: (2026)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2025)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2025)
Learning from Failures in Multi-Attempt Reinforcement Learning
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
von: Chung, Stephen, et al.
Veröffentlicht: (2025)
Tree-Based Leakage Inspection and Control in Concept Bottleneck Models
von: Ragkousis, Angelos, et al.
Veröffentlicht: (2024)
von: Ragkousis, Angelos, et al.
Veröffentlicht: (2024)
GLIDE-RL: Grounded Language Instruction through DEmonstration in RL
von: Kharyal, Chaitanya, et al.
Veröffentlicht: (2024)
von: Kharyal, Chaitanya, et al.
Veröffentlicht: (2024)
Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories
von: Benac, Leo, et al.
Veröffentlicht: (2024)
von: Benac, Leo, et al.
Veröffentlicht: (2024)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
Failure Modes of Maximum Entropy RLHF
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2025)
von: Çağatan, Ömer Veysel, et al.
Veröffentlicht: (2025)
Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
Failure Modes of LLMs for Causal Reasoning on Narratives
von: Yamin, Khurram, et al.
Veröffentlicht: (2024)
von: Yamin, Khurram, et al.
Veröffentlicht: (2024)
Diagnosing Structural Failures in LLM-Based Evidence Extraction for Meta-Analysis
von: Tan, Zhiyin, et al.
Veröffentlicht: (2026)
von: Tan, Zhiyin, et al.
Veröffentlicht: (2026)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
von: Shi, Zhengyan, et al.
Veröffentlicht: (2024)
von: Shi, Zhengyan, et al.
Veröffentlicht: (2024)
CoG: Controllable Graph Reasoning via Relational Blueprints and Failure-Aware Refinement over Knowledge Graphs
von: Liu, Yuanxiang, et al.
Veröffentlicht: (2026)
von: Liu, Yuanxiang, et al.
Veröffentlicht: (2026)
A Theoretical Understanding of Self-Correction through In-context Alignment
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning
von: Deng, Hexuan, et al.
Veröffentlicht: (2025)
von: Deng, Hexuan, et al.
Veröffentlicht: (2025)
TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning
von: Shandilya, Shivam, et al.
Veröffentlicht: (2024)
von: Shandilya, Shivam, et al.
Veröffentlicht: (2024)
From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization
von: Zhou, Chenxi, et al.
Veröffentlicht: (2026)
von: Zhou, Chenxi, et al.
Veröffentlicht: (2026)
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
Large Language Model Reasoning Failures
von: Song, Peiyang, et al.
Veröffentlicht: (2026)
von: Song, Peiyang, et al.
Veröffentlicht: (2026)
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
von: Zhang, Kongcheng, et al.
Veröffentlicht: (2025)
von: Zhang, Kongcheng, et al.
Veröffentlicht: (2025)
Efficient Detection of Intermittent Job Failures Using Few-Shot Learning
von: Aïdasso, Henri, et al.
Veröffentlicht: (2025)
von: Aïdasso, Henri, et al.
Veröffentlicht: (2025)
Understanding the Effects of Iterative Prompting on Truthfulness
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2024)
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2024)
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
von: Wang, Zhenting, et al.
Veröffentlicht: (2025)
von: Wang, Zhenting, et al.
Veröffentlicht: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
Do regularization methods for shortcut mitigation work as intended?
von: Hong, Haoyang, et al.
Veröffentlicht: (2025)
von: Hong, Haoyang, et al.
Veröffentlicht: (2025)
DualDiffusion: A Speculative Decoding Strategy for Masked Diffusion Models
von: Goyal, Satyam, et al.
Veröffentlicht: (2026)
von: Goyal, Satyam, et al.
Veröffentlicht: (2026)
Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization Failure
von: Wang, Boshi, et al.
Veröffentlicht: (2025)
von: Wang, Boshi, et al.
Veröffentlicht: (2025)
Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
von: Wu, Haoze, et al.
Veröffentlicht: (2025)
von: Wu, Haoze, et al.
Veröffentlicht: (2025)
PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
von: Qiu, Zhaopeng, et al.
Veröffentlicht: (2026)
von: Qiu, Zhaopeng, et al.
Veröffentlicht: (2026)
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
von: Lin, Zihan, et al.
Veröffentlicht: (2026)
von: Lin, Zihan, et al.
Veröffentlicht: (2026)
Mass-Producing Failures of Multimodal Systems with Language Models
von: Tong, Shengbang, et al.
Veröffentlicht: (2023)
von: Tong, Shengbang, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
von: Bou, Matthieu, et al.
Veröffentlicht: (2025) -
Insights from the Inverse: Reconstructing LLM Training Goals Through Inverse Reinforcement Learning
von: Joselowitz, Jared, et al.
Veröffentlicht: (2024) -
Solving the Inverse Alignment Problem for Efficient RLHF
von: Krishna, Shambhavi, et al.
Veröffentlicht: (2024) -
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
von: Ji, Xiaotong, et al.
Veröffentlicht: (2025) -
Accurate Failure Prediction in Agents Does Not Imply Effective Failure Prevention
von: Vasudev, Rakshith, et al.
Veröffentlicht: (2026)