Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Dylan, Srivastava, Pragya, Dragan, Anca, Laidlaw, Cassidy |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
by: Laidlaw, Cassidy, et al.
Published: (2024)
by: Laidlaw, Cassidy, et al.
Published: (2024)
Out of style: Misadventures with LLMs and code style transfer
by: Munson, Karl, et al.
Published: (2024)
by: Munson, Karl, et al.
Published: (2024)
Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
by: Jin, Haolin, et al.
Published: (2025)
by: Jin, Haolin, et al.
Published: (2025)
FailureMem: A Failure-Aware Multimodal Framework for Autonomous Software Repair
by: Ma, Ruize, et al.
Published: (2026)
by: Ma, Ruize, et al.
Published: (2026)
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
by: Roig, JV
Published: (2025)
by: Roig, JV
Published: (2025)
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2026)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
by: Wu, Fan, et al.
Published: (2026)
by: Wu, Fan, et al.
Published: (2026)
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
by: Singha, Ananya, et al.
Published: (2025)
by: Singha, Ananya, et al.
Published: (2025)
Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
by: Lee, Yunseo, et al.
Published: (2025)
by: Lee, Yunseo, et al.
Published: (2025)
GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
by: Jiang, Shufan, et al.
Published: (2026)
by: Jiang, Shufan, et al.
Published: (2026)
Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
LibRec: Benchmarking Retrieval-Augmented LLMs for Library Migration Recommendations
by: Han, Junxiao, et al.
Published: (2025)
by: Han, Junxiao, et al.
Published: (2025)
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
by: Kokane, Shirley, et al.
Published: (2024)
by: Kokane, Shirley, et al.
Published: (2024)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
by: Hu, Ruida, et al.
Published: (2025)
by: Hu, Ruida, et al.
Published: (2025)
Monitoring Monitorability
by: Guan, Melody Y., et al.
Published: (2025)
by: Guan, Melody Y., et al.
Published: (2025)
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs
by: Tan, Qitao, et al.
Published: (2026)
by: Tan, Qitao, et al.
Published: (2026)
Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment
by: Yamani, Asma, et al.
Published: (2025)
by: Yamani, Asma, et al.
Published: (2025)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
LLMs as Continuous Learners: Improving the Reproduction of Defective Code in Software Issues
by: Lin, Yalan, et al.
Published: (2024)
by: Lin, Yalan, et al.
Published: (2024)
Out of Distribution Detection in Self-adaptive Robots with AI-powered Digital Twins
by: Isaku, Erblin, et al.
Published: (2025)
by: Isaku, Erblin, et al.
Published: (2025)
CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
by: Xie, Danning, et al.
Published: (2025)
by: Xie, Danning, et al.
Published: (2025)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
by: Zhu, Hongda, et al.
Published: (2025)
by: Zhu, Hongda, et al.
Published: (2025)
Helping LLMs Improve Code Generation Using Feedback from Testing and Static Analysis
by: Dolcetti, Greta, et al.
Published: (2024)
by: Dolcetti, Greta, et al.
Published: (2024)
$\mathbb{USCD}$: Improving Code Generation of LLMs by Uncertainty-Aware Selective Contrastive Decoding
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
LLMs are All You Need? Improving Fuzz Testing for MOJO with Large Language Models
by: Huang, Linghan, et al.
Published: (2025)
by: Huang, Linghan, et al.
Published: (2025)
Are LLMs Ready for TOON? Benchmarking Structural Correctness-Sustainability Trade-offs in Novel Structured Output Formats
by: Masciari, Elio, et al.
Published: (2026)
by: Masciari, Elio, et al.
Published: (2026)
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
by: Pan, Chenkai, et al.
Published: (2026)
by: Pan, Chenkai, et al.
Published: (2026)
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
by: Laidlaw, Cassidy, et al.
Published: (2023)
by: Laidlaw, Cassidy, et al.
Published: (2023)
Energy-Aware Code Generation with LLMs: Benchmarking Small vs. Large Language Models for Sustainable AI Programming
by: Ashraf, Humza, et al.
Published: (2025)
by: Ashraf, Humza, et al.
Published: (2025)
Willful Disobedience: Automatically Detecting Failures in Agentic Traces
by: Sharma, Reshabh K, et al.
Published: (2026)
by: Sharma, Reshabh K, et al.
Published: (2026)
Demystifying the Lifecycle of Failures in Platform-Orchestrated Agentic Workflows
by: Ma, Xuyan, et al.
Published: (2025)
by: Ma, Xuyan, et al.
Published: (2025)
LLM-Based Automated Diagnosis Of Integration Test Failures At Google
by: Ziftci, Celal, et al.
Published: (2026)
by: Ziftci, Celal, et al.
Published: (2026)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
by: Xu, Qiao, et al.
Published: (2026)
by: Xu, Qiao, et al.
Published: (2026)
Efficient Failure Management for Multi-Agent Systems with Reasoning Trace Representation
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents
by: Zhao, Chenyu, et al.
Published: (2026)
by: Zhao, Chenyu, et al.
Published: (2026)
Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories
by: Majgaonkar, Oorja, et al.
Published: (2025)
by: Majgaonkar, Oorja, et al.
Published: (2025)
LLMs: A Game-Changer for Software Engineers?
by: Haque, Md Asraful
Published: (2024)
by: Haque, Md Asraful
Published: (2024)
RuntimeSlicer: Towards Generalizable Unified Runtime State Representation for Failure Management
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
Seven Failure Points When Engineering a Retrieval Augmented Generation System
by: Barnett, Scott, et al.
Published: (2024)
by: Barnett, Scott, et al.
Published: (2024)
Learning to Debug: LLM-Organized Knowledge Trees for Solving RTL Assertion Failures
by: Bai, Yunsheng, et al.
Published: (2025)
by: Bai, Yunsheng, et al.
Published: (2025)
Similar Items
-
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
by: Laidlaw, Cassidy, et al.
Published: (2024) -
Out of style: Misadventures with LLMs and code style transfer
by: Munson, Karl, et al.
Published: (2024) -
Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
by: Jin, Haolin, et al.
Published: (2025) -
FailureMem: A Failure-Aware Multimodal Framework for Autonomous Software Repair
by: Ma, Ruize, et al.
Published: (2026) -
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
by: Roig, JV
Published: (2025)