Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Feng, Dylan, Srivastava, Pragya, Dragan, Anca, Laidlaw, Cassidy |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
par: Laidlaw, Cassidy, et autres
Publié: (2024)
par: Laidlaw, Cassidy, et autres
Publié: (2024)
Out of style: Misadventures with LLMs and code style transfer
par: Munson, Karl, et autres
Publié: (2024)
par: Munson, Karl, et autres
Publié: (2024)
Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
par: Jin, Haolin, et autres
Publié: (2025)
par: Jin, Haolin, et autres
Publié: (2025)
FailureMem: A Failure-Aware Multimodal Framework for Autonomous Software Repair
par: Ma, Ruize, et autres
Publié: (2026)
par: Ma, Ruize, et autres
Publié: (2026)
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
par: Roig, JV
Publié: (2025)
par: Roig, JV
Publié: (2025)
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2026)
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2026)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
par: Wu, Fan, et autres
Publié: (2026)
par: Wu, Fan, et autres
Publié: (2026)
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
par: Singha, Ananya, et autres
Publié: (2025)
par: Singha, Ananya, et autres
Publié: (2025)
Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
par: Lee, Yunseo, et autres
Publié: (2025)
par: Lee, Yunseo, et autres
Publié: (2025)
GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
par: Jiang, Shufan, et autres
Publié: (2026)
par: Jiang, Shufan, et autres
Publié: (2026)
Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering
par: Alebachew, Yoseph Berhanu, et autres
Publié: (2026)
par: Alebachew, Yoseph Berhanu, et autres
Publié: (2026)
LibRec: Benchmarking Retrieval-Augmented LLMs for Library Migration Recommendations
par: Han, Junxiao, et autres
Publié: (2025)
par: Han, Junxiao, et autres
Publié: (2025)
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
par: Kokane, Shirley, et autres
Publié: (2024)
par: Kokane, Shirley, et autres
Publié: (2024)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
par: Hu, Ruida, et autres
Publié: (2025)
par: Hu, Ruida, et autres
Publié: (2025)
Monitoring Monitorability
par: Guan, Melody Y., et autres
Publié: (2025)
par: Guan, Melody Y., et autres
Publié: (2025)
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs
par: Tan, Qitao, et autres
Publié: (2026)
par: Tan, Qitao, et autres
Publié: (2026)
Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment
par: Yamani, Asma, et autres
Publié: (2025)
par: Yamani, Asma, et autres
Publié: (2025)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
par: Ngassom, Sylvain Kouemo, et autres
Publié: (2024)
par: Ngassom, Sylvain Kouemo, et autres
Publié: (2024)
LLMs as Continuous Learners: Improving the Reproduction of Defective Code in Software Issues
par: Lin, Yalan, et autres
Publié: (2024)
par: Lin, Yalan, et autres
Publié: (2024)
Out of Distribution Detection in Self-adaptive Robots with AI-powered Digital Twins
par: Isaku, Erblin, et autres
Publié: (2025)
par: Isaku, Erblin, et autres
Publié: (2025)
CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
par: Xie, Danning, et autres
Publié: (2025)
par: Xie, Danning, et autres
Publié: (2025)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
par: Zhu, Hongda, et autres
Publié: (2025)
par: Zhu, Hongda, et autres
Publié: (2025)
Helping LLMs Improve Code Generation Using Feedback from Testing and Static Analysis
par: Dolcetti, Greta, et autres
Publié: (2024)
par: Dolcetti, Greta, et autres
Publié: (2024)
$\mathbb{USCD}$: Improving Code Generation of LLMs by Uncertainty-Aware Selective Contrastive Decoding
par: Wang, Shuai, et autres
Publié: (2024)
par: Wang, Shuai, et autres
Publié: (2024)
LLMs are All You Need? Improving Fuzz Testing for MOJO with Large Language Models
par: Huang, Linghan, et autres
Publié: (2025)
par: Huang, Linghan, et autres
Publié: (2025)
Are LLMs Ready for TOON? Benchmarking Structural Correctness-Sustainability Trade-offs in Novel Structured Output Formats
par: Masciari, Elio, et autres
Publié: (2026)
par: Masciari, Elio, et autres
Publié: (2026)
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
par: Pan, Chenkai, et autres
Publié: (2026)
par: Pan, Chenkai, et autres
Publié: (2026)
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
par: Laidlaw, Cassidy, et autres
Publié: (2023)
par: Laidlaw, Cassidy, et autres
Publié: (2023)
Energy-Aware Code Generation with LLMs: Benchmarking Small vs. Large Language Models for Sustainable AI Programming
par: Ashraf, Humza, et autres
Publié: (2025)
par: Ashraf, Humza, et autres
Publié: (2025)
Willful Disobedience: Automatically Detecting Failures in Agentic Traces
par: Sharma, Reshabh K, et autres
Publié: (2026)
par: Sharma, Reshabh K, et autres
Publié: (2026)
Demystifying the Lifecycle of Failures in Platform-Orchestrated Agentic Workflows
par: Ma, Xuyan, et autres
Publié: (2025)
par: Ma, Xuyan, et autres
Publié: (2025)
LLM-Based Automated Diagnosis Of Integration Test Failures At Google
par: Ziftci, Celal, et autres
Publié: (2026)
par: Ziftci, Celal, et autres
Publié: (2026)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
par: Xu, Qiao, et autres
Publié: (2026)
par: Xu, Qiao, et autres
Publié: (2026)
Efficient Failure Management for Multi-Agent Systems with Reasoning Trace Representation
par: Zhang, Lingzhe, et autres
Publié: (2026)
par: Zhang, Lingzhe, et autres
Publié: (2026)
Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents
par: Zhao, Chenyu, et autres
Publié: (2026)
par: Zhao, Chenyu, et autres
Publié: (2026)
Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories
par: Majgaonkar, Oorja, et autres
Publié: (2025)
par: Majgaonkar, Oorja, et autres
Publié: (2025)
LLMs: A Game-Changer for Software Engineers?
par: Haque, Md Asraful
Publié: (2024)
par: Haque, Md Asraful
Publié: (2024)
RuntimeSlicer: Towards Generalizable Unified Runtime State Representation for Failure Management
par: Zhang, Lingzhe, et autres
Publié: (2026)
par: Zhang, Lingzhe, et autres
Publié: (2026)
Seven Failure Points When Engineering a Retrieval Augmented Generation System
par: Barnett, Scott, et autres
Publié: (2024)
par: Barnett, Scott, et autres
Publié: (2024)
Learning to Debug: LLM-Organized Knowledge Trees for Solving RTL Assertion Failures
par: Bai, Yunsheng, et autres
Publié: (2025)
par: Bai, Yunsheng, et autres
Publié: (2025)
Documents similaires
-
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
par: Laidlaw, Cassidy, et autres
Publié: (2024) -
Out of style: Misadventures with LLMs and code style transfer
par: Munson, Karl, et autres
Publié: (2024) -
Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
par: Jin, Haolin, et autres
Publié: (2025) -
FailureMem: A Failure-Aware Multimodal Framework for Autonomous Software Repair
par: Ma, Ruize, et autres
Publié: (2026) -
How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations
par: Roig, JV
Publié: (2025)