AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Shu, Hu, Jingyu, Li, Tong, Yan, Hanqi, Wang, Wenxuan, Wang, Di |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Stability Trap: Evaluating the Reliability of LLM-Based Instruction Adherence Auditing
von: Shergadwala, Murtuza N.
Veröffentlicht: (2026)
von: Shergadwala, Murtuza N.
Veröffentlicht: (2026)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
von: Zhang, Chenchen, et al.
Veröffentlicht: (2025)
von: Zhang, Chenchen, et al.
Veröffentlicht: (2025)
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
von: Wang, Wenxuan
Veröffentlicht: (2024)
von: Wang, Wenxuan
Veröffentlicht: (2024)
FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications
von: Zhang, Yunfan, et al.
Veröffentlicht: (2026)
von: Zhang, Yunfan, et al.
Veröffentlicht: (2026)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025)
von: Chou, Jason, et al.
Veröffentlicht: (2025)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design
von: Qiu, Ruidi, et al.
Veröffentlicht: (2024)
von: Qiu, Ruidi, et al.
Veröffentlicht: (2024)
Evaluating and Achieving Controllable Code Completion in Code LLM
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Wink: Recovering from Misbehaviors in Coding Agents
von: Nanda, Rahul, et al.
Veröffentlicht: (2026)
von: Nanda, Rahul, et al.
Veröffentlicht: (2026)
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
von: Xiang, Yanzheng, et al.
Veröffentlicht: (2025)
von: Xiang, Yanzheng, et al.
Veröffentlicht: (2025)
AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications
von: Shen, Leming, et al.
Veröffentlicht: (2025)
von: Shen, Leming, et al.
Veröffentlicht: (2025)
Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks
von: Chen, Kexin, et al.
Veröffentlicht: (2024)
von: Chen, Kexin, et al.
Veröffentlicht: (2024)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
von: Jain, Naman, et al.
Veröffentlicht: (2024)
von: Jain, Naman, et al.
Veröffentlicht: (2024)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
von: Li, Shuqing, et al.
Veröffentlicht: (2025)
von: Li, Shuqing, et al.
Veröffentlicht: (2025)
Showing LLM-Generated Code Selectively Based on Confidence of LLMs
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
CodeRAG-Bench: Can Retrieval Augment Code Generation?
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
von: Ma, Zeyao, et al.
Veröffentlicht: (2024)
von: Ma, Zeyao, et al.
Veröffentlicht: (2024)
Rigor, Reliability, and Reproducibility Matter: A Decade-Scale Survey of 572 Code Benchmarks
von: Cao, Jialun, et al.
Veröffentlicht: (2025)
von: Cao, Jialun, et al.
Veröffentlicht: (2025)
Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems
von: Martinez, Matias, et al.
Veröffentlicht: (2025)
von: Martinez, Matias, et al.
Veröffentlicht: (2025)
Comparing Developer and LLM Biases in Code Evaluation
von: Mittal, Aditya, et al.
Veröffentlicht: (2026)
von: Mittal, Aditya, et al.
Veröffentlicht: (2026)
FeatBench: Towards More Realistic Evaluation of Feature-level Code Generation
von: Chen, Haorui, et al.
Veröffentlicht: (2025)
von: Chen, Haorui, et al.
Veröffentlicht: (2025)
Genetic Auto-prompt Learning for Pre-trained Code Intelligence Language Models
von: Feng, Chengzhe, et al.
Veröffentlicht: (2024)
von: Feng, Chengzhe, et al.
Veröffentlicht: (2024)
Code Fingerprints: Disentangled Attribution of LLM-Generated Code
von: Guo, Jiaxun, et al.
Veröffentlicht: (2026)
von: Guo, Jiaxun, et al.
Veröffentlicht: (2026)
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation
von: Pysklo, Hubert M., et al.
Veröffentlicht: (2026)
von: Pysklo, Hubert M., et al.
Veröffentlicht: (2026)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
von: Chi, Wayne, et al.
Veröffentlicht: (2026)
von: Chi, Wayne, et al.
Veröffentlicht: (2026)
Reasoning Runtime Behavior of a Program with LLM: How Far Are We?
von: Chen, Junkai, et al.
Veröffentlicht: (2024)
von: Chen, Junkai, et al.
Veröffentlicht: (2024)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
von: Lai, Peng, et al.
Veröffentlicht: (2026)
von: Lai, Peng, et al.
Veröffentlicht: (2026)
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
von: Diddee, Harshita, et al.
Veröffentlicht: (2026)
von: Diddee, Harshita, et al.
Veröffentlicht: (2026)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
CommitBench: A Benchmark for Commit Message Generation
von: Schall, Maximilian, et al.
Veröffentlicht: (2024)
von: Schall, Maximilian, et al.
Veröffentlicht: (2024)
Stingy Context: 18:1 Hierarchical Code Compression for LLM Auto-Coding
von: Ostby, David Linus
Veröffentlicht: (2026)
von: Ostby, David Linus
Veröffentlicht: (2026)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Stability Trap: Evaluating the Reliability of LLM-Based Instruction Adherence Auditing
von: Shergadwala, Murtuza N.
Veröffentlicht: (2026) -
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
von: Zhang, Chenchen, et al.
Veröffentlicht: (2025) -
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
von: Wang, Wenxuan
Veröffentlicht: (2024) -
FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications
von: Zhang, Yunfan, et al.
Veröffentlicht: (2026) -
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025)