RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Atinafu, Yonas, Cohen, Robin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Improving LLM Performance Through Black-Box Online Tuning: A Case for Adding System Specs to Factsheets for Trusted AI
par: Atinafu, Yonas, et autres
Publié: (2026)
par: Atinafu, Yonas, et autres
Publié: (2026)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
par: Thaman, Kunvar
Publié: (2026)
par: Thaman, Kunvar
Publié: (2026)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
par: Reworr, et autres
Publié: (2024)
par: Reworr, et autres
Publié: (2024)
LLM Agents can Autonomously Hack Websites
par: Fang, Richard, et autres
Publié: (2024)
par: Fang, Richard, et autres
Publié: (2024)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
par: Roth, Amit, et autres
Publié: (2026)
par: Roth, Amit, et autres
Publié: (2026)
ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
par: Liu, Zexi, et autres
Publié: (2025)
par: Liu, Zexi, et autres
Publié: (2025)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
par: Zhao, Bingchen, et autres
Publié: (2026)
par: Zhao, Bingchen, et autres
Publié: (2026)
PixLift: Accelerating Web Browsing via AI Upscaling
par: Atinafu, Yonas, et autres
Publié: (2025)
par: Atinafu, Yonas, et autres
Publié: (2025)
Hacking CTFs with Plain Agents
par: Turtayev, Rustem, et autres
Publié: (2024)
par: Turtayev, Rustem, et autres
Publié: (2024)
Reward Hacking as Equilibrium under Finite Evaluation
par: Wang, Jiacheng, et autres
Publié: (2026)
par: Wang, Jiacheng, et autres
Publié: (2026)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
par: Taylor, Mia, et autres
Publié: (2025)
par: Taylor, Mia, et autres
Publié: (2025)
HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
par: Muzsai, Lajos, et autres
Publié: (2024)
par: Muzsai, Lajos, et autres
Publié: (2024)
Evaluation and Benchmarking of LLM Agents: A Survey
par: Mohammadi, Mahmoud, et autres
Publié: (2025)
par: Mohammadi, Mahmoud, et autres
Publié: (2025)
Reward Hacking in Rubric-Based Reinforcement Learning
par: Mahmoud, Anas, et autres
Publié: (2026)
par: Mahmoud, Anas, et autres
Publié: (2026)
MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents
par: Wang, Shouju, et autres
Publié: (2026)
par: Wang, Shouju, et autres
Publié: (2026)
Reward Hacking Mitigation using Verifiable Composite Rewards
par: Tarek, Mirza Farhan Bin, et autres
Publié: (2025)
par: Tarek, Mirza Farhan Bin, et autres
Publié: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
par: Hatgis-Kessell, Stephane, et autres
Publié: (2025)
par: Hatgis-Kessell, Stephane, et autres
Publié: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
par: Fu, Jiayi, et autres
Publié: (2025)
par: Fu, Jiayi, et autres
Publié: (2025)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
par: Deshpande, Darshan, et autres
Publié: (2026)
par: Deshpande, Darshan, et autres
Publié: (2026)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
par: Qiu, Jielin, et autres
Publié: (2025)
par: Qiu, Jielin, et autres
Publié: (2025)
Fighting AI with AI: AI-Agent Augmented DNS Blocking of LLM Services during Student Evaluations
par: Kassa, Yonas, et autres
Publié: (2026)
par: Kassa, Yonas, et autres
Publié: (2026)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
par: Beigi, Mohammad, et autres
Publié: (2026)
par: Beigi, Mohammad, et autres
Publié: (2026)
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
par: Da, Jeff, et autres
Publié: (2025)
par: Da, Jeff, et autres
Publié: (2025)
Proof-of-Use: Mitigating Tool-Call Hacking in Deep Research Agents
par: Ma, SHengjie, et autres
Publié: (2025)
par: Ma, SHengjie, et autres
Publié: (2025)
Playing NetHack with LLMs: Potential & Limitations as Zero-Shot Agents
par: Jeurissen, Dominik, et autres
Publié: (2024)
par: Jeurissen, Dominik, et autres
Publié: (2024)
Spontaneous Reward Hacking in Iterative Self-Refinement
par: Pan, Jane, et autres
Publié: (2024)
par: Pan, Jane, et autres
Publié: (2024)
RRO: LLM Agent Optimization Through Rising Reward Trajectories
par: Wang, Zilong, et autres
Publié: (2025)
par: Wang, Zilong, et autres
Publié: (2025)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
par: Wang, Chaoqi, et autres
Publié: (2025)
par: Wang, Chaoqi, et autres
Publié: (2025)
Prompt Flow Integrity to Prevent Privilege Escalation in LLM Agents
par: Kim, Juhee, et autres
Publié: (2025)
par: Kim, Juhee, et autres
Publié: (2025)
The End of Reward Engineering: How LLMs Are Redefining Multi-Agent Coordination
par: Su, Haoran, et autres
Publié: (2026)
par: Su, Haoran, et autres
Publié: (2026)
AgentDevel: Reframing Self-Evolving LLM Agents as Release Engineering
par: Zhang, Di
Publié: (2026)
par: Zhang, Di
Publié: (2026)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
par: Deng, Shihan, et autres
Publié: (2024)
par: Deng, Shihan, et autres
Publié: (2024)
AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
par: Jiang, Tanqiu, et autres
Publié: (2026)
par: Jiang, Tanqiu, et autres
Publié: (2026)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
par: Lù, Xing Han, et autres
Publié: (2025)
par: Lù, Xing Han, et autres
Publié: (2025)
Evaluate-as-Action: Self-Evaluated Process Rewards for Retrieval-Augmented Agents
par: Shu, Jiangming, et autres
Publié: (2026)
par: Shu, Jiangming, et autres
Publié: (2026)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
par: Ding, Liang
Publié: (2026)
par: Ding, Liang
Publié: (2026)
SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems
par: Yang, Zonglin, et autres
Publié: (2026)
par: Yang, Zonglin, et autres
Publié: (2026)
AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML
par: Trirat, Patara, et autres
Publié: (2024)
par: Trirat, Patara, et autres
Publié: (2024)
Benchmarking LLM Agents for Wealth-Management Workflows
par: Milsom, Rory
Publié: (2025)
par: Milsom, Rory
Publié: (2025)
Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
par: Wang, Yiding, et autres
Publié: (2025)
par: Wang, Yiding, et autres
Publié: (2025)
Documents similaires
-
Improving LLM Performance Through Black-Box Online Tuning: A Case for Adding System Specs to Factsheets for Trusted AI
par: Atinafu, Yonas, et autres
Publié: (2026) -
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
par: Thaman, Kunvar
Publié: (2026) -
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
par: Reworr, et autres
Publié: (2024) -
LLM Agents can Autonomously Hack Websites
par: Fang, Richard, et autres
Publié: (2024) -
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
par: Roth, Amit, et autres
Publié: (2026)