CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Haojia, Tan, Xiaoyu, Qin, Yulei, Xu, Zihan, Shi, Yuchen, Li, Zongyi, Li, Gang, Cai, Shaofei, Cai, Siqi, Fu, Chaoyou, Li, Ke, Sun, Xing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
Training-Free Group Relative Policy Optimization
von: Cai, Yuzheng, et al.
Veröffentlicht: (2025)
von: Cai, Yuzheng, et al.
Veröffentlicht: (2025)
Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
von: Qin, Yulei, et al.
Veröffentlicht: (2025)
von: Qin, Yulei, et al.
Veröffentlicht: (2025)
Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization
von: Shi, Yuchen, et al.
Veröffentlicht: (2025)
von: Shi, Yuchen, et al.
Veröffentlicht: (2025)
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
von: Shen, Haiyang, et al.
Veröffentlicht: (2024)
von: Shen, Haiyang, et al.
Veröffentlicht: (2024)
QuanBench: Benchmarking Quantum Code Generation with Large Language Models
von: Guo, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Guo, Xiaoyu, et al.
Veröffentlicht: (2025)
Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
von: Qin, Yulei, et al.
Veröffentlicht: (2025)
von: Qin, Yulei, et al.
Veröffentlicht: (2025)
AL-Bench: A Benchmark for Automatic Logging
von: Tan, Boyin, et al.
Veröffentlicht: (2025)
von: Tan, Boyin, et al.
Veröffentlicht: (2025)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
PyBench: Evaluating LLM Agent on various real-world coding tasks
von: Zhang, Yaolun, et al.
Veröffentlicht: (2024)
von: Zhang, Yaolun, et al.
Veröffentlicht: (2024)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
von: Li, Dawei, et al.
Veröffentlicht: (2026)
von: Li, Dawei, et al.
Veröffentlicht: (2026)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
von: Xiao, Yijia, et al.
Veröffentlicht: (2025)
von: Xiao, Yijia, et al.
Veröffentlicht: (2025)
ConCovUp: Effective Agent-Based Test Driver Generation for Concurrency Testing
von: Cai, Yuandao, et al.
Veröffentlicht: (2026)
von: Cai, Yuandao, et al.
Veröffentlicht: (2026)
RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
von: Merrill, Mike A., et al.
Veröffentlicht: (2026)
von: Merrill, Mike A., et al.
Veröffentlicht: (2026)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
von: Liu, Zhou, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
UniCode: Augmenting Evaluation for Code Reasoning
von: Zheng, Xinyue, et al.
Veröffentlicht: (2025)
von: Zheng, Xinyue, et al.
Veröffentlicht: (2025)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
von: Sun, Jiazheng, et al.
Veröffentlicht: (2026)
von: Sun, Jiazheng, et al.
Veröffentlicht: (2026)
FlowAgent: Achieving Compliance and Flexibility for Workflow Agents
von: Shi, Yuchen, et al.
Veröffentlicht: (2025)
von: Shi, Yuchen, et al.
Veröffentlicht: (2025)
PBT-Bench: Benchmarking AI Agents on Property-Based Testing
von: Jing, Lucas, et al.
Veröffentlicht: (2026)
von: Jing, Lucas, et al.
Veröffentlicht: (2026)
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
von: Guan, Hao, et al.
Veröffentlicht: (2026)
von: Guan, Hao, et al.
Veröffentlicht: (2026)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
von: Chen, Junkai, et al.
Veröffentlicht: (2025)
von: Chen, Junkai, et al.
Veröffentlicht: (2025)
SysTradeBench: An Iterative Build-Test-Patch Benchmark for Strategy-to-Code Trading Systems with Drift-Aware Diagnostics
von: Cao, Yuchen, et al.
Veröffentlicht: (2026)
von: Cao, Yuchen, et al.
Veröffentlicht: (2026)
Empowering RepoQA-Agent based on Reinforcement Learning Driven by Monte-carlo Tree Search
von: Li, Guochang, et al.
Veröffentlicht: (2025)
von: Li, Guochang, et al.
Veröffentlicht: (2025)
VecIntrinBench: Benchmarking Cross-Architecture Intrinsic Code Migration for RISC-V Vector
von: Han, Liutong, et al.
Veröffentlicht: (2025)
von: Han, Liutong, et al.
Veröffentlicht: (2025)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025)
von: Chou, Jason, et al.
Veröffentlicht: (2025)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
von: Guo, Hanyang, et al.
Veröffentlicht: (2025)
von: Guo, Hanyang, et al.
Veröffentlicht: (2025)
Enhancing Security in Third-Party Library Reuse -- Comprehensive Detection of 1-day Vulnerability through Code Patch Analysis
von: Xu, Shangzhi, et al.
Veröffentlicht: (2024)
von: Xu, Shangzhi, et al.
Veröffentlicht: (2024)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale
von: Cai, Yangxiao, et al.
Veröffentlicht: (2025)
von: Cai, Yangxiao, et al.
Veröffentlicht: (2025)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11)
von: Amin, Nishil, et al.
Veröffentlicht: (2026)
von: Amin, Nishil, et al.
Veröffentlicht: (2026)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
RepoMod-Bench: A Benchmark for Code Repository Modernization via Implementation-Agnostic Testing
von: Li, Xuefeng, et al.
Veröffentlicht: (2026)
von: Li, Xuefeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
von: Cai, Shaofei, et al.
Veröffentlicht: (2025) -
Training-Free Group Relative Policy Optimization
von: Cai, Yuzheng, et al.
Veröffentlicht: (2025) -
Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
von: Qin, Yulei, et al.
Veröffentlicht: (2025) -
Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization
von: Shi, Yuchen, et al.
Veröffentlicht: (2025) -
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
von: Shen, Haiyang, et al.
Veröffentlicht: (2024)