DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zhou, Han, Zhaoyang, Yan, Guochen, Liang, Hao, Zeng, Bohan, Chen, Xing, Song, Yuanfeng, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
von: Zhang, Zehua, et al.
Veröffentlicht: (2025)
von: Zhang, Zehua, et al.
Veröffentlicht: (2025)
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
von: Li, Chenxin, et al.
Veröffentlicht: (2026)
von: Li, Chenxin, et al.
Veröffentlicht: (2026)
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
von: Chi, Wayne, et al.
Veröffentlicht: (2025)
von: Chi, Wayne, et al.
Veröffentlicht: (2025)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
von: Han, Tingxu, et al.
Veröffentlicht: (2026)
von: Han, Tingxu, et al.
Veröffentlicht: (2026)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
von: Gao, Zeyu, et al.
Veröffentlicht: (2025)
von: Gao, Zeyu, et al.
Veröffentlicht: (2025)
SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
von: Ma, Zeyao, et al.
Veröffentlicht: (2024)
von: Ma, Zeyao, et al.
Veröffentlicht: (2024)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
VecIntrinBench: Benchmarking Cross-Architecture Intrinsic Code Migration for RISC-V Vector
von: Han, Liutong, et al.
Veröffentlicht: (2025)
von: Han, Liutong, et al.
Veröffentlicht: (2025)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows
von: Muna, Rabeya Khatun, et al.
Veröffentlicht: (2026)
von: Muna, Rabeya Khatun, et al.
Veröffentlicht: (2026)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
World of Workflows: A Benchmark for Bringing World Models to Enterprise Systems
von: Gupta, Lakshya, et al.
Veröffentlicht: (2026)
von: Gupta, Lakshya, et al.
Veröffentlicht: (2026)
Beyond the YAML File: Understanding Real-World GitHub Actions Workflow Adoption
von: Khatami, Ali, et al.
Veröffentlicht: (2026)
von: Khatami, Ali, et al.
Veröffentlicht: (2026)
PerfBench: Can Agents Resolve Real-World Performance Bugs?
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
von: Xu, Yisen, et al.
Veröffentlicht: (2026)
von: Xu, Yisen, et al.
Veröffentlicht: (2026)
CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System
von: Hu, Li, et al.
Veröffentlicht: (2025)
von: Hu, Li, et al.
Veröffentlicht: (2025)
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
von: Shen, Haiyang, et al.
Veröffentlicht: (2024)
von: Shen, Haiyang, et al.
Veröffentlicht: (2024)
Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development
von: Zeng, Zhengran, et al.
Veröffentlicht: (2025)
von: Zeng, Zhengran, et al.
Veröffentlicht: (2025)
AgentRaft: Automated Detection of Data Over-Exposure in LLM Agents
von: Lin, Yixi, et al.
Veröffentlicht: (2026)
von: Lin, Yixi, et al.
Veröffentlicht: (2026)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
von: Yang, Jie, et al.
Veröffentlicht: (2026)
von: Yang, Jie, et al.
Veröffentlicht: (2026)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
von: Xu, Qiao, et al.
Veröffentlicht: (2026)
von: Xu, Qiao, et al.
Veröffentlicht: (2026)
SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection
von: Ahmed, Md Basim Uddin, et al.
Veröffentlicht: (2025)
von: Ahmed, Md Basim Uddin, et al.
Veröffentlicht: (2025)
LLM-Agents Driven Automated Simulation Testing and Analysis of small Uncrewed Aerial Systems
von: Duvvuru, Venkata Sai Aswath, et al.
Veröffentlicht: (2025)
von: Duvvuru, Venkata Sai Aswath, et al.
Veröffentlicht: (2025)
IDE-Bench: Evaluating Large Language Models as IDE Agents on Real-World Software Engineering Tasks
von: Mateega, Spencer, et al.
Veröffentlicht: (2026)
von: Mateega, Spencer, et al.
Veröffentlicht: (2026)
Decompile-Bench: Million-Scale Binary-Source Function Pairs for Real-World Binary Decompilation
von: Tan, Hanzhuo, et al.
Veröffentlicht: (2025)
von: Tan, Hanzhuo, et al.
Veröffentlicht: (2025)
A Benchmark for Language Models in Real-World System Building
von: Jin, Weilin, et al.
Veröffentlicht: (2026)
von: Jin, Weilin, et al.
Veröffentlicht: (2026)
Heimdallr: Characterizing and Detecting LLM-Induced Security Risks in GitHub CI Workflows
von: Ruan, Bonan, et al.
Veröffentlicht: (2026)
von: Ruan, Bonan, et al.
Veröffentlicht: (2026)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
von: Pereira, Kristen, et al.
Veröffentlicht: (2026)
von: Pereira, Kristen, et al.
Veröffentlicht: (2026)
Engineering AI Agents for Clinical Workflows: A Case Study in Architecture,MLOps, and Governance
von: Lopes, Cláudio Lúcio do Val, et al.
Veröffentlicht: (2026)
von: Lopes, Cláudio Lúcio do Val, et al.
Veröffentlicht: (2026)
Benchmarking and Studying the LLM-based Code Review
von: Zeng, Zhengran, et al.
Veröffentlicht: (2025)
von: Zeng, Zhengran, et al.
Veröffentlicht: (2025)
Statistical Independence Aware Caching for LLM Workflows
von: Dai, Yihan, et al.
Veröffentlicht: (2025)
von: Dai, Yihan, et al.
Veröffentlicht: (2025)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
von: Lu, Pengrui, et al.
Veröffentlicht: (2026)
von: Lu, Pengrui, et al.
Veröffentlicht: (2026)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
von: Zhang, Zehua, et al.
Veröffentlicht: (2025) -
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
von: Li, Chenxin, et al.
Veröffentlicht: (2026) -
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
von: Chi, Wayne, et al.
Veröffentlicht: (2025) -
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
von: Han, Tingxu, et al.
Veröffentlicht: (2026) -
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
von: Gao, Zeyu, et al.
Veröffentlicht: (2025)