DCA-Bench: A Benchmark for Dataset Curation Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Benhao, Yu, Yingzhuo, Huang, Jin, Zhang, Xingjian, Ma, Jiaqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AAPMT: AGI Assessment Through Prompt and Metric Transformer
von: Huang, Benhao
Veröffentlicht: (2024)
von: Huang, Benhao
Veröffentlicht: (2024)
Can LLMs Effectively Leverage Graph Structural Information through Prompts, and Why?
von: Huang, Jin, et al.
Veröffentlicht: (2023)
von: Huang, Jin, et al.
Veröffentlicht: (2023)
Bench-CoE: a Framework for Collaboration of Experts from Benchmark
von: Wang, Yuanshuai, et al.
Veröffentlicht: (2024)
von: Wang, Yuanshuai, et al.
Veröffentlicht: (2024)
FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases
von: Zhang, Xingjian, et al.
Veröffentlicht: (2026)
von: Zhang, Xingjian, et al.
Veröffentlicht: (2026)
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
von: Huang, Jiahao, et al.
Veröffentlicht: (2026)
von: Huang, Jiahao, et al.
Veröffentlicht: (2026)
RL-MSA: a Reinforcement Learning-based Multi-line bus Scheduling Approach
von: Liu, Yingzhuo
Veröffentlicht: (2024)
von: Liu, Yingzhuo
Veröffentlicht: (2024)
MASSW: A New Dataset and Benchmark Tasks for AI-Assisted Scientific Workflows
von: Zhang, Xingjian, et al.
Veröffentlicht: (2024)
von: Zhang, Xingjian, et al.
Veröffentlicht: (2024)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
von: Zhang, Hanrong, et al.
Veröffentlicht: (2024)
von: Zhang, Hanrong, et al.
Veröffentlicht: (2024)
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
von: Dong, Haonan, et al.
Veröffentlicht: (2026)
von: Dong, Haonan, et al.
Veröffentlicht: (2026)
FedRS-Bench: Realistic Federated Learning Datasets and Benchmarks in Remote Sensing
von: Zhao, Haodong, et al.
Veröffentlicht: (2025)
von: Zhao, Haodong, et al.
Veröffentlicht: (2025)
AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
von: Shang, Yu, et al.
Veröffentlicht: (2025)
von: Shang, Yu, et al.
Veröffentlicht: (2025)
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
von: Yu, Bo, et al.
Veröffentlicht: (2026)
von: Yu, Bo, et al.
Veröffentlicht: (2026)
SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval
von: Li, Ningyuan, et al.
Veröffentlicht: (2026)
von: Li, Ningyuan, et al.
Veröffentlicht: (2026)
SoundnessBench: A Soundness Benchmark for Neural Network Verifiers
von: Zhou, Xingjian, et al.
Veröffentlicht: (2024)
von: Zhou, Xingjian, et al.
Veröffentlicht: (2024)
MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2025)
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
GUI-360$^\circ$: A Comprehensive Dataset and Benchmark for Computer-Using Agents
von: Mu, Jian, et al.
Veröffentlicht: (2025)
von: Mu, Jian, et al.
Veröffentlicht: (2025)
ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
von: Jin, Tengjun, et al.
Veröffentlicht: (2025)
von: Jin, Tengjun, et al.
Veröffentlicht: (2025)
IGL-Bench: Establishing the Comprehensive Benchmark for Imbalanced Graph Learning
von: Qin, Jiawen, et al.
Veröffentlicht: (2024)
von: Qin, Jiawen, et al.
Veröffentlicht: (2024)
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
von: Xiong, Lei, et al.
Veröffentlicht: (2026)
von: Xiong, Lei, et al.
Veröffentlicht: (2026)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
von: Yin, Sheng, et al.
Veröffentlicht: (2024)
von: Yin, Sheng, et al.
Veröffentlicht: (2024)
Defining and Extracting generalizable interaction primitives from DNNs
von: Chen, Lu, et al.
Veröffentlicht: (2024)
von: Chen, Lu, et al.
Veröffentlicht: (2024)
WritingBench: A Comprehensive Benchmark for Generative Writing
von: Wu, Yuning, et al.
Veröffentlicht: (2025)
von: Wu, Yuning, et al.
Veröffentlicht: (2025)
PSPA-Bench: A Personalized Benchmark for Smartphone GUI Agent
von: Nie, Hongyi, et al.
Veröffentlicht: (2026)
von: Nie, Hongyi, et al.
Veröffentlicht: (2026)
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
von: Choi, Chanyeol, et al.
Veröffentlicht: (2025)
von: Choi, Chanyeol, et al.
Veröffentlicht: (2025)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
von: Song, Yuanyi, et al.
Veröffentlicht: (2025)
von: Song, Yuanyi, et al.
Veröffentlicht: (2025)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
von: Styles, Olly, et al.
Veröffentlicht: (2024)
von: Styles, Olly, et al.
Veröffentlicht: (2024)
MedConsultBench: A Full-Cycle, Fine-Grained, Process-Aware Benchmark for Medical Consultation Agents
von: Qiao, Chuhan, et al.
Veröffentlicht: (2026)
von: Qiao, Chuhan, et al.
Veröffentlicht: (2026)
RelBench: A Benchmark for Deep Learning on Relational Databases
von: Robinson, Joshua, et al.
Veröffentlicht: (2024)
von: Robinson, Joshua, et al.
Veröffentlicht: (2024)
ELT-Bench-Verified: Benchmark Quality Issues Underestimate AI Agent Capabilities
von: Zanoli, Christopher, et al.
Veröffentlicht: (2026)
von: Zanoli, Christopher, et al.
Veröffentlicht: (2026)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
von: Zhou, Yifan, et al.
Veröffentlicht: (2026)
von: Zhou, Yifan, et al.
Veröffentlicht: (2026)
Escaping the Context Bottleneck: Active Context Curation for LLM Agents via Reinforcement Learning
von: Li, Xiaozhe, et al.
Veröffentlicht: (2026)
von: Li, Xiaozhe, et al.
Veröffentlicht: (2026)
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
von: Shi, Wentao, et al.
Veröffentlicht: (2026)
von: Shi, Wentao, et al.
Veröffentlicht: (2026)
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents
von: Shen, Haiyang, et al.
Veröffentlicht: (2024)
von: Shen, Haiyang, et al.
Veröffentlicht: (2024)
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
von: Levy, Ido, et al.
Veröffentlicht: (2024)
von: Levy, Ido, et al.
Veröffentlicht: (2024)
InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents
von: Zhu, Zhenghao, et al.
Veröffentlicht: (2025)
von: Zhu, Zhenghao, et al.
Veröffentlicht: (2025)
AgentSearchBench: A Benchmark for AI Agent Search in the Wild
von: Wu, Bin, et al.
Veröffentlicht: (2026)
von: Wu, Bin, et al.
Veröffentlicht: (2026)
Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry
von: Peng, Run, et al.
Veröffentlicht: (2025)
von: Peng, Run, et al.
Veröffentlicht: (2025)
SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation
von: He, Yibo, et al.
Veröffentlicht: (2025)
von: He, Yibo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AAPMT: AGI Assessment Through Prompt and Metric Transformer
von: Huang, Benhao
Veröffentlicht: (2024) -
Can LLMs Effectively Leverage Graph Structural Information through Prompts, and Why?
von: Huang, Jin, et al.
Veröffentlicht: (2023) -
Bench-CoE: a Framework for Collaboration of Experts from Benchmark
von: Wang, Yuanshuai, et al.
Veröffentlicht: (2024) -
FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases
von: Zhang, Xingjian, et al.
Veröffentlicht: (2026) -
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
von: Huang, Jiahao, et al.
Veröffentlicht: (2026)