AutomationBench
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shepard, Daniel, Salimans, Robin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Automating Document Intelligence in Statutory City Planning
par: Malmqvist, Lars, et autres
Publié: (2026)
par: Malmqvist, Lars, et autres
Publié: (2026)
Adopting Large Language Models to Automated System Integration
par: Pesl, Robin D.
Publié: (2025)
par: Pesl, Robin D.
Publié: (2025)
Multistep Distillation of Diffusion Models via Moment Matching
par: Salimans, Tim, et autres
Publié: (2024)
par: Salimans, Tim, et autres
Publié: (2024)
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
par: Jotautaitė, Monika, et autres
Publié: (2026)
par: Jotautaitė, Monika, et autres
Publié: (2026)
TaskBench: Benchmarking Large Language Models for Task Automation
par: Shen, Yongliang, et autres
Publié: (2023)
par: Shen, Yongliang, et autres
Publié: (2023)
Bench4KE: Benchmarking Automated Competency Question Generation
par: Lippolis, Anna Sofia, et autres
Publié: (2025)
par: Lippolis, Anna Sofia, et autres
Publié: (2025)
LiveMedBench: A Contamination-Free Medical Benchmark for LLMs with Automated Rubric Evaluation
par: Yan, Zhiling, et autres
Publié: (2026)
par: Yan, Zhiling, et autres
Publié: (2026)
BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting
par: Wang, Zhensheng, et autres
Publié: (2026)
par: Wang, Zhensheng, et autres
Publié: (2026)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
par: Patel, Liana, et autres
Publié: (2025)
par: Patel, Liana, et autres
Publié: (2025)
FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)
par: Priyanshu, Aman, et autres
Publié: (2024)
par: Priyanshu, Aman, et autres
Publié: (2024)
IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations
par: Fu, Deqing, et autres
Publié: (2024)
par: Fu, Deqing, et autres
Publié: (2024)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
par: Oliva, Gustavo A., et autres
Publié: (2025)
par: Oliva, Gustavo A., et autres
Publié: (2025)
AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance
par: Patel, Dhaval, et autres
Publié: (2025)
par: Patel, Dhaval, et autres
Publié: (2025)
Automated Generation and Tagging of Knowledge Components from Multiple-Choice Questions
par: Moore, Steven, et autres
Publié: (2024)
par: Moore, Steven, et autres
Publié: (2024)
Counterfactual Reasoning in Automated Planning
par: Pozanco, Alberto, et autres
Publié: (2026)
par: Pozanco, Alberto, et autres
Publié: (2026)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
par: Rank, Ben, et autres
Publié: (2026)
par: Rank, Ben, et autres
Publié: (2026)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
par: Tu, Xinming, et autres
Publié: (2026)
par: Tu, Xinming, et autres
Publié: (2026)
DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering
par: Li, Yinsheng, et autres
Publié: (2025)
par: Li, Yinsheng, et autres
Publié: (2025)
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
par: Chen, Haolin, et autres
Publié: (2026)
par: Chen, Haolin, et autres
Publié: (2026)
Deep Research Bench: Evaluating AI Web Research Agents
par: FutureSearch, et autres
Publié: (2025)
par: FutureSearch, et autres
Publié: (2025)
ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
par: Jin, Tengjun, et autres
Publié: (2025)
par: Jin, Tengjun, et autres
Publié: (2025)
RamanBench: A Large-Scale Benchmark for Machine Learning on Raman Spectroscopy
par: Koddenbrock, Mario, et autres
Publié: (2026)
par: Koddenbrock, Mario, et autres
Publié: (2026)
EM Distillation for One-step Diffusion Models
par: Xie, Sirui, et autres
Publié: (2024)
par: Xie, Sirui, et autres
Publié: (2024)
MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
par: Kim, Hyunjun, et autres
Publié: (2025)
par: Kim, Hyunjun, et autres
Publié: (2025)
TeamBench: Evaluating Agent Coordination under Enforced Role Separation
par: Kim, Yubin, et autres
Publié: (2026)
par: Kim, Yubin, et autres
Publié: (2026)
TaoBench: Do Automated Theorem Prover LLMs Generalize Beyond MathLib?
par: Taylor, Alexander K, et autres
Publié: (2026)
par: Taylor, Alexander K, et autres
Publié: (2026)
WorkstreamBench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
par: Yen, Thomson, et autres
Publié: (2026)
par: Yen, Thomson, et autres
Publié: (2026)
A Graph-Attentive LSTM Model for Malicious URL Detection
par: Hossain, Md. Ifthekhar, et autres
Publié: (2025)
par: Hossain, Md. Ifthekhar, et autres
Publié: (2025)
Riemann-Bench: A Benchmark for Moonshot Mathematics
par: Garre, Suhaas, et autres
Publié: (2026)
par: Garre, Suhaas, et autres
Publié: (2026)
JobBench: Aligning Agent Work With Human Will
par: Li, Yuetai, et autres
Publié: (2026)
par: Li, Yuetai, et autres
Publié: (2026)
Sales Research Agent and Sales Research Bench
par: Bhol, Deepanjan
Publié: (2025)
par: Bhol, Deepanjan
Publié: (2025)
REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning
par: Taechoyotin, Pawin, et autres
Publié: (2025)
par: Taechoyotin, Pawin, et autres
Publié: (2025)
Framing AI System Benchmarking as a Learning Task: FlexBench and the Open MLPerf Dataset
par: Fursin, Grigori, et autres
Publié: (2025)
par: Fursin, Grigori, et autres
Publié: (2025)
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
par: Zhao, Enyu, et autres
Publié: (2025)
par: Zhao, Enyu, et autres
Publié: (2025)
Automating Thought of Search: A Journey Towards Soundness and Completeness
par: Cao, Daniel, et autres
Publié: (2024)
par: Cao, Daniel, et autres
Publié: (2024)
ReEfBench: Quantifying the Reasoning Efficiency of LLMs
par: Fu, Zhizhang, et autres
Publié: (2026)
par: Fu, Zhizhang, et autres
Publié: (2026)
FIRE-Bench: Evaluating Agents on the Rediscovery of Scientific Insights
par: Wang, Zhen, et autres
Publié: (2026)
par: Wang, Zhen, et autres
Publié: (2026)
ConvexBench: Can LLMs Recognize Convex Functions?
par: Liu, Yepeng, et autres
Publié: (2026)
par: Liu, Yepeng, et autres
Publié: (2026)
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
par: Bai, Songlin, et autres
Publié: (2026)
par: Bai, Songlin, et autres
Publié: (2026)
CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics
par: Liu, Junqi, et autres
Publié: (2025)
par: Liu, Junqi, et autres
Publié: (2025)
Documents similaires
-
Automating Document Intelligence in Statutory City Planning
par: Malmqvist, Lars, et autres
Publié: (2026) -
Adopting Large Language Models to Automated System Integration
par: Pesl, Robin D.
Publié: (2025) -
Multistep Distillation of Diffusion Models via Moment Matching
par: Salimans, Tim, et autres
Publié: (2024) -
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
par: Jotautaitė, Monika, et autres
Publié: (2026) -
TaskBench: Benchmarking Large Language Models for Task Automation
par: Shen, Yongliang, et autres
Publié: (2023)