Automating Benchmark Design
Fuente:
arXiv
Saved in:
| Main Authors: | Dsouza, Amanda, Vishwakarma, Harit, Qi, Zhengyang, Bauer, Justin, Pham, Derek, Walshe, Thomas, Parchami, Armin, Sala, Frederic, Varma, Paroma |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
by: Bauer, Justin, et al.
Published: (2026)
by: Bauer, Justin, et al.
Published: (2026)
RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics
by: Qi, Zhengyang, et al.
Published: (2026)
by: Qi, Zhengyang, et al.
Published: (2026)
EnvBench: A Benchmark for Automated Environment Setup
by: Eliseeva, Aleksandra, et al.
Published: (2025)
by: Eliseeva, Aleksandra, et al.
Published: (2025)
Can LLMs Generate Architectural Design Decisions? -An Exploratory Empirical study
by: Dhar, Rudra, et al.
Published: (2024)
by: Dhar, Rudra, et al.
Published: (2024)
Graph-Free Root Cause Analysis
by: Pham, Luan
Published: (2026)
by: Pham, Luan
Published: (2026)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
LLM-Enhanced Log Anomaly Detection: A Comprehensive Benchmark of Large Language Models for Automated System Diagnostics
by: Patel, Disha
Published: (2026)
by: Patel, Disha
Published: (2026)
DRAFT-ing Architectural Design Decisions using LLMs
by: Dhar, Rudra, et al.
Published: (2025)
by: Dhar, Rudra, et al.
Published: (2025)
RocketPPA: Code-Level Power, Performance, and Area Prediction via LLM and Mixture of Experts
by: Abdollahi, Armin, et al.
Published: (2025)
by: Abdollahi, Armin, et al.
Published: (2025)
You Only Train Once: A Flexible Training Framework for Code Vulnerability Detection Driven by Vul-Vector
by: Tian, Bowen, et al.
Published: (2025)
by: Tian, Bowen, et al.
Published: (2025)
Automated Machine Learning: A Case Study on Non-Intrusive Appliance Load Monitoring
by: Moin, Armin, et al.
Published: (2022)
by: Moin, Armin, et al.
Published: (2022)
Enhancing Architecture Frameworks by Including Modern Stakeholders and their Views/Viewpoints
by: Moin, Armin, et al.
Published: (2023)
by: Moin, Armin, et al.
Published: (2023)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
by: Kuntz, Thomas, et al.
Published: (2025)
by: Kuntz, Thomas, et al.
Published: (2025)
Monitizer: Automating Design and Evaluation of Neural Network Monitors
by: Azeem, Muqsit, et al.
Published: (2024)
by: Azeem, Muqsit, et al.
Published: (2024)
Automated Trustworthiness Testing for Machine Learning Classifiers
by: Cho, Steven, et al.
Published: (2024)
by: Cho, Steven, et al.
Published: (2024)
Automating Formal Verification with Reinforcement Learning and Recursive Inference
by: Tan, Max
Published: (2026)
by: Tan, Max
Published: (2026)
Automated Modernization of Machine Learning Engineering Notebooks for Reproducibility
by: Jin, Bihui, et al.
Published: (2026)
by: Jin, Bihui, et al.
Published: (2026)
Integrating Large Language Models for Automated Structural Analysis
by: Liang, Haoran, et al.
Published: (2025)
by: Liang, Haoran, et al.
Published: (2025)
Automating API Documentation with LLMs: A BERTopic Approach
by: Naghshzan, AmirHossein
Published: (2025)
by: Naghshzan, AmirHossein
Published: (2025)
AutoSpec: Automated Generation of Neural Network Specifications
by: Jin, Shuowei, et al.
Published: (2024)
by: Jin, Shuowei, et al.
Published: (2024)
The Limits of Long-Context Reasoning in Automated Bug Fixing
by: Raju, Ravi, et al.
Published: (2026)
by: Raju, Ravi, et al.
Published: (2026)
Polynomiogram: An Integrated Framework for Root Visualization and Generative Art
by: Nguyen, Hoang Duc, et al.
Published: (2025)
by: Nguyen, Hoang Duc, et al.
Published: (2025)
Towards Predicting Multi-Vulnerability Attack Chains in Software Supply Chains from Software Bill of Materials Graphs
by: Baird, Laura, et al.
Published: (2026)
by: Baird, Laura, et al.
Published: (2026)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
by: Huang, Tzu-Heng, et al.
Published: (2025)
by: Huang, Tzu-Heng, et al.
Published: (2025)
Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures
by: Lacroix, Nicolas, et al.
Published: (2026)
by: Lacroix, Nicolas, et al.
Published: (2026)
CA2: Code-Aware Agent for Automated Game Testing
by: Adaikkappan, Valliappan Chidambaram, et al.
Published: (2026)
by: Adaikkappan, Valliappan Chidambaram, et al.
Published: (2026)
Automating REST API Postman Test Cases Using LLM
by: Sri, S Deepika, et al.
Published: (2024)
by: Sri, S Deepika, et al.
Published: (2024)
QualiTagger: Automating software quality detection in issue trackers
by: Shivashankar, Karthik, et al.
Published: (2025)
by: Shivashankar, Karthik, et al.
Published: (2025)
pyAKI -- An Open Source Solution to Automated KDIGO classification
by: Porschen, Christian, et al.
Published: (2024)
by: Porschen, Christian, et al.
Published: (2024)
Teaching an Online Multi-Institutional Research Level Software Engineering Course with Industry -- an Experience Report
by: Jalote, Pankaj, et al.
Published: (2025)
by: Jalote, Pankaj, et al.
Published: (2025)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
by: Jiang, Yuancheng, et al.
Published: (2025)
by: Jiang, Yuancheng, et al.
Published: (2025)
Exploring Code Language Models for Automated HLS-based Hardware Generation: Benchmark, Infrastructure and Analysis
by: Gai, Jiahao, et al.
Published: (2025)
by: Gai, Jiahao, et al.
Published: (2025)
Cataloguing Hugging Face Models to Software Engineering Activities: Automation and Findings
by: González, Alexandra, et al.
Published: (2025)
by: González, Alexandra, et al.
Published: (2025)
Automated Program Repair: Emerging trends pose and expose problems for benchmarks
by: Renzullo, Joseph, et al.
Published: (2024)
by: Renzullo, Joseph, et al.
Published: (2024)
Constrained Adversarial Learning for Automated Software Testing: a literature review
by: Vitorino, João, et al.
Published: (2023)
by: Vitorino, João, et al.
Published: (2023)
PrismaDV: Automated Task-Aware Data Unit Test Generation
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
Automating the Training and Deployment of Models in MLOps by Integrating Systems with Machine Learning
by: Liang, Penghao, et al.
Published: (2024)
by: Liang, Penghao, et al.
Published: (2024)
ChIRAAG: ChatGPT Informed Rapid and Automated Assertion Generation
by: Mali, Bhabesh, et al.
Published: (2024)
by: Mali, Bhabesh, et al.
Published: (2024)
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
by: Li, Yuangang, et al.
Published: (2026)
by: Li, Yuangang, et al.
Published: (2026)
Similar Items
-
Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
by: Bauer, Justin, et al.
Published: (2026) -
RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics
by: Qi, Zhengyang, et al.
Published: (2026) -
EnvBench: A Benchmark for Automated Environment Setup
by: Eliseeva, Aleksandra, et al.
Published: (2025) -
Can LLMs Generate Architectural Design Decisions? -An Exploratory Empirical study
by: Dhar, Rudra, et al.
Published: (2024) -
Graph-Free Root Cause Analysis
by: Pham, Luan
Published: (2026)