ML Research Benchmark
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Kenney, Matthew |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
par: Gan, Eric, et autres
Publié: (2026)
par: Gan, Eric, et autres
Publié: (2026)
Benchmarking Edge AI Platforms for High-Performance ML Inference
par: Jayanth, Rakshith, et autres
Publié: (2024)
par: Jayanth, Rakshith, et autres
Publié: (2024)
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
par: Atinafu, Yonas, et autres
Publié: (2026)
par: Atinafu, Yonas, et autres
Publié: (2026)
Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization
par: Jia, Hangyi, et autres
Publié: (2025)
par: Jia, Hangyi, et autres
Publié: (2025)
The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
par: Chung, Jae-Won, et autres
Publié: (2025)
par: Chung, Jae-Won, et autres
Publié: (2025)
Wake Vision: A Tailored Dataset and Benchmark Suite for TinyML Computer Vision Applications
par: Banbury, Colby, et autres
Publié: (2024)
par: Banbury, Colby, et autres
Publié: (2024)
FormalML: A Benchmark for Evaluating Formal Subgoal Completion in Machine Learning Theory
par: Yang, Xiao-Wen, et autres
Publié: (2025)
par: Yang, Xiao-Wen, et autres
Publié: (2025)
Design Principles for Falsifiable, Replicable and Reproducible Empirical ML Research
par: Vranješ, Daniel, et autres
Publié: (2024)
par: Vranješ, Daniel, et autres
Publié: (2024)
HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI
par: Pricope, Tidor-Vlad
Publié: (2025)
par: Pricope, Tidor-Vlad
Publié: (2025)
Codenames as a Benchmark for Large Language Models
par: Stephenson, Matthew, et autres
Publié: (2024)
par: Stephenson, Matthew, et autres
Publié: (2024)
Pre-Hoc Predictions in AutoML: Leveraging LLMs to Enhance Model Selection and Benchmarking for Tabular datasets
par: Belkhiter, Yannis, et autres
Publié: (2025)
par: Belkhiter, Yannis, et autres
Publié: (2025)
AdaptoML-UX: An Adaptive User-centered GUI-based AutoML Toolkit for Non-AI Experts and HCI Researchers
par: Gomaa, Amr, et autres
Publié: (2024)
par: Gomaa, Amr, et autres
Publié: (2024)
AutoML Systems For Medical Imaging
par: Jidney, Tasmia Tahmida, et autres
Publié: (2023)
par: Jidney, Tasmia Tahmida, et autres
Publié: (2023)
Towards Knowledgeable Deep Research: Framework and Benchmark
par: Liu, Wenxuan, et autres
Publié: (2026)
par: Liu, Wenxuan, et autres
Publié: (2026)
ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
par: Chittepu, Yaswanth, et autres
Publié: (2025)
par: Chittepu, Yaswanth, et autres
Publié: (2025)
ZeroML: A Next Generation AutoML Language
par: Mahmud, Monirul Islam
Publié: (2025)
par: Mahmud, Monirul Islam
Publié: (2025)
Certified ML Object Detection for Surveillance Missions
par: Belcaid, Mohammed, et autres
Publié: (2024)
par: Belcaid, Mohammed, et autres
Publié: (2024)
pAI/MSc: ML Theory Research with Humans on the Loop
par: Abdelmoneum, Mahmoud, et autres
Publié: (2026)
par: Abdelmoneum, Mahmoud, et autres
Publié: (2026)
Implementation of airborne ML models with semantics preservation
par: Valot, Nicolas, et autres
Publié: (2025)
par: Valot, Nicolas, et autres
Publié: (2025)
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
par: Wang, Jiayu, et autres
Publié: (2025)
par: Wang, Jiayu, et autres
Publié: (2025)
DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent
par: Wu, Tongzhou, et autres
Publié: (2026)
par: Wu, Tongzhou, et autres
Publié: (2026)
Benchmark Transparency: Measuring the Impact of Data on Evaluation
par: Kovatchev, Venelin, et autres
Publié: (2024)
par: Kovatchev, Venelin, et autres
Publié: (2024)
ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows
par: Padigela, Harshith, et autres
Publié: (2025)
par: Padigela, Harshith, et autres
Publié: (2025)
LoCoML: A Framework for Real-World ML Inference Pipelines
par: Maddireddy, Kritin, et autres
Publié: (2025)
par: Maddireddy, Kritin, et autres
Publié: (2025)
SafePickle: Robust and Generic ML Detection of Malicious Pickle-based ML Models
par: Ohayon, Hillel, et autres
Publié: (2026)
par: Ohayon, Hillel, et autres
Publié: (2026)
Re$^2$Math: Benchmarking Theorem Retrieval in Research-Level Mathematics
par: Lyu, Zicheng, et autres
Publié: (2026)
par: Lyu, Zicheng, et autres
Publié: (2026)
DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
par: Huang, Yukun, et autres
Publié: (2026)
par: Huang, Yukun, et autres
Publié: (2026)
CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
par: Queen, Owen, et autres
Publié: (2025)
par: Queen, Owen, et autres
Publié: (2025)
QMBench: A Research Level Benchmark for Quantum Materials Research
par: Wang, Yanzhen, et autres
Publié: (2025)
par: Wang, Yanzhen, et autres
Publié: (2025)
Accelerating IoV Intrusion Detection: Benchmarking GPU-Accelerated vs CPU-Based ML Libraries
par: Çolhak, Furkan, et autres
Publié: (2025)
par: Çolhak, Furkan, et autres
Publié: (2025)
CubicML: Automated ML for Large ML Systems Co-design with ML Prediction of Performance
par: Wen, Wei, et autres
Publié: (2024)
par: Wen, Wei, et autres
Publié: (2024)
Redefining Finance: The Influence of Artificial Intelligence (AI) and Machine Learning (ML)
par: Kumar, Animesh
Publié: (2024)
par: Kumar, Animesh
Publié: (2024)
More Questions than Answers? Lessons from Integrating Explainable AI into a Cyber-AI Tool
par: Suh, Ashley, et autres
Publié: (2024)
par: Suh, Ashley, et autres
Publié: (2024)
CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers
par: Pan, Haining, et autres
Publié: (2025)
par: Pan, Haining, et autres
Publié: (2025)
NarraBench: A Comprehensive Framework for Narrative Benchmarking
par: Hamilton, Sil, et autres
Publié: (2025)
par: Hamilton, Sil, et autres
Publié: (2025)
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
par: Xiong, Lei, et autres
Publié: (2026)
par: Xiong, Lei, et autres
Publié: (2026)
How to design a dataset compliant with an ML-based system ODD?
par: Cappi, Cyril, et autres
Publié: (2024)
par: Cappi, Cyril, et autres
Publié: (2024)
ML-SceGen: A Multi-level Scenario Generation Framework
par: Xiao, Yicheng, et autres
Publié: (2025)
par: Xiao, Yicheng, et autres
Publié: (2025)
Automatic Mapping of AutomationML Files to Ontologies for Graph Queries and Validation
par: Westermann, Tom, et autres
Publié: (2025)
par: Westermann, Tom, et autres
Publié: (2025)
SCUBA: Salesforce Computer Use Benchmark
par: Dai, Yutong, et autres
Publié: (2025)
par: Dai, Yutong, et autres
Publié: (2025)
Documents similaires
-
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
par: Gan, Eric, et autres
Publié: (2026) -
Benchmarking Edge AI Platforms for High-Performance ML Inference
par: Jayanth, Rakshith, et autres
Publié: (2024) -
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
par: Atinafu, Yonas, et autres
Publié: (2026) -
Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization
par: Jia, Hangyi, et autres
Publié: (2025) -
The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
par: Chung, Jae-Won, et autres
Publié: (2025)