FML-bench: Benchmarking Machine Learning Agents for Scientific Research
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zou, Qiran, Lam, Hou Hei, Zhao, Wenhao, Tang, Yiming, Chen, Tingting, Yu, Samson, Zhang, Tianyi, Liu, Chang, Ji, Xiangyang, Liu, Dianbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
von: Zou, Qiran, et al.
Veröffentlicht: (2026)
von: Zou, Qiran, et al.
Veröffentlicht: (2026)
Representation Collapsing Problems in Vector Quantization
von: Zhao, Wenhao, et al.
Veröffentlicht: (2024)
von: Zhao, Wenhao, et al.
Veröffentlicht: (2024)
Mitigating Premature Discretization with Progressive Quantization for Robust Vector Tokenization
von: Zhao, Wenhao, et al.
Veröffentlicht: (2026)
von: Zhao, Wenhao, et al.
Veröffentlicht: (2026)
Early Quantization Shrinks Codebook: A Simple Fix for Diversity-Preserving Tokenization
von: Zhao, Wenhao, et al.
Veröffentlicht: (2026)
von: Zhao, Wenhao, et al.
Veröffentlicht: (2026)
HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds
von: Chen, Tingting, et al.
Veröffentlicht: (2025)
von: Chen, Tingting, et al.
Veröffentlicht: (2025)
Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery
von: Chen, Tingting, et al.
Veröffentlicht: (2025)
von: Chen, Tingting, et al.
Veröffentlicht: (2025)
ParCo: Part-Coordinating Text-to-Motion Synthesis
von: Zou, Qiran, et al.
Veröffentlicht: (2024)
von: Zou, Qiran, et al.
Veröffentlicht: (2024)
How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models
von: Wu, Yudi, et al.
Veröffentlicht: (2025)
von: Wu, Yudi, et al.
Veröffentlicht: (2025)
Uncertainty-Based Extensible Codebook for Discrete Federated Learning in Heterogeneous Data Silos
von: Zhang, Tianyi, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2024)
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
von: Chan, Jun Shern, et al.
Veröffentlicht: (2024)
von: Chan, Jun Shern, et al.
Veröffentlicht: (2024)
CXR-LanIC: Language-Grounded Interpretable Classifier for Chest X-Ray Diagnosis
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
von: Toledo, Edan, et al.
Veröffentlicht: (2025)
von: Toledo, Edan, et al.
Veröffentlicht: (2025)
OpFML: Pipeline for ML-based Operational Forecasting
von: Alvi, Shahbaz, et al.
Veröffentlicht: (2026)
von: Alvi, Shahbaz, et al.
Veröffentlicht: (2026)
SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation
von: Vetcha, Nitin, et al.
Veröffentlicht: (2026)
von: Vetcha, Nitin, et al.
Veröffentlicht: (2026)
BarlowTwins-CXR : Enhancing Chest X-Ray abnormality localization in heterogeneous data with cross-domain self-supervised learning
von: Sheng, Haoyue, et al.
Veröffentlicht: (2024)
von: Sheng, Haoyue, et al.
Veröffentlicht: (2024)
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
STORI: A Benchmark and Taxonomy for Stochastic Environments
von: Barsainyan, Aryan Amit, et al.
Veröffentlicht: (2025)
von: Barsainyan, Aryan Amit, et al.
Veröffentlicht: (2025)
Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
von: Zan, Daoguang, et al.
Veröffentlicht: (2025)
von: Zan, Daoguang, et al.
Veröffentlicht: (2025)
CL-bench: A Benchmark for Context Learning
von: Dou, Shihan, et al.
Veröffentlicht: (2026)
von: Dou, Shihan, et al.
Veröffentlicht: (2026)
FairFML: Fair Federated Machine Learning with a Case Study on Reducing Gender Disparities in Cardiac Arrest Outcome Prediction
von: Li, Siqi, et al.
Veröffentlicht: (2024)
von: Li, Siqi, et al.
Veröffentlicht: (2024)
PMFL: Partial Meta-Federated Learning for heterogeneous tasks and its applications on real-world medical records
von: Zhang, Tianyi, et al.
Veröffentlicht: (2021)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2021)
LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth
von: Zeng, Weihao, et al.
Veröffentlicht: (2026)
von: Zeng, Weihao, et al.
Veröffentlicht: (2026)
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
von: Xiong, Lei, et al.
Veröffentlicht: (2026)
von: Xiong, Lei, et al.
Veröffentlicht: (2026)
Data-Dependent Smoothing for Protein Discovery with Walk-Jump Sampling
von: Anumasa, Srinivas, et al.
Veröffentlicht: (2025)
von: Anumasa, Srinivas, et al.
Veröffentlicht: (2025)
EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
von: Liu, Shaoyu, et al.
Veröffentlicht: (2025)
von: Liu, Shaoyu, et al.
Veröffentlicht: (2025)
Effect of Foam Filling and Holes on In‐Plane Compressive Failure Behavior of FML Face Sheet Grid Sandwich Structures
von: Fengling Zhao, et al.
Veröffentlicht: (2026)
von: Fengling Zhao, et al.
Veröffentlicht: (2026)
CodeUnlearn: Amortized Zero-Shot Machine Unlearning in Language Models Using Discrete Concept
von: Wu, YuXuan, et al.
Veröffentlicht: (2024)
von: Wu, YuXuan, et al.
Veröffentlicht: (2024)
$τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
von: Lee, Hwiwon, et al.
Veröffentlicht: (2025)
von: Lee, Hwiwon, et al.
Veröffentlicht: (2025)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
von: Zhao, Haiyan, et al.
Veröffentlicht: (2025)
SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
von: Zan, Daoguang, et al.
Veröffentlicht: (2024)
von: Zan, Daoguang, et al.
Veröffentlicht: (2024)
TimeMachine-bench: A Benchmark for Evaluating Model Capabilities in Repository-Level Migration Tasks
von: Fujii, Ryo, et al.
Veröffentlicht: (2026)
von: Fujii, Ryo, et al.
Veröffentlicht: (2026)
AI Research Agents Narrow Scientific Exploration
von: Tang, Yixuan, et al.
Veröffentlicht: (2026)
von: Tang, Yixuan, et al.
Veröffentlicht: (2026)
Physical Reasoning and Object Planning for Household Embodied Agents
von: Agrawal, Ayush, et al.
Veröffentlicht: (2023)
von: Agrawal, Ayush, et al.
Veröffentlicht: (2023)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
Enhance Eye Disease Detection using Learnable Probabilistic Discrete Latents in Machine Learning Architectures
von: Prabhakaran, Anirudh, et al.
Veröffentlicht: (2024)
von: Prabhakaran, Anirudh, et al.
Veröffentlicht: (2024)
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
von: Ai, Kuangshi, et al.
Veröffentlicht: (2026)
von: Ai, Kuangshi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
von: Zou, Qiran, et al.
Veröffentlicht: (2026) -
Representation Collapsing Problems in Vector Quantization
von: Zhao, Wenhao, et al.
Veröffentlicht: (2024) -
Mitigating Premature Discretization with Progressive Quantization for Robust Vector Tokenization
von: Zhao, Wenhao, et al.
Veröffentlicht: (2026) -
Early Quantization Shrinks Codebook: A Simple Fix for Diversity-Preserving Tokenization
von: Zhao, Wenhao, et al.
Veröffentlicht: (2026) -
HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds
von: Chen, Tingting, et al.
Veröffentlicht: (2025)