FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Qiran, Lam, Hou Hei, Zhao, Wenhao, Chen, Tingting, Tang, Yiming, Yu, Samson, Zhu, Yingtao, Anumasa, Srinivas, Zhang, Zufeng, Zhang, Tianyi, Liu, Chang, Jiang, Zhengyao, Goyal, Anirudh, Liu, Dianbo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FML-bench: Benchmarking Machine Learning Agents for Scientific Research
by: Zou, Qiran, et al.
Published: (2025)
by: Zou, Qiran, et al.
Published: (2025)
Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery
by: Chen, Tingting, et al.
Published: (2025)
by: Chen, Tingting, et al.
Published: (2025)
Multi-Novelty: Improve the Diversity and Novelty of Contents Generated by Large Language Models via inference-time Multi-Views Brainstorming
by: Lagzian, Arash, et al.
Published: (2025)
by: Lagzian, Arash, et al.
Published: (2025)
Data-Dependent Smoothing for Protein Discovery with Walk-Jump Sampling
by: Anumasa, Srinivas, et al.
Published: (2025)
by: Anumasa, Srinivas, et al.
Published: (2025)
Laplacian Score Sharpening for Mitigating Hallucination in Diffusion Models
by: C, Barath Chandran., et al.
Published: (2025)
by: C, Barath Chandran., et al.
Published: (2025)
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Representation Collapsing Problems in Vector Quantization
by: Zhao, Wenhao, et al.
Published: (2024)
by: Zhao, Wenhao, et al.
Published: (2024)
Mitigating Premature Discretization with Progressive Quantization for Robust Vector Tokenization
by: Zhao, Wenhao, et al.
Published: (2026)
by: Zhao, Wenhao, et al.
Published: (2026)
HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds
by: Chen, Tingting, et al.
Published: (2025)
by: Chen, Tingting, et al.
Published: (2025)
Early Quantization Shrinks Codebook: A Simple Fix for Diversity-Preserving Tokenization
by: Zhao, Wenhao, et al.
Published: (2026)
by: Zhao, Wenhao, et al.
Published: (2026)
Physical Reasoning and Object Planning for Household Embodied Agents
by: Agrawal, Ayush, et al.
Published: (2023)
by: Agrawal, Ayush, et al.
Published: (2023)
Uncertainty-Based Extensible Codebook for Discrete Federated Learning in Heterogeneous Data Silos
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
Navigating heterogeneous protein landscapes through geometry-aware smoothing
by: Anumasa, Srinivas, et al.
Published: (2026)
by: Anumasa, Srinivas, et al.
Published: (2026)
Comment on ‘Effects of nurse‐led self‐care interventions on health outcomes among people with heart failure: A systematic review and meta‐analysis’
by: Xiaoxia Liu, et al.
Published: (2024)
by: Xiaoxia Liu, et al.
Published: (2024)
PMFL: Partial Meta-Federated Learning for heterogeneous tasks and its applications on real-world medical records
by: Zhang, Tianyi, et al.
Published: (2021)
by: Zhang, Tianyi, et al.
Published: (2021)
VDRive: Leveraging Reinforced VLA and Diffusion Policy for End-to-end Autonomous Driving
by: Guo, Ziang, et al.
Published: (2025)
by: Guo, Ziang, et al.
Published: (2025)
AI-generated data contamination erodes pathological variability and diagnostic reliability
by: He, Hongyu, et al.
Published: (2026)
by: He, Hongyu, et al.
Published: (2026)
Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models
by: Wu, Yudi, et al.
Published: (2025)
by: Wu, Yudi, et al.
Published: (2025)
How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
by: Saini, Harshvardhan, et al.
Published: (2026)
by: Saini, Harshvardhan, et al.
Published: (2026)
OpFML: Pipeline for ML-based Operational Forecasting
by: Alvi, Shahbaz, et al.
Published: (2026)
by: Alvi, Shahbaz, et al.
Published: (2026)
BarlowTwins-CXR : Enhancing Chest X-Ray abnormality localization in heterogeneous data with cross-domain self-supervised learning
by: Sheng, Haoyue, et al.
Published: (2024)
by: Sheng, Haoyue, et al.
Published: (2024)
River-LLM: Large Language Model Seamless Exit Based on KV Share
by: Shen, Yingtao, et al.
Published: (2026)
by: Shen, Yingtao, et al.
Published: (2026)
SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
by: Antoniades, Antonis, et al.
Published: (2024)
by: Antoniades, Antonis, et al.
Published: (2024)
Enhance Eye Disease Detection using Learnable Probabilistic Discrete Latents in Machine Learning Architectures
by: Prabhakaran, Anirudh, et al.
Published: (2024)
by: Prabhakaran, Anirudh, et al.
Published: (2024)
CXR-LanIC: Language-Grounded Interpretable Classifier for Chest X-Ray Diagnosis
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
by: Saini, Harshvardhan, et al.
Published: (2026)
by: Saini, Harshvardhan, et al.
Published: (2026)
Difficult Examples Hurt Unsupervised Contrastive Learning: A Theoretical Perspective
by: Zhang, Yi-Ge, et al.
Published: (2025)
by: Zhang, Yi-Ge, et al.
Published: (2025)
LitSearch: A Retrieval Benchmark for Scientific Literature Search
by: Ajith, Anirudh, et al.
Published: (2024)
by: Ajith, Anirudh, et al.
Published: (2024)
Masked Generative Priors Improve World Models Sequence Modelling Capabilities
by: Meo, Cristian, et al.
Published: (2024)
by: Meo, Cristian, et al.
Published: (2024)
SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation
by: Vetcha, Nitin, et al.
Published: (2026)
by: Vetcha, Nitin, et al.
Published: (2026)
Interaction induced splitting of Dirac monopoles in the topological Thouless pumping of strongly interacting Bosons and SU($N$) Fermions
by: Lam, Hei, et al.
Published: (2024)
by: Lam, Hei, et al.
Published: (2024)
SWE-bench Goes Live!
by: Zhang, Linghao, et al.
Published: (2025)
by: Zhang, Linghao, et al.
Published: (2025)
ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
by: Luo, Yitong, et al.
Published: (2025)
by: Luo, Yitong, et al.
Published: (2025)
More diverse more adaptive: Comprehensive Multi-task Learning for Improved LLM Domain Adaptation in E-commerce
by: Piao, Tong, et al.
Published: (2025)
by: Piao, Tong, et al.
Published: (2025)
Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
by: Zan, Daoguang, et al.
Published: (2025)
by: Zan, Daoguang, et al.
Published: (2025)
Molecular dynamics simulations of the solubility of chitosan grafted polyacrylamide and its adsorption mechanism with kaolinite: Impact of length and distribution of branched‐chain
by: Wei Zhao, et al.
Published: (2024)
by: Wei Zhao, et al.
Published: (2024)
A Sudakov Decomposition in Riemannian Manifolds with Positive Curvature
by: Huang, Zhengyao
Published: (2026)
by: Huang, Zhengyao
Published: (2026)
A Unified Training Process for Fake News Detection based on Fine-Tuned BERT Model
by: Tida, Vijay Srinivas, et al.
Published: (2022)
by: Tida, Vijay Srinivas, et al.
Published: (2022)
FairFML: Fair Federated Machine Learning with a Case Study on Reducing Gender Disparities in Cardiac Arrest Outcome Prediction
by: Li, Siqi, et al.
Published: (2024)
by: Li, Siqi, et al.
Published: (2024)
Similar Items
-
FML-bench: Benchmarking Machine Learning Agents for Scientific Research
by: Zou, Qiran, et al.
Published: (2025) -
Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery
by: Chen, Tingting, et al.
Published: (2025) -
Multi-Novelty: Improve the Diversity and Novelty of Contents Generated by Large Language Models via inference-time Multi-Views Brainstorming
by: Lagzian, Arash, et al.
Published: (2025) -
Data-Dependent Smoothing for Protein Discovery with Walk-Jump Sampling
by: Anumasa, Srinivas, et al.
Published: (2025) -
Laplacian Score Sharpening for Mitigating Hallucination in Diffusion Models
by: C, Barath Chandran., et al.
Published: (2025)