MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Hui, Xiong, Miao, Lu, Yujie, Han, Wei, Deng, Ailin, He, Yufei, Wu, Jiaying, Li, Yibo, Liu, Yue, Hooi, Bryan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ConfTuner: Training Large Language Models to Express Their Confidence Verbally
por: Li, Yibo, et al.
Publicado: (2025)
por: Li, Yibo, et al.
Publicado: (2025)
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
por: Li, Yibo, et al.
Publicado: (2026)
por: Li, Yibo, et al.
Publicado: (2026)
Proximity-Informed Calibration for Deep Neural Networks
por: Xiong, Miao, et al.
Publicado: (2023)
por: Xiong, Miao, et al.
Publicado: (2023)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
por: Deng, Ailin, et al.
Publicado: (2024)
por: Deng, Ailin, et al.
Publicado: (2024)
Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?
por: He, Yufei, et al.
Publicado: (2025)
por: He, Yufei, et al.
Publicado: (2025)
Words or Vision: Do Vision-Language Models Have Blind Faith in Text?
por: Deng, Ailin, et al.
Publicado: (2025)
por: Deng, Ailin, et al.
Publicado: (2025)
Towards Realistic Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions
por: Guo, Qianyun, et al.
Publicado: (2026)
por: Guo, Qianyun, et al.
Publicado: (2026)
Can Knowledge Graphs Make Large Language Models More Trustworthy? An Empirical Study Over Open-ended Question Answering
por: Sui, Yuan, et al.
Publicado: (2024)
por: Sui, Yuan, et al.
Publicado: (2024)
ID$^3$: Identity-Preserving-yet-Diversified Diffusion Models for Synthetic Face Recognition
por: Li, Shen, et al.
Publicado: (2024)
por: Li, Shen, et al.
Publicado: (2024)
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
por: Lou, Zhanzhi, et al.
Publicado: (2026)
por: Lou, Zhanzhi, et al.
Publicado: (2026)
MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents
por: Li, Ruochen, et al.
Publicado: (2024)
por: Li, Ruochen, et al.
Publicado: (2024)
Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style Attacks
por: Wu, Jiaying, et al.
Publicado: (2023)
por: Wu, Jiaying, et al.
Publicado: (2023)
APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents
por: Li, Yibo, et al.
Publicado: (2026)
por: Li, Yibo, et al.
Publicado: (2026)
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
por: Yang, Xianglin, et al.
Publicado: (2026)
por: Yang, Xianglin, et al.
Publicado: (2026)
MCU: An Evaluation Framework for Open-Ended Game Agents
por: Zheng, Xinyue, et al.
Publicado: (2023)
por: Zheng, Xinyue, et al.
Publicado: (2023)
UniGraph: Learning a Unified Cross-Domain Foundation Model for Text-Attributed Graphs
por: He, Yufei, et al.
Publicado: (2024)
por: He, Yufei, et al.
Publicado: (2024)
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
por: Cao, Tri, et al.
Publicado: (2025)
por: Cao, Tri, et al.
Publicado: (2025)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
por: Xiong, Miao, et al.
Publicado: (2023)
por: Xiong, Miao, et al.
Publicado: (2023)
TACT: Mitigating Overthinking and Overacting in Coding Agents via Activation Steering
por: Sui, Yuan, et al.
Publicado: (2026)
por: Sui, Yuan, et al.
Publicado: (2026)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
por: Chen, Yulin, et al.
Publicado: (2026)
por: Chen, Yulin, et al.
Publicado: (2026)
Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
por: Zhang, Jenny, et al.
Publicado: (2025)
por: Zhang, Jenny, et al.
Publicado: (2025)
FlowReasoner: Reinforcing Query-Level Meta-Agents
por: Gao, Hongcheng, et al.
Publicado: (2025)
por: Gao, Hongcheng, et al.
Publicado: (2025)
EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
por: He, Yufei, et al.
Publicado: (2025)
por: He, Yufei, et al.
Publicado: (2025)
UniGraph2: Learning a Unified Embedding Space to Bind Multimodal Graphs
por: He, Yufei, et al.
Publicado: (2025)
por: He, Yufei, et al.
Publicado: (2025)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
por: Cao, Tri, et al.
Publicado: (2026)
por: Cao, Tri, et al.
Publicado: (2026)
Enhancing Multi-Agent Debate System Performance via Confidence Expression
por: Lin, Zijie, et al.
Publicado: (2025)
por: Lin, Zijie, et al.
Publicado: (2025)
FlipAttack: Jailbreak LLMs via Flipping
por: Liu, Yue, et al.
Publicado: (2024)
por: Liu, Yue, et al.
Publicado: (2024)
AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
por: Ying, Lance, et al.
Publicado: (2026)
por: Ying, Lance, et al.
Publicado: (2026)
InfoQuest: Evaluating Multi-Turn Dialogue Agents for Open-Ended Conversations with Hidden Context
por: de Oliveira, Bryan L. M., et al.
Publicado: (2025)
por: de Oliveira, Bryan L. M., et al.
Publicado: (2025)
NTSFormer: A Self-Teaching Graph Transformer for Multimodal Isolated Cold-Start Node Classification
por: Hu, Jun, et al.
Publicado: (2025)
por: Hu, Jun, et al.
Publicado: (2025)
AEL: Agent Evolving Learning for Open-Ended Environments
por: Xu, Wujiang, et al.
Publicado: (2026)
por: Xu, Wujiang, et al.
Publicado: (2026)
Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation
por: Chen, Nuo, et al.
Publicado: (2026)
por: Chen, Nuo, et al.
Publicado: (2026)
Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation
por: Sui, Yuan, et al.
Publicado: (2026)
por: Sui, Yuan, et al.
Publicado: (2026)
Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
por: Sui, Yuan, et al.
Publicado: (2025)
por: Sui, Yuan, et al.
Publicado: (2025)
KLong: Training LLM Agent for Extremely Long-horizon Tasks
por: Liu, Yue, et al.
Publicado: (2026)
por: Liu, Yue, et al.
Publicado: (2026)
AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents
por: Somasekharan, Nithin, et al.
Publicado: (2026)
por: Somasekharan, Nithin, et al.
Publicado: (2026)
Robust Agents in Open-Ended Worlds
por: Samvelyan, Mikayel
Publicado: (2025)
por: Samvelyan, Mikayel
Publicado: (2025)
AgentCPM-Report: Interleaving Drafting and Deepening for Open-Ended Deep Research
por: Li, Yishan, et al.
Publicado: (2026)
por: Li, Yishan, et al.
Publicado: (2026)
EvoClinician: A Self-Evolving Agent for Multi-Turn Medical Diagnosis via Test-Time Evolutionary Learning
por: He, Yufei, et al.
Publicado: (2026)
por: He, Yufei, et al.
Publicado: (2026)
CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery
por: Qu, Ao, et al.
Publicado: (2026)
por: Qu, Ao, et al.
Publicado: (2026)
Ejemplares similares
-
ConfTuner: Training Large Language Models to Express Their Confidence Verbally
por: Li, Yibo, et al.
Publicado: (2025) -
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
por: Li, Yibo, et al.
Publicado: (2026) -
Proximity-Informed Calibration for Deep Neural Networks
por: Xiong, Miao, et al.
Publicado: (2023) -
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
por: Deng, Ailin, et al.
Publicado: (2024) -
Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?
por: He, Yufei, et al.
Publicado: (2025)