MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yunxiang, Khalifa, Muhammad, Bhushan, Shitanshu, Murphy, Grant D, Logeswaran, Lajanugen, Kim, Jaekyeom, Lee, Moontae, Lee, Honglak, Wang, Lu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
di: Zhang, Yunxiang, et al.
Pubblicazione: (2024)
di: Zhang, Yunxiang, et al.
Pubblicazione: (2024)
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
di: Khalifa, Muhammad, et al.
Pubblicazione: (2026)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2026)
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
di: Khalifa, Muhammad, et al.
Pubblicazione: (2023)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2023)
Process Reward Models That Think
di: Khalifa, Muhammad, et al.
Pubblicazione: (2025)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2025)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
di: Kim, Jaekyeom, et al.
Pubblicazione: (2024)
di: Kim, Jaekyeom, et al.
Pubblicazione: (2024)
Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
di: Logeswaran, Lajanugen, et al.
Pubblicazione: (2026)
di: Logeswaran, Lajanugen, et al.
Pubblicazione: (2026)
Selective LoRA for Visual Tokens and Attention Heads
di: Luo, Tiange, et al.
Pubblicazione: (2025)
di: Luo, Tiange, et al.
Pubblicazione: (2025)
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
di: Fu, Yao, et al.
Pubblicazione: (2024)
di: Fu, Yao, et al.
Pubblicazione: (2024)
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
di: Shen, Siqi, et al.
Pubblicazione: (2024)
di: Shen, Siqi, et al.
Pubblicazione: (2024)
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
di: Shen, Siqi, et al.
Pubblicazione: (2025)
di: Shen, Siqi, et al.
Pubblicazione: (2025)
Visual Test-time Scaling for GUI Agent Grounding
di: Luo, Tiange, et al.
Pubblicazione: (2025)
di: Luo, Tiange, et al.
Pubblicazione: (2025)
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
di: Zhang, Lechen, et al.
Pubblicazione: (2024)
di: Zhang, Lechen, et al.
Pubblicazione: (2024)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
di: Zheng, Mingqian, et al.
Pubblicazione: (2023)
di: Zheng, Mingqian, et al.
Pubblicazione: (2023)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
di: Zhang, Lechen, et al.
Pubblicazione: (2025)
di: Zhang, Lechen, et al.
Pubblicazione: (2025)
LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
di: Zou, Kaijian, et al.
Pubblicazione: (2025)
di: Zou, Kaijian, et al.
Pubblicazione: (2025)
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
di: Jang, Yunseok, et al.
Pubblicazione: (2025)
di: Jang, Yunseok, et al.
Pubblicazione: (2025)
You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments
di: Shu, Bangzhao, et al.
Pubblicazione: (2023)
di: Shu, Bangzhao, et al.
Pubblicazione: (2023)
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
di: Han, Janghoon, et al.
Pubblicazione: (2025)
di: Han, Janghoon, et al.
Pubblicazione: (2025)
Towards Diverse Evaluation of Class Incremental Learning: A Representation Learning Perspective
di: Cha, Sungmin, et al.
Pubblicazione: (2022)
di: Cha, Sungmin, et al.
Pubblicazione: (2022)
Learning to Unlearn: Instance-wise Unlearning for Pre-trained Classifiers
di: Cha, Sungmin, et al.
Pubblicazione: (2023)
di: Cha, Sungmin, et al.
Pubblicazione: (2023)
Source-Aware Training Enables Knowledge Attribution in Language Models
di: Khalifa, Muhammad, et al.
Pubblicazione: (2024)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2024)
Interactive and Expressive Code-Augmented Planning with Large Language Models
di: Liu, Anthony Z., et al.
Pubblicazione: (2024)
di: Liu, Anthony Z., et al.
Pubblicazione: (2024)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
di: Kim, Geon-Hyeong, et al.
Pubblicazione: (2025)
di: Kim, Geon-Hyeong, et al.
Pubblicazione: (2025)
Do Not Trust Licenses You See: Dataset Compliance Requires Massive-Scale AI-Powered Lifecycle Tracing
di: Kim, Jaekyeom, et al.
Pubblicazione: (2025)
di: Kim, Jaekyeom, et al.
Pubblicazione: (2025)
Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines
di: Carvalho, Wilka, et al.
Pubblicazione: (2025)
di: Carvalho, Wilka, et al.
Pubblicazione: (2025)
If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs
di: Khalifa, Muhammad, et al.
Pubblicazione: (2024)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2024)
When Is Enough Not Enough? Illusory Completion in Search Agents
di: Ko, Dayoon, et al.
Pubblicazione: (2026)
di: Ko, Dayoon, et al.
Pubblicazione: (2026)
Training-free Detection of AI-generated images via Cropping Robustness
di: Choi, Sungik, et al.
Pubblicazione: (2025)
di: Choi, Sungik, et al.
Pubblicazione: (2025)
Practical and Reproducible Symbolic Music Generation by Large Language Models with Structural Embeddings
di: Rhyu, Seungyeon, et al.
Pubblicazione: (2024)
di: Rhyu, Seungyeon, et al.
Pubblicazione: (2024)
Learning to Ideate for Machine Learning Engineering Agents
di: Zhang, Yunxiang, et al.
Pubblicazione: (2026)
di: Zhang, Yunxiang, et al.
Pubblicazione: (2026)
LG AI Research & KAIST at EHRSQL 2024: Self-Training Large Language Models with Pseudo-Labeled Unanswerable Questions for a Reliable Text-to-SQL System on EHRs
di: Jo, Yongrae, et al.
Pubblicazione: (2024)
di: Jo, Yongrae, et al.
Pubblicazione: (2024)
AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
di: Ok, Hyunjong, et al.
Pubblicazione: (2025)
di: Ok, Hyunjong, et al.
Pubblicazione: (2025)
Schützende Bewältigung
di: Logeswaran, Araththy
Pubblicazione: (2022)
di: Logeswaran, Araththy
Pubblicazione: (2022)
Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language Models
di: Kim, Jiyeon, et al.
Pubblicazione: (2026)
di: Kim, Jiyeon, et al.
Pubblicazione: (2026)
Can Vision-Language Models Solve the Shell Game?
di: Liu, Tiedong, et al.
Pubblicazione: (2026)
di: Liu, Tiedong, et al.
Pubblicazione: (2026)
Probing Visual Language Priors in VLMs
di: Luo, Tiange, et al.
Pubblicazione: (2024)
di: Luo, Tiange, et al.
Pubblicazione: (2024)
FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?
di: Mittal, Chinmay, et al.
Pubblicazione: (2024)
di: Mittal, Chinmay, et al.
Pubblicazione: (2024)
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
di: Yoran, Ori, et al.
Pubblicazione: (2024)
di: Yoran, Ori, et al.
Pubblicazione: (2024)
LitCab: Lightweight Language Model Calibration over Short- and Long-form Responses
di: Liu, Xin, et al.
Pubblicazione: (2023)
di: Liu, Xin, et al.
Pubblicazione: (2023)
Counterfactual Voting Adjustment for Quality Assessment and Fairer Voting in Online Platforms with Helpfulness Evaluation
di: Liu, Chang, et al.
Pubblicazione: (2025)
di: Liu, Chang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
di: Zhang, Yunxiang, et al.
Pubblicazione: (2024) -
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
di: Khalifa, Muhammad, et al.
Pubblicazione: (2026) -
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
di: Khalifa, Muhammad, et al.
Pubblicazione: (2023) -
Process Reward Models That Think
di: Khalifa, Muhammad, et al.
Pubblicazione: (2025) -
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
di: Kim, Jaekyeom, et al.
Pubblicazione: (2024)