LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yan, Shuo, Li, Ruochen, Luo, Ziming, Wang, Zimu, Li, Daoyang, Jing, Liqiang, He, Kaiyu, Wu, Peilin, Michalopoulos, George, Zhang, Yue, Zhang, Ziyang, Zhang, Mian, Chen, Zhiyu, Du, Xinya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
IDEA: Enhancing the Rule Learning Ability of Large Language Model Agent through Induction, Deduction, and Abduction
von: He, Kaiyu, et al.
Veröffentlicht: (2024)
von: He, Kaiyu, et al.
Veröffentlicht: (2024)
Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers
von: He, Kaiyu, et al.
Veröffentlicht: (2026)
von: He, Kaiyu, et al.
Veröffentlicht: (2026)
GEAR: A General Evaluation Framework for Abductive Reasoning
von: He, Kaiyu, et al.
Veröffentlicht: (2025)
von: He, Kaiyu, et al.
Veröffentlicht: (2025)
HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
LDC: Learning to Generate Research Idea with Dynamic Control
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control
von: Li, Ruosen, et al.
Veröffentlicht: (2025)
von: Li, Ruosen, et al.
Veröffentlicht: (2025)
MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2023)
von: Jing, Liqiang, et al.
Veröffentlicht: (2023)
FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning
von: Li, Ruosen, et al.
Veröffentlicht: (2024)
von: Li, Ruosen, et al.
Veröffentlicht: (2024)
Do Retrieval-Augmented Language Models Adapt to Varying User Needs?
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
von: Wu, Peilin, et al.
Veröffentlicht: (2025)
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering
von: Li, Ruosen, et al.
Veröffentlicht: (2024)
von: Li, Ruosen, et al.
Veröffentlicht: (2024)
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs
von: Yan, Bowen, et al.
Veröffentlicht: (2024)
von: Yan, Bowen, et al.
Veröffentlicht: (2024)
AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
von: Wu, Peilin, et al.
Veröffentlicht: (2026)
von: Wu, Peilin, et al.
Veröffentlicht: (2026)
Document-level Causal Relation Extraction with Knowledge-guided Binary Question Answering
von: Wang, Zimu, et al.
Veröffentlicht: (2024)
von: Wang, Zimu, et al.
Veröffentlicht: (2024)
LLM4SR: A Survey on Large Language Models for Scientific Research
von: Luo, Ziming, et al.
Veröffentlicht: (2025)
von: Luo, Ziming, et al.
Veröffentlicht: (2025)
PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment
von: Li, Yantao, et al.
Veröffentlicht: (2026)
von: Li, Yantao, et al.
Veröffentlicht: (2026)
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
von: Li, Xiaozhe, et al.
Veröffentlicht: (2025)
von: Li, Xiaozhe, et al.
Veröffentlicht: (2025)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
M3-BENCH: Process-Aware Evaluation of LLM Agents' Social Behaviors in Mixed-Motive Games
von: Xie, Sixiong, et al.
Veröffentlicht: (2026)
von: Xie, Sixiong, et al.
Veröffentlicht: (2026)
DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
von: Li, Xiaozhe, et al.
Veröffentlicht: (2026)
von: Li, Xiaozhe, et al.
Veröffentlicht: (2026)
A Dynamic Principal Agent Problem with One-sided Commitment
von: Zhang, Jianfeng, et al.
Veröffentlicht: (2022)
von: Zhang, Jianfeng, et al.
Veröffentlicht: (2022)
Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research
von: Zhang, Qianqian, et al.
Veröffentlicht: (2025)
von: Zhang, Qianqian, et al.
Veröffentlicht: (2025)
When Less Latent Leads to Better Relay: Information-Preserving Compression for Latent Multi-Agent LLM Collaboration
von: Li, Yiping, et al.
Veröffentlicht: (2026)
von: Li, Yiping, et al.
Veröffentlicht: (2026)
A Compact One-Way Fault-Tolerant Optical Quantum Computation
von: Du, Peilin, et al.
Veröffentlicht: (2025)
von: Du, Peilin, et al.
Veröffentlicht: (2025)
Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis
von: Li, Daoyang, et al.
Veröffentlicht: (2024)
von: Li, Daoyang, et al.
Veröffentlicht: (2024)
Counterfactual Graph for Multi-Agent LLM Calibration
von: Huang, Jiatan, et al.
Veröffentlicht: (2026)
von: Huang, Jiatan, et al.
Veröffentlicht: (2026)
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
von: Chang, Yue, et al.
Veröffentlicht: (2024)
von: Chang, Yue, et al.
Veröffentlicht: (2024)
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
The role of special and differential treatment for developing countries in GATT and the World Trade Organization / Constantine Michalopoulos
von: Michalopoulos, Constantine
Veröffentlicht: (2000)
von: Michalopoulos, Constantine
Veröffentlicht: (2000)
The integration of transition economies into the world trading system / Constantine Michalopoulos
von: Michalopoulos, Constantine
Veröffentlicht: (1999)
von: Michalopoulos, Constantine
Veröffentlicht: (1999)
World bank programs for adjustment and growth / Constantine Michalopoulos
von: Michalopoulos, Constantine
Veröffentlicht: (1987)
von: Michalopoulos, Constantine
Veröffentlicht: (1987)
Trade policy and market access issues for veveloping countries : implications for the millennium round / Constantine Michalopoulos
von: Michalopoulos, Constantine
Veröffentlicht: (1999)
von: Michalopoulos, Constantine
Veröffentlicht: (1999)
Trade performance and policy in the new independent states / Constantine Michalopoulos, David G. Tarr
von: Michalopoulos, Constantine
Veröffentlicht: (1996)
von: Michalopoulos, Constantine
Veröffentlicht: (1996)
Payments and finance problems in the commonwealth of independent states / Constantine Michalopoulos
von: Michalopoulos, Constantine
Veröffentlicht: (1996)
von: Michalopoulos, Constantine
Veröffentlicht: (1996)
Services trade in the Balkans / Constantine Michalopoulos, Vasileios Panousopoulos
von: Michalopoulos, Constantine
von: Michalopoulos, Constantine
The economics of customs union in the commonwealth of independent states / Constantine Michalopoulos, David Tarr
von: Michalopoulos, Constantine
Veröffentlicht: (1997)
von: Michalopoulos, Constantine
Veröffentlicht: (1997)
Ähnliche Einträge
-
IDEA: Enhancing the Rule Learning Ability of Large Language Model Agent through Induction, Deduction, and Abduction
von: He, Kaiyu, et al.
Veröffentlicht: (2024) -
Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers
von: He, Kaiyu, et al.
Veröffentlicht: (2026) -
GEAR: A General Evaluation Framework for Abductive Reasoning
von: He, Kaiyu, et al.
Veröffentlicht: (2025) -
HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
von: Wu, Peilin, et al.
Veröffentlicht: (2025) -
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty
von: Wu, Peilin, et al.
Veröffentlicht: (2025)