ResearchGym: Evaluating Language Model Agents on Real-World AI Research
Fuente:
arXiv
Guardado en:
| Autores principales: | Garikaparthi, Aniketh, Patwardhan, Manasi, Cohan, Arman |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery
por: Garikaparthi, Aniketh, et al.
Publicado: (2025)
por: Garikaparthi, Aniketh, et al.
Publicado: (2025)
REVERE: Reflective Evolving Research Engineer for Scientific Workflows
por: Gangireddi, Balaji Dinesh, et al.
Publicado: (2026)
por: Gangireddi, Balaji Dinesh, et al.
Publicado: (2026)
Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation
por: Mule, Srujan P, et al.
Publicado: (2026)
por: Mule, Srujan P, et al.
Publicado: (2026)
MIR: Methodology Inspiration Retrieval for Scientific Research Problems
por: Garikaparthi, Aniketh, et al.
Publicado: (2025)
por: Garikaparthi, Aniketh, et al.
Publicado: (2025)
Can LLMs Perceive Time? An Empirical Investigation
por: Garikaparthi, Aniketh
Publicado: (2026)
por: Garikaparthi, Aniketh
Publicado: (2026)
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
por: Zhao, Yilun, et al.
Publicado: (2025)
por: Zhao, Yilun, et al.
Publicado: (2025)
Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
por: Wu, Sihong, et al.
Publicado: (2026)
por: Wu, Sihong, et al.
Publicado: (2026)
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
por: Xu, Zhijian, et al.
Publicado: (2025)
por: Xu, Zhijian, et al.
Publicado: (2025)
SciMDR: Advancing Scientific Multimodal Document Reasoning
por: Chen, Ziyu, et al.
Publicado: (2026)
por: Chen, Ziyu, et al.
Publicado: (2026)
RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation
por: Wu, Sihong, et al.
Publicado: (2026)
por: Wu, Sihong, et al.
Publicado: (2026)
ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies
por: Gandhi, Shubham, et al.
Publicado: (2025)
por: Gandhi, Shubham, et al.
Publicado: (2025)
Defend: Automated Rebuttals for Peer Review with Minimal Author Guidance
por: Khatri, Jyotsana, et al.
Publicado: (2026)
por: Khatri, Jyotsana, et al.
Publicado: (2026)
Acceleron: A Tool to Accelerate Research Ideation
por: Nigam, Harshit, et al.
Publicado: (2024)
por: Nigam, Harshit, et al.
Publicado: (2024)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
por: Wang, Zhun, et al.
Publicado: (2025)
por: Wang, Zhun, et al.
Publicado: (2025)
WorldGym: World Model as An Environment for Policy Evaluation
por: Quevedo, Julian, et al.
Publicado: (2025)
por: Quevedo, Julian, et al.
Publicado: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
ANCHOR: Branch-Point Data Generation for GUI Agents
por: Wei, Jinbiao, et al.
Publicado: (2026)
por: Wei, Jinbiao, et al.
Publicado: (2026)
The BrowserGym Ecosystem for Web Agent Research
por: De Chezelles, Thibault Le Sellier, et al.
Publicado: (2024)
por: De Chezelles, Thibault Le Sellier, et al.
Publicado: (2024)
One-shot Optimized Steering Vectors Mediate Safety-relevant Behaviors in LLMs
por: Dunefsky, Jacob, et al.
Publicado: (2025)
por: Dunefsky, Jacob, et al.
Publicado: (2025)
TutorGym: A Testbed for Evaluating AI Agents as Tutors and Students
por: Weitekamp, Daniel, et al.
Publicado: (2025)
por: Weitekamp, Daniel, et al.
Publicado: (2025)
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
por: Wei, Jinbiao, et al.
Publicado: (2026)
por: Wei, Jinbiao, et al.
Publicado: (2026)
Real-World Gaps in AI Governance Research
por: Strauss, Ilan, et al.
Publicado: (2025)
por: Strauss, Ilan, et al.
Publicado: (2025)
Step-level Optimization for Efficient Computer-use Agents
por: Wei, Jinbiao, et al.
Publicado: (2026)
por: Wei, Jinbiao, et al.
Publicado: (2026)
Bayesian Calibration of Win Rate Estimation with LLM Evaluators
por: Gao, Yicheng, et al.
Publicado: (2024)
por: Gao, Yicheng, et al.
Publicado: (2024)
ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research
por: Shen, Hao, et al.
Publicado: (2026)
por: Shen, Hao, et al.
Publicado: (2026)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
por: Wang, Zhun, et al.
Publicado: (2026)
por: Wang, Zhun, et al.
Publicado: (2026)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
por: Deng, Chunyuan, et al.
Publicado: (2023)
por: Deng, Chunyuan, et al.
Publicado: (2023)
Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
por: Jain, Riddhi, et al.
Publicado: (2025)
por: Jain, Riddhi, et al.
Publicado: (2025)
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
por: Tang, Xiangru, et al.
Publicado: (2023)
por: Tang, Xiangru, et al.
Publicado: (2023)
Deep Research Bench: Evaluating AI Web Research Agents
por: FutureSearch, et al.
Publicado: (2025)
por: FutureSearch, et al.
Publicado: (2025)
On the Benefits of Fine-Grained Loss Truncation: A Case Study on Factuality in Summarization
por: Flores, Lorenzo Jaime Yu, et al.
Publicado: (2024)
por: Flores, Lorenzo Jaime Yu, et al.
Publicado: (2024)
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
por: Zhang, Xichen, et al.
Publicado: (2026)
por: Zhang, Xichen, et al.
Publicado: (2026)
PersonaGym: Evaluating Persona Agents and LLMs
por: Samuel, Vinay, et al.
Publicado: (2024)
por: Samuel, Vinay, et al.
Publicado: (2024)
AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
por: Xi, Zhiheng, et al.
Publicado: (2024)
por: Xi, Zhiheng, et al.
Publicado: (2024)
Calibrating Long-form Generations from Large Language Models
por: Huang, Yukun, et al.
Publicado: (2024)
por: Huang, Yukun, et al.
Publicado: (2024)
FormGym: Doing Paperwork with Agents
por: Toles, Matthew, et al.
Publicado: (2025)
por: Toles, Matthew, et al.
Publicado: (2025)
PaperBench: Evaluating AI's Ability to Replicate AI Research
por: Starace, Giulio, et al.
Publicado: (2025)
por: Starace, Giulio, et al.
Publicado: (2025)
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
por: Lee, Jinu, et al.
Publicado: (2025)
por: Lee, Jinu, et al.
Publicado: (2025)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
por: Vijayvargiya, Sanidhya, et al.
Publicado: (2025)
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
por: Patwardhan, Tejal, et al.
Publicado: (2025)
por: Patwardhan, Tejal, et al.
Publicado: (2025)
Ejemplares similares
-
IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery
por: Garikaparthi, Aniketh, et al.
Publicado: (2025) -
REVERE: Reflective Evolving Research Engineer for Scientific Workflows
por: Gangireddi, Balaji Dinesh, et al.
Publicado: (2026) -
Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation
por: Mule, Srujan P, et al.
Publicado: (2026) -
MIR: Methodology Inspiration Retrieval for Scientific Research Problems
por: Garikaparthi, Aniketh, et al.
Publicado: (2025) -
Can LLMs Perceive Time? An Empirical Investigation
por: Garikaparthi, Aniketh
Publicado: (2026)