InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yunze, Fu, Dayuan, Si, Weiye, Huang, Zhen, Jiang, Mohan, Li, Keyu, Xia, Shijie, Sun, Jie, Xu, Tianze, Hu, Xiangkun, Lu, Pengrui, Cai, Xiaojie, Ye, Lyumanshan, Zhu, Wenhong, Xiao, Yang, Liu, Pengfei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interaction as Intelligence Part II: Asynchronous Human-Agent Rollout for Long-Horizon Task Training
by: Fu, Dayuan, et al.
Published: (2025)
by: Fu, Dayuan, et al.
Published: (2025)
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
by: Li, Keyu, et al.
Published: (2026)
by: Li, Keyu, et al.
Published: (2026)
ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry
by: Xu, Tianze, et al.
Published: (2025)
by: Xu, Tianze, et al.
Published: (2025)
DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
by: Zheng, Yuxiang, et al.
Published: (2025)
by: Zheng, Yuxiang, et al.
Published: (2025)
DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery
by: Li, Keyu, et al.
Published: (2025)
by: Li, Keyu, et al.
Published: (2025)
Context Engineering 2.0: The Context of Context Engineering
by: Hua, Qishuo, et al.
Published: (2025)
by: Hua, Qishuo, et al.
Published: (2025)
daVinci-Agency: Unlocking Long-Horizon Agency Data-Efficiently
by: Jiang, Mohan, et al.
Published: (2026)
by: Jiang, Mohan, et al.
Published: (2026)
Interaction as Intelligence: Deep Research With Human-AI Partnership
by: Ye, Lyumanshan, et al.
Published: (2025)
by: Ye, Lyumanshan, et al.
Published: (2025)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
by: Lu, Pengrui, et al.
Published: (2026)
by: Lu, Pengrui, et al.
Published: (2026)
daVinci-Dev: Agent-native Mid-training for Software Engineering
by: Zeng, Ji, et al.
Published: (2026)
by: Zeng, Ji, et al.
Published: (2026)
AlphaGo Moment for Model Architecture Discovery
by: Liu, Yixiu, et al.
Published: (2025)
by: Liu, Yixiu, et al.
Published: (2025)
PreAct: Prediction Enhances Agent's Planning Ability
by: Fu, Dayuan, et al.
Published: (2024)
by: Fu, Dayuan, et al.
Published: (2024)
daVinci-LLM:Towards the Science of Pretraining
by: Qin, Yiwei, et al.
Published: (2026)
by: Qin, Yiwei, et al.
Published: (2026)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World
by: He, Yanheng, et al.
Published: (2024)
by: He, Yanheng, et al.
Published: (2024)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
by: Tu, Xinming, et al.
Published: (2026)
by: Tu, Xinming, et al.
Published: (2026)
Gamifying Green: Sustainable Innovation Through Digital Platform Ecosystems
by: Pengfei Fu, et al.
Published: (2025)
by: Pengfei Fu, et al.
Published: (2025)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
by: Xu, Tianze, et al.
Published: (2026)
by: Xu, Tianze, et al.
Published: (2026)
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
by: Huang, Jiahao, et al.
Published: (2026)
by: Huang, Jiahao, et al.
Published: (2026)
Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System
by: Du, Haikuo, et al.
Published: (2025)
by: Du, Haikuo, et al.
Published: (2025)
OlympicArena Medal Ranks: Who Is the Most Intelligent AI So Far?
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
Innovative study for enhanced performance of the photogalvanic cells for solar energy conversion and storage
by: Mohan Lal
Published: (2025)
by: Mohan Lal
Published: (2025)
Innovative Antibacterial Magneto‐Conductive Multifunctional Hydrogel From Biocompatible Materials
by: Muzammal Hussain, et al.
Published: (2025)
by: Muzammal Hussain, et al.
Published: (2025)
Retail Investors' Environmental Attention and Green Innovation—Evidence From Interactive Platforms
by: En Xie, et al.
Published: (2025)
by: En Xie, et al.
Published: (2025)
TRIZ Agents: A Multi-Agent LLM Approach for TRIZ-Based Innovation
by: Szczepanik, Kamil, et al.
Published: (2025)
by: Szczepanik, Kamil, et al.
Published: (2025)
daVinci-Env: Open SWE Environment Synthesis at Scale
by: Fu, Dayuan, et al.
Published: (2026)
by: Fu, Dayuan, et al.
Published: (2026)
Deep FinResearch Bench: Evaluating AI's Ability to Conduct Professional Financial Investment Research
by: Haque, Mirazul, et al.
Published: (2026)
by: Haque, Mirazul, et al.
Published: (2026)
Enhancing Electrode Performance through Triple Distribution Modulation of Active Material, Conductive Agent, and Porosity
by: Renjie He, et al.
Published: (2024)
by: Renjie He, et al.
Published: (2024)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
by: Zheng, Junhao, et al.
Published: (2025)
by: Zheng, Junhao, et al.
Published: (2025)
MSI-Agent: Incorporating Multi-Scale Insight into Embodied Agents for Superior Planning and Decision-Making
by: Fu, Dayuan, et al.
Published: (2024)
by: Fu, Dayuan, et al.
Published: (2024)
GAI: Generative Agents for Innovation
by: Sato, Masahiro
Published: (2024)
by: Sato, Masahiro
Published: (2024)
SR-Scientist: Scientific Equation Discovery With Agentic AI
by: Xia, Shijie, et al.
Published: (2025)
by: Xia, Shijie, et al.
Published: (2025)
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
by: Xu, Zelai, et al.
Published: (2025)
by: Xu, Zelai, et al.
Published: (2025)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
by: Prathifkumar, Thanosan, et al.
Published: (2025)
by: Prathifkumar, Thanosan, et al.
Published: (2025)
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
by: Liu, Xuan, et al.
Published: (2026)
by: Liu, Xuan, et al.
Published: (2026)
Mapping Innovation Networks: A Network-Based Approach to Actor Heterogeneity in National Innovation Systems
by: Jeong, Dawoon, et al.
Published: (2025)
by: Jeong, Dawoon, et al.
Published: (2025)
Data Darwinism Part I: Unlocking the Value of Scientific Data for Pre-training
by: Qin, Yiwei, et al.
Published: (2026)
by: Qin, Yiwei, et al.
Published: (2026)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
by: Hou, Yixuan, et al.
Published: (2025)
by: Hou, Yixuan, et al.
Published: (2025)
"Ghost of the past": identifying and resolving privacy leakage from LLM's memory through proactive user interaction
by: Zhang, Shuning, et al.
Published: (2024)
by: Zhang, Shuning, et al.
Published: (2024)
Similar Items
-
Interaction as Intelligence Part II: Asynchronous Human-Agent Rollout for Long-Horizon Task Training
by: Fu, Dayuan, et al.
Published: (2025) -
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts
by: Li, Keyu, et al.
Published: (2026) -
ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry
by: Xu, Tianze, et al.
Published: (2025) -
DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
by: Zheng, Yuxiang, et al.
Published: (2025) -
DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery
by: Li, Keyu, et al.
Published: (2025)