The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Si, Chenglei, Hashimoto, Tatsunori, Yang, Diyi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909660192178176
author Si, Chenglei
Hashimoto, Tatsunori
Yang, Diyi
author_facet Si, Chenglei
Hashimoto, Tatsunori
Yang, Diyi
contents Large Language Models (LLMs) have shown promise in accelerating the scientific research pipeline. A key capability for this process is the ability to generate novel research ideas, and prior studies have found settings in which LLM-generated research ideas were judged as more novel than human-expert ideas. However, a good idea should not simply appear to be novel, it should also result in better research after being executed. To test whether AI-generated ideas lead to better research outcomes, we conduct an execution study by recruiting 43 expert researchers to execute randomly-assigned ideas, either written by experts or generated by an LLM. Each expert spent over 100 hours implementing the idea and wrote a 4-page short paper to document the experiments. All the executed projects are then reviewed blindly by expert NLP researchers. Comparing the review scores of the same ideas before and after execution, the scores of the LLM-generated ideas decrease significantly more than expert-written ideas on all evaluation metrics (novelty, excitement, effectiveness, and overall; p < 0.05), closing the gap between LLM and human ideas observed at the ideation stage. When comparing the aggregated review scores from the execution study, we even observe that for many metrics there is a flip in rankings where human ideas score higher than LLM ideas. This ideation-execution gap highlights the limitations of current LLMs in generating truly effective research ideas and the challenge of evaluating research ideas in the absence of execution outcomes.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20803
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
Si, Chenglei
Hashimoto, Tatsunori
Yang, Diyi
Computation and Language
Artificial Intelligence
Computers and Society
Human-Computer Interaction
Machine Learning
Large Language Models (LLMs) have shown promise in accelerating the scientific research pipeline. A key capability for this process is the ability to generate novel research ideas, and prior studies have found settings in which LLM-generated research ideas were judged as more novel than human-expert ideas. However, a good idea should not simply appear to be novel, it should also result in better research after being executed. To test whether AI-generated ideas lead to better research outcomes, we conduct an execution study by recruiting 43 expert researchers to execute randomly-assigned ideas, either written by experts or generated by an LLM. Each expert spent over 100 hours implementing the idea and wrote a 4-page short paper to document the experiments. All the executed projects are then reviewed blindly by expert NLP researchers. Comparing the review scores of the same ideas before and after execution, the scores of the LLM-generated ideas decrease significantly more than expert-written ideas on all evaluation metrics (novelty, excitement, effectiveness, and overall; p < 0.05), closing the gap between LLM and human ideas observed at the ideation stage. When comparing the aggregated review scores from the execution study, we even observe that for many metrics there is a flip in rankings where human ideas score higher than LLM ideas. This ideation-execution gap highlights the limitations of current LLMs in generating truly effective research ideas and the challenge of evaluating research ideas in the absence of execution outcomes.
title The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
topic Computation and Language
Artificial Intelligence
Computers and Society
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2506.20803