The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
Fuente:
arXiv
Saved in:
| Main Authors: | Si, Chenglei, Hashimoto, Tatsunori, Yang, Diyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
Towards Execution-Grounded Automated AI Research
by: Si, Chenglei, et al.
Published: (2026)
by: Si, Chenglei, et al.
Published: (2026)
DiscoverLLM: From Executing Intents to Discovering Them
by: Kim, Tae Soo, et al.
Published: (2026)
by: Kim, Tae Soo, et al.
Published: (2026)
Can Large Language Models Unlock Novel Scientific Research Ideas?
by: Kumar, Sandeep, et al.
Published: (2024)
by: Kumar, Sandeep, et al.
Published: (2024)
Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
by: Shao, Yijia, et al.
Published: (2025)
by: Shao, Yijia, et al.
Published: (2025)
Properties and Challenges of LLM-Generated Explanations
by: Kunz, Jenny, et al.
Published: (2024)
by: Kunz, Jenny, et al.
Published: (2024)
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
by: Abels, Axel, et al.
Published: (2025)
by: Abels, Axel, et al.
Published: (2025)
ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios
by: Phutane, Mahika, et al.
Published: (2025)
by: Phutane, Mahika, et al.
Published: (2025)
Augmenting Research Ideation with Data: An Empirical Investigation in Social Science
by: Liu, Xiao, et al.
Published: (2025)
by: Liu, Xiao, et al.
Published: (2025)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
Synthetic Reader Panels: Tournament-Based Ideation with LLM Personas for Autonomous Publishing
by: Zimmerman, Fred
Published: (2026)
by: Zimmerman, Fred
Published: (2026)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
by: Sturgeon, Benjamin, et al.
Published: (2025)
by: Sturgeon, Benjamin, et al.
Published: (2025)
EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety
by: Qiu, Jiahao, et al.
Published: (2025)
by: Qiu, Jiahao, et al.
Published: (2025)
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
by: Singh, Anikait, et al.
Published: (2025)
by: Singh, Anikait, et al.
Published: (2025)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
by: Chiu, Yu Ying, et al.
Published: (2025)
by: Chiu, Yu Ying, et al.
Published: (2025)
How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment
by: Ashkinaze, Joshua, et al.
Published: (2024)
by: Ashkinaze, Joshua, et al.
Published: (2024)
Llms, Virtual Users, and Bias: Predicting Any Survey Question Without Human Data
by: Sinacola, Enzo, et al.
Published: (2025)
by: Sinacola, Enzo, et al.
Published: (2025)
LLM4PM: A case study on using Large Language Models for Process Modeling in Enterprise Organizations
by: Ziche, Clara, et al.
Published: (2024)
by: Ziche, Clara, et al.
Published: (2024)
Who is a Better Matchmaker? Human vs. Algorithmic Judge Assignment in a High-Stakes Startup Competition
by: Xi, Sarina, et al.
Published: (2025)
by: Xi, Sarina, et al.
Published: (2025)
EduAgent: Generative Student Agents in Learning
by: Xu, Songlin, et al.
Published: (2024)
by: Xu, Songlin, et al.
Published: (2024)
Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
ProgressGym: Alignment with a Millennium of Moral Progress
by: Qiu, Tianyi, et al.
Published: (2024)
by: Qiu, Tianyi, et al.
Published: (2024)
The Ontological Dissonance Hypothesis: AI-Triggered Delusional Ideation as Folie a Deux Technologique
by: Lipinska, Izabela, et al.
Published: (2025)
by: Lipinska, Izabela, et al.
Published: (2025)
Making Language Models Better Tool Learners with Execution Feedback
by: Qiao, Shuofei, et al.
Published: (2023)
by: Qiao, Shuofei, et al.
Published: (2023)
Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
by: Jain, Raunak
Published: (2025)
by: Jain, Raunak
Published: (2025)
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
by: Seshadri, Preethi, et al.
Published: (2026)
by: Seshadri, Preethi, et al.
Published: (2026)
MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance
by: Xu, Jia, et al.
Published: (2025)
by: Xu, Jia, et al.
Published: (2025)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
by: Baidya, Avinash, et al.
Published: (2025)
by: Baidya, Avinash, et al.
Published: (2025)
Examining and Addressing Barriers to Diversity in LLM-Generated Ideas
by: Deng, Yuting, et al.
Published: (2026)
by: Deng, Yuting, et al.
Published: (2026)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
by: Ngueajio, Mikel K., et al.
Published: (2025)
by: Ngueajio, Mikel K., et al.
Published: (2025)
AmarDoctor: An AI-Driven, Multilingual, Voice-Interactive Digital Health Application for Primary Care Triage and Patient Management to Bridge the Digital Health Divide for Bengali Speakers
by: Nahar, Nazmun, et al.
Published: (2025)
by: Nahar, Nazmun, et al.
Published: (2025)
From Measurement to Expertise: Empathetic Expert Adapters for Context-Based Empathy in Conversational AI Agents
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs
by: Sun, Huaman, et al.
Published: (2023)
by: Sun, Huaman, et al.
Published: (2023)
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
by: Chiu, Yu Ying, et al.
Published: (2025)
by: Chiu, Yu Ying, et al.
Published: (2025)
Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation
by: Samory, Mattia, et al.
Published: (2025)
by: Samory, Mattia, et al.
Published: (2025)
What are human values, and how do we align AI to them?
by: Klingefjord, Oliver, et al.
Published: (2024)
by: Klingefjord, Oliver, et al.
Published: (2024)
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
by: Bhattacharya, Haimanti, et al.
Published: (2024)
by: Bhattacharya, Haimanti, et al.
Published: (2024)
Large Language Models Can Infer Personality from Free-Form User Interactions
by: Peters, Heinrich, et al.
Published: (2024)
by: Peters, Heinrich, et al.
Published: (2024)
Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior
by: Ardebili, Ali Aghazadeh, et al.
Published: (2026)
by: Ardebili, Ali Aghazadeh, et al.
Published: (2026)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
by: Zheng, Mingqian, et al.
Published: (2023)
by: Zheng, Mingqian, et al.
Published: (2023)
Similar Items
-
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
by: Si, Chenglei, et al.
Published: (2024) -
Towards Execution-Grounded Automated AI Research
by: Si, Chenglei, et al.
Published: (2026) -
DiscoverLLM: From Executing Intents to Discovering Them
by: Kim, Tae Soo, et al.
Published: (2026) -
Can Large Language Models Unlock Novel Scientific Research Ideas?
by: Kumar, Sandeep, et al.
Published: (2024) -
Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
by: Shao, Yijia, et al.
Published: (2025)