What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
Fuente:
arXiv
Saved in:
| Main Authors: | Audran-Reiss, Alexis, Armengol-Estapé, Jordi, Hambardzumyan, Karen, Budhiraja, Amar, Josifoski, Martin, Toledo, Edan, Hazra, Rishi, Magka, Despoina, Shvartsman, Michael, Pathak, Parth, Kao, Justine T, Cipolina-Kun, Lucia, Gauri, Bhavul, Gagnon-Audet, Jean-Christophe, Tewolde, Emanuel, Zhang, Jenny, Cohen, Taco, Adi, Yossi, Shavrina, Tatiana, Bachrach, Yoram |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
by: Maiti, Shalini, et al.
Published: (2025)
by: Maiti, Shalini, et al.
Published: (2025)
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
by: Lupidi, Alisia, et al.
Published: (2026)
by: Lupidi, Alisia, et al.
Published: (2026)
AIRA_2: Overcoming Bottlenecks in AI Research Agents
by: Hambardzumyan, Karen, et al.
Published: (2026)
by: Hambardzumyan, Karen, et al.
Published: (2026)
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
by: Toledo, Edan, et al.
Published: (2025)
by: Toledo, Edan, et al.
Published: (2025)
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
by: Zhao, Bingchen, et al.
Published: (2025)
by: Zhao, Bingchen, et al.
Published: (2025)
APRES: An Agentic Paper Revision and Evaluation System
by: Zhao, Bingchen, et al.
Published: (2026)
by: Zhao, Bingchen, et al.
Published: (2026)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
by: Pepe, Alberto, et al.
Published: (2026)
by: Pepe, Alberto, et al.
Published: (2026)
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
by: Nathani, Deepak, et al.
Published: (2025)
by: Nathani, Deepak, et al.
Published: (2025)
Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
Bootstrapping Task Spaces for Self-Improvement
by: Jiang, Minqi, et al.
Published: (2025)
by: Jiang, Minqi, et al.
Published: (2025)
Scaling and Distilling Transformer Models for sEMG
by: Mehlman, Nicholas, et al.
Published: (2025)
by: Mehlman, Nicholas, et al.
Published: (2025)
Neural Mean-Field Games: Extending Mean-Field Game Theory with Neural Stochastic Differential Equations
by: Thöni, Anna C. M., et al.
Published: (2025)
by: Thöni, Anna C. M., et al.
Published: (2025)
LLMs versus the Halting Problem: Characterizing Program Termination Reasoning
by: Sultan, Oren, et al.
Published: (2026)
by: Sultan, Oren, et al.
Published: (2026)
Diverse AI Personas Can Mitigate the Homogenization Effect in Human-AI Collaborative Ideation
by: Wan, Yun, et al.
Published: (2025)
by: Wan, Yun, et al.
Published: (2025)
A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs
by: Kalra, Dayal Singh, et al.
Published: (2026)
by: Kalra, Dayal Singh, et al.
Published: (2026)
Scaling Small Agents Through Strategy Auctions
by: Alazraki, Lisa, et al.
Published: (2026)
by: Alazraki, Lisa, et al.
Published: (2026)
Crowd IQ -- Aggregating Opinions to Boost Performance
by: Kosinski, Michal, et al.
Published: (2024)
by: Kosinski, Michal, et al.
Published: (2024)
SLaDe: A Portable Small Language Model Decompiler for Optimized Assembly
by: Armengol-Estapé, Jordi, et al.
Published: (2023)
by: Armengol-Estapé, Jordi, et al.
Published: (2023)
Social Financing Alternatives for Social Economy Enterprises: Coop57
by: Glòria Estapé-Dubreuil
Published: (2014)
by: Glòria Estapé-Dubreuil
Published: (2014)
From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars
by: Kornilov, Albert, et al.
Published: (2024)
by: Kornilov, Albert, et al.
Published: (2024)
On Some Extensions of the Boué-Dupuis Variational Formula
by: Budhiraja, A.
Published: (2024)
by: Budhiraja, A.
Published: (2024)
Las actividades productivas y su relación con la contaminación del agua de la Microcuenca Negroyacu, en Guaranda, Ecuador
by: Carlos Taco-Taco
Published: (2017)
by: Carlos Taco-Taco
Published: (2017)
The Finiteness Principle for the boundary values of $C^2$-functions
by: Shvartsman, Pavel
Published: (2024)
by: Shvartsman, Pavel
Published: (2024)
Efficient Algorithms for Lipschitz Selections of Set-Valued Mappings in ${\bf R}^2$: long version
by: Shvartsman, Pavel
Published: (2025)
by: Shvartsman, Pavel
Published: (2025)
Mind the Web: The Security of Web Use Agents
by: Shapira, Avishag, et al.
Published: (2025)
by: Shapira, Avishag, et al.
Published: (2025)
DPs, Phi-features and Tense in the Context of Abyssinian (Eritrean and Ethiopian) Semitic Languages
by: Tewolde, Tesfay
Published: (2022)
by: Tewolde, Tesfay
Published: (2022)
julieaudet/cell-manufacturing: HDDE
by: Julie Audet
Published: (2026)
by: Julie Audet
Published: (2026)
GIS in schools / Richard Audet and Gail Ludwig
by: Audet, Richard
by: Audet, Richard
Feature Likelihood Divergence: Evaluating the Generalization of Generative Models Using Samples
by: Jiralerspong, Marco, et al.
Published: (2023)
by: Jiralerspong, Marco, et al.
Published: (2023)
Modelling Chemical Reaction Networks using Neural Ordinary Differential Equations
by: Thöni, Anna C. M., et al.
Published: (2025)
by: Thöni, Anna C. M., et al.
Published: (2025)
Don't Transform the Code, Code the Transforms: Towards Precise Code Rewriting using LLMs
by: Cummins, Chris, et al.
Published: (2024)
by: Cummins, Chris, et al.
Published: (2024)
DUAS FACES DO PODER
by: Peter Bachrach
Published: (2011)
by: Peter Bachrach
Published: (2011)
Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play
by: Cipolina-Kun, Lucia, et al.
Published: (2025)
by: Cipolina-Kun, Lucia, et al.
Published: (2025)
Asking the Right Questions: Improving Reasoning with Generated Stepping Stones
by: Hu, Hengyuan, et al.
Published: (2026)
by: Hu, Hengyuan, et al.
Published: (2026)
Irreducibility and locus of complex roots of polynomials related to Fermat's Last Theorem
by: Karapetyan, Hayk, et al.
Published: (2025)
by: Karapetyan, Hayk, et al.
Published: (2025)
On ideal class groups of totally degenerate number rings
by: Hambardzumyan, Ruben, et al.
Published: (2025)
by: Hambardzumyan, Ruben, et al.
Published: (2025)
Graphs, Disjoint Matchings and Some Inequalities
by: Hambardzumyan, Lianna, et al.
Published: (2015)
by: Hambardzumyan, Lianna, et al.
Published: (2015)
Real Time Fatigue Crack Growth Monitoring Using High Precision Control and Data Acquisition Systems
by: Hambardzumyan, Arev, et al.
Published: (2025)
by: Hambardzumyan, Arev, et al.
Published: (2025)
Training AI Co-Scientists Using Rubric Rewards
by: Goel, Shashwat, et al.
Published: (2025)
by: Goel, Shashwat, et al.
Published: (2025)
Constructing surfaces with first Steklov eigenvalue of arbitrarily large multiplicity
by: Audet-Beaumont, Samuel
Published: (2024)
by: Audet-Beaumont, Samuel
Published: (2024)
Similar Items
-
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
by: Maiti, Shalini, et al.
Published: (2025) -
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
by: Lupidi, Alisia, et al.
Published: (2026) -
AIRA_2: Overcoming Bottlenecks in AI Research Agents
by: Hambardzumyan, Karen, et al.
Published: (2026) -
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
by: Toledo, Edan, et al.
Published: (2025) -
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
by: Zhao, Bingchen, et al.
Published: (2025)