AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Lupidi, Alisia, Gauri, Bhavul, Foster, Thomas Simon, Omari, Bassel Al, Magka, Despoina, Pepe, Alberto, Audran-Reiss, Alexis, Aghamelu, Muna, Baldwin, Nicolas, Cipolina-Kun, Lucia, Gagnon-Audet, Jean-Christophe, Leow, Chee Hau, Lefdal, Sandra, Mossalam, Hossam, Moudgil, Abhinav, Nazir, Saba, Tewolde, Emanuel, Urrego, Isabel, Estape, Jordi Armengol, Budhiraja, Amar, Chaurasia, Gaurav, Charnalia, Abhishek, Dunfield, Derek, Hambardzumyan, Karen, Izcovich, Daniel, Josifoski, Martin, Mediratta, Ishita, Niu, Kelvin, Pathak, Parth, Shvartsman, Michael, Toledo, Edan, Protopopov, Anton, Raileanu, Roberta, Miller, Alexander, Shavrina, Tatiana, Foerster, Jakob, Bachrach, Yoram |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
by: Audran-Reiss, Alexis, et al.
Published: (2025)
by: Audran-Reiss, Alexis, et al.
Published: (2025)
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
by: Maiti, Shalini, et al.
Published: (2025)
by: Maiti, Shalini, et al.
Published: (2025)
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
by: Zhao, Bingchen, et al.
Published: (2025)
by: Zhao, Bingchen, et al.
Published: (2025)
AIRA_2: Overcoming Bottlenecks in AI Research Agents
by: Hambardzumyan, Karen, et al.
Published: (2026)
by: Hambardzumyan, Karen, et al.
Published: (2026)
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
by: Toledo, Edan, et al.
Published: (2025)
by: Toledo, Edan, et al.
Published: (2025)
APRES: An Agentic Paper Revision and Evaluation System
by: Zhao, Bingchen, et al.
Published: (2026)
by: Zhao, Bingchen, et al.
Published: (2026)
The Generalization Gap in Offline Reinforcement Learning
by: Mediratta, Ishita, et al.
Published: (2023)
by: Mediratta, Ishita, et al.
Published: (2023)
Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
by: Lupidi, Alisia, et al.
Published: (2024)
by: Lupidi, Alisia, et al.
Published: (2024)
A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs
by: Kalra, Dayal Singh, et al.
Published: (2026)
by: Kalra, Dayal Singh, et al.
Published: (2026)
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
by: Nathani, Deepak, et al.
Published: (2025)
by: Nathani, Deepak, et al.
Published: (2025)
Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
by: Alexey, Protopopov
Published: (2026)
by: Alexey, Protopopov
Published: (2026)
A Robust Remote Photoplethysmography Method
by: Protopopov, Alexey
Published: (2025)
by: Protopopov, Alexey
Published: (2025)
A Robust Camera-based Method for Breath Rate Measurement
by: Protopopov, Alexey
Published: (2025)
by: Protopopov, Alexey
Published: (2025)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
by: Kirk, Robert, et al.
Published: (2023)
by: Kirk, Robert, et al.
Published: (2023)
Flattening subtyping by eta expansion
by: Dunfield, Jana
Published: (2024)
by: Dunfield, Jana
Published: (2024)
Epistemic Dissonance and Modal Boundaries
by: Raileanu, Dragos
Published: (2025)
by: Raileanu, Dragos
Published: (2025)
Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design
by: Pepe, Alberto, et al.
Published: (2026)
by: Pepe, Alberto, et al.
Published: (2026)
Social Financing Alternatives for Social Economy Enterprises: Coop57
by: Glòria Estapé-Dubreuil
Published: (2014)
by: Glòria Estapé-Dubreuil
Published: (2014)
From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars
by: Kornilov, Albert, et al.
Published: (2024)
by: Kornilov, Albert, et al.
Published: (2024)
On Some Extensions of the Boué-Dupuis Variational Formula
by: Budhiraja, A.
Published: (2024)
by: Budhiraja, A.
Published: (2024)
How to choose a best kidney for pediatric kidney transplant recipients
by: Asha Moudgil, et al.
Published: (2024)
by: Asha Moudgil, et al.
Published: (2024)
The Finiteness Principle for the boundary values of $C^2$-functions
by: Shvartsman, Pavel
Published: (2024)
by: Shvartsman, Pavel
Published: (2024)
Efficient Algorithms for Lipschitz Selections of Set-Valued Mappings in ${\bf R}^2$: long version
by: Shvartsman, Pavel
Published: (2025)
by: Shvartsman, Pavel
Published: (2025)
DPs, Phi-features and Tense in the Context of Abyssinian (Eritrean and Ethiopian) Semitic Languages
by: Tewolde, Tesfay
Published: (2022)
by: Tewolde, Tesfay
Published: (2022)
AIRS-assisted Vehicular Networks with Rate-Splitting SWIPT Receivers: Joint Trajectory and Communication Design
by: Nam, Gyoungyoon, et al.
Published: (2024)
by: Nam, Gyoungyoon, et al.
Published: (2024)
julieaudet/cell-manufacturing: HDDE
by: Julie Audet
Published: (2026)
by: Julie Audet
Published: (2026)
GIS in schools / Richard Audet and Gail Ludwig
by: Audet, Richard
by: Audet, Richard
Scaling and Distilling Transformer Models for sEMG
by: Mehlman, Nicholas, et al.
Published: (2025)
by: Mehlman, Nicholas, et al.
Published: (2025)
Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
by: Nathanson, Samuel, et al.
Published: (2025)
by: Nathanson, Samuel, et al.
Published: (2025)
Can LLMs Extract Frame-Semantic Arguments?
by: Devasier, Jacob, et al.
Published: (2025)
by: Devasier, Jacob, et al.
Published: (2025)
Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play
by: Cipolina-Kun, Lucia, et al.
Published: (2025)
by: Cipolina-Kun, Lucia, et al.
Published: (2025)
Irreducibility and locus of complex roots of polynomials related to Fermat's Last Theorem
by: Karapetyan, Hayk, et al.
Published: (2025)
by: Karapetyan, Hayk, et al.
Published: (2025)
On ideal class groups of totally degenerate number rings
by: Hambardzumyan, Ruben, et al.
Published: (2025)
by: Hambardzumyan, Ruben, et al.
Published: (2025)
Graphs, Disjoint Matchings and Some Inequalities
by: Hambardzumyan, Lianna, et al.
Published: (2015)
by: Hambardzumyan, Lianna, et al.
Published: (2015)
Real Time Fatigue Crack Growth Monitoring Using High Precision Control and Data Acquisition Systems
by: Hambardzumyan, Arev, et al.
Published: (2025)
by: Hambardzumyan, Arev, et al.
Published: (2025)
Supposedly Equivalent Facts That Aren't? Entity Frequency in Pre-training Induces Asymmetry in LLMs
by: He, Yuan, et al.
Published: (2025)
by: He, Yuan, et al.
Published: (2025)
Constructing surfaces with first Steklov eigenvalue of arbitrarily large multiplicity
by: Audet-Beaumont, Samuel
Published: (2024)
by: Audet-Beaumont, Samuel
Published: (2024)
A unified Casson-Lin invariant for the real forms of SL(2)
by: Dunfield, Nathan M., et al.
Published: (2022)
by: Dunfield, Nathan M., et al.
Published: (2022)
Ribbon concordances and slice obstructions: experiments and examples
by: Dunfield, Nathan M., et al.
Published: (2025)
by: Dunfield, Nathan M., et al.
Published: (2025)
Similar Items
-
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
by: Audran-Reiss, Alexis, et al.
Published: (2025) -
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
by: Maiti, Shalini, et al.
Published: (2025) -
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
by: Zhao, Bingchen, et al.
Published: (2025) -
AIRA_2: Overcoming Bottlenecks in AI Research Agents
by: Hambardzumyan, Karen, et al.
Published: (2026) -
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
by: Toledo, Edan, et al.
Published: (2025)