AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bragg, Jonathan, D'Arcy, Mike, Balepur, Nishant, Bareket, Dan, Dalvi, Bhavana, Feldman, Sergey, Haddad, Dany, Hwang, Jena D., Jansen, Peter, Kishore, Varsha, Majumder, Bodhisattwa Prasad, Naik, Aakanksha, Rahamimov, Sigal, Richardson, Kyle, Singh, Amanpreet, Surana, Harshit, Tiktinsky, Aryeh, Vasu, Rosni, Wiener, Guy, Anastasiades, Chloe, Candra, Stefan, Dunkelberger, Jason, Emery, Dan, Evans, Rob, Hamada, Malachi, Huff, Regan, Kinney, Rodney, Latzke, Matt, Lochner, Jaron, Lozano-Aguilera, Ruben, Nguyen, Cecile, Rao, Smita, Tanaka, Amber, Vlahos, Brooke, Clark, Peter, Downey, Doug, Goldberg, Yoav, Sabharwal, Ashish, Weld, Daniel S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding Usage and Engagement in AI-Powered Scientific Research Tools: The Asta Interaction Dataset
von: Haddad, Dany, et al.
Veröffentlicht: (2026)
von: Haddad, Dany, et al.
Veröffentlicht: (2026)
Ai2 Scholar QA: Organized Literature Synthesis with Attribution
von: Singh, Amanpreet, et al.
Veröffentlicht: (2025)
von: Singh, Amanpreet, et al.
Veröffentlicht: (2025)
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
von: Hwang, Jena D., et al.
Veröffentlicht: (2026)
von: Hwang, Jena D., et al.
Veröffentlicht: (2026)
Data-driven Discovery with Large Generative Models
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024)
DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
Metric-valued regression
von: Cohen, Dan Tsir, et al.
Veröffentlicht: (2022)
von: Cohen, Dan Tsir, et al.
Veröffentlicht: (2022)
Latent Factor Models Meets Instructions: Goal-conditioned Latent Factor Discovery without Task Supervision
von: Xie, Zhouhang, et al.
Veröffentlicht: (2025)
von: Xie, Zhouhang, et al.
Veröffentlicht: (2025)
AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
LitPivot: Developing Well-Situated Research Ideas Through Dynamic Contextualization and Critique within the Literature Landscape
von: Kambhamettu, Hita, et al.
Veröffentlicht: (2026)
von: Kambhamettu, Hita, et al.
Veröffentlicht: (2026)
HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance
von: Vasu, Rosni, et al.
Veröffentlicht: (2025)
von: Vasu, Rosni, et al.
Veröffentlicht: (2025)
Cocoa: Co-Planning and Co-Execution with AI Agents
von: Feng, K. J. Kevin, et al.
Veröffentlicht: (2024)
von: Feng, K. J. Kevin, et al.
Veröffentlicht: (2024)
CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
von: Jansen, Peter, et al.
Veröffentlicht: (2025)
von: Jansen, Peter, et al.
Veröffentlicht: (2025)
Do Pretrained Contextual Language Models Distinguish between Hebrew Homograph Analyses?
von: Shmidman, Avi, et al.
Veröffentlicht: (2024)
von: Shmidman, Avi, et al.
Veröffentlicht: (2024)
IdeaSynth: Iterative Research Idea Development Through Evolving and Composing Idea Facets with Literature-Grounded Feedback
von: Pu, Kevin, et al.
Veröffentlicht: (2024)
von: Pu, Kevin, et al.
Veröffentlicht: (2024)
E-learning open seminar on "Human–centered artificial intelligence in education: From theory to practice"
von: Anastasiades, Panagiotes, et al.
Veröffentlicht: (2025)
von: Anastasiades, Panagiotes, et al.
Veröffentlicht: (2025)
Realizable Bayes-Consistency for General Metric Losses
von: Cohen, Dan Tsir, et al.
Veröffentlicht: (2026)
von: Cohen, Dan Tsir, et al.
Veröffentlicht: (2026)
HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation
von: Vasu, Rosni, et al.
Veröffentlicht: (2025)
von: Vasu, Rosni, et al.
Veröffentlicht: (2025)
Dataset celebrity endorser and purchase intention
von: Candra, Sevenpri
Veröffentlicht: (2025)
von: Candra, Sevenpri
Veröffentlicht: (2025)
Strong MHD Turbulence and Coherent Structures as Drivers of Cosmic Particle Acceleration
von: Vlahos, Loukas
Veröffentlicht: (2026)
von: Vlahos, Loukas
Veröffentlicht: (2026)
The Teachers' Mental Health Literacy Scale
von: Candra Skrzypek
Veröffentlicht: (2024)
von: Candra Skrzypek
Veröffentlicht: (2024)
American wilds
von: Vlahos, James
von: Vlahos, James
Next weekend
von: Vlahos, James
von: Vlahos, James
Digital Grassroots Organizing: How Residents Are Shaping Local Civic Participation
von: Nick Vlahos
Veröffentlicht: (2025)
von: Nick Vlahos
Veröffentlicht: (2025)
Omakase: proactive assistance with actionable suggestions for evolving scientific research projects
von: Siangliulue, Pao, et al.
Veröffentlicht: (2026)
von: Siangliulue, Pao, et al.
Veröffentlicht: (2026)
Papers-to-Posts: Supporting Detailed Long-Document Summarization with an Interactive LLM-Powered Source Outline
von: Radensky, Marissa, et al.
Veröffentlicht: (2024)
von: Radensky, Marissa, et al.
Veröffentlicht: (2024)
Spectral upward radiance and retrieved cloud properties
von: Krisna, Trismono Candra
Veröffentlicht: (2018)
von: Krisna, Trismono Candra
Veröffentlicht: (2018)
PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected Papers
von: Lee, Yoonjoo, et al.
Veröffentlicht: (2024)
von: Lee, Yoonjoo, et al.
Veröffentlicht: (2024)
El Público y la biblioteca : metodologías para la difusión de la lectura / edición Grazia Asta y Paolo Federighi ; colaboración Grazia Asta ... [et all.]
On the tensorization of the variational distance
von: Kontorovich, Aryeh
Veröffentlicht: (2024)
von: Kontorovich, Aryeh
Veröffentlicht: (2024)
Por qué somos liberales
von: Neier, Aryeh
Veröffentlicht: (2006)
von: Neier, Aryeh
Veröffentlicht: (2006)
A homogenization principle for total variation
von: Kontorovich, Aryeh
Veröffentlicht: (2026)
von: Kontorovich, Aryeh
Veröffentlicht: (2026)
Representation Learning on a Random Lattice
von: Brill, Aryeh
Veröffentlicht: (2025)
von: Brill, Aryeh
Veröffentlicht: (2025)
Decoupling Maximal Inequalities
von: Kontorovich, Aryeh
Veröffentlicht: (2023)
von: Kontorovich, Aryeh
Veröffentlicht: (2023)
TV homogenization inequalities
von: Kontorovich, Aryeh
Veröffentlicht: (2026)
von: Kontorovich, Aryeh
Veröffentlicht: (2026)
Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills
von: Nottingham, Kolby, et al.
Veröffentlicht: (2024)
von: Nottingham, Kolby, et al.
Veröffentlicht: (2024)
Multi-physics Preconditioning for Thermally Activated Batteries
von: Phillips, Malachi
Veröffentlicht: (2026)
von: Phillips, Malachi
Veröffentlicht: (2026)
Reseña de "Los ecos de Mathias Goeritz. Catálogo de la exposición" Ferruccio Asta (coord.) y "Los ecos de Mathias Goeritz" Rodrigez Prampolini y Ferruccio Asta.
von: Manuel Rocha Iturbide
Veröffentlicht: (1998)
von: Manuel Rocha Iturbide
Veröffentlicht: (1998)
Transcendental Optimization in Geometric Programming via Power Series Approximations
von: singh, Amanpreet
Veröffentlicht: (2025)
von: singh, Amanpreet
Veröffentlicht: (2025)
Ähnliche Einträge
-
Understanding Usage and Engagement in AI-Powered Scientific Research Tools: The Asta Interaction Dataset
von: Haddad, Dany, et al.
Veröffentlicht: (2026) -
Ai2 Scholar QA: Organized Literature Synthesis with Attribution
von: Singh, Amanpreet, et al.
Veröffentlicht: (2025) -
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
von: Hwang, Jena D., et al.
Veröffentlicht: (2026) -
Data-driven Discovery with Large Generative Models
von: Majumder, Bodhisattwa Prasad, et al.
Veröffentlicht: (2024) -
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)