Skip to content
Universidad del Mar SIBUMAR Descubridor Institucional UMAR
  • Inicio
  • Búsqueda avanzada
  • Explorar
  • Login
    • English
    • Deutsch
    • Español
    • Français
    • Italiano
Advanced
  • AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
Cover Image

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bragg, Jonathan, D'Arcy, Mike, Balepur, Nishant, Bareket, Dan, Dalvi, Bhavana, Feldman, Sergey, Haddad, Dany, Hwang, Jena D., Jansen, Peter, Kishore, Varsha, Majumder, Bodhisattwa Prasad, Naik, Aakanksha, Rahamimov, Sigal, Richardson, Kyle, Singh, Amanpreet, Surana, Harshit, Tiktinsky, Aryeh, Vasu, Rosni, Wiener, Guy, Anastasiades, Chloe, Candra, Stefan, Dunkelberger, Jason, Emery, Dan, Evans, Rob, Hamada, Malachi, Huff, Regan, Kinney, Rodney, Latzke, Matt, Lochner, Jaron, Lozano-Aguilera, Ruben, Nguyen, Cecile, Rao, Smita, Tanaka, Amber, Vlahos, Brooke, Clark, Peter, Downey, Doug, Goldberg, Yoav, Sabharwal, Ashish, Weld, Daniel S.
Format: Preprint
Published: 2025
Subjects:
Artificial Intelligence
Computation and Language
Online Access:
Acceder al recurso
Tags: Add Tag
No Tags, Be the first to tag this record!
  • Cite this
  • Text this
  • Email this
  • Print
  • Export Record
    • Export to RefWorks
    • Export to EndNoteWeb
    • Export to EndNote
  • Save to List
  • Permanent link
  • Holdings
  • Description
  • Comments
  • Similar Items
  • Staff View

Internet

https://arxiv.org/abs/2510.21652

Similar Items

  • Understanding Usage and Engagement in AI-Powered Scientific Research Tools: The Asta Interaction Dataset
    by: Haddad, Dany, et al.
    Published: (2026)
  • Ai2 Scholar QA: Organized Literature Synthesis with Attribution
    by: Singh, Amanpreet, et al.
    Published: (2025)
  • Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
    by: Hwang, Jena D., et al.
    Published: (2026)
  • Data-driven Discovery with Large Generative Models
    by: Majumder, Bodhisattwa Prasad, et al.
    Published: (2024)
  • Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
    by: Balepur, Nishant, et al.
    Published: (2026)
Universidad del Mar
Universidad del MarSistema Bibliotecario de la Universidad del MarDescubridor Institucional UMARImplementación y desarrollo: Mtro. Carlos Alonso Albores Pérez
InicioBúsqueda avanzadaExplorar
Visitas al Descubridor: 33,245© 2026 Universidad del Mar