AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
Fuente:
arXiv
Salvato in:
| Autori principali: | Yoran, Ori, Amouyal, Samuel Joseph, Malaviya, Chaitanya, Bogin, Ben, Press, Ofir, Berant, Jonathan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Answering Questions by Meta-Reasoning over Multiple Chains of Thought
di: Yoran, Ori, et al.
Pubblicazione: (2023)
di: Yoran, Ori, et al.
Pubblicazione: (2023)
Making Retrieval-Augmented Language Models Robust to Irrelevant Context
di: Yoran, Ori, et al.
Pubblicazione: (2023)
di: Yoran, Ori, et al.
Pubblicazione: (2023)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
di: Ivgi, Maor, et al.
Pubblicazione: (2024)
di: Ivgi, Maor, et al.
Pubblicazione: (2024)
Large Language Models for Psycholinguistic Plausibility Pretesting
di: Amouyal, Samuel Joseph, et al.
Pubblicazione: (2024)
di: Amouyal, Samuel Joseph, et al.
Pubblicazione: (2024)
When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models
di: Amouyal, Samuel Joseph, et al.
Pubblicazione: (2025)
di: Amouyal, Samuel Joseph, et al.
Pubblicazione: (2025)
Comparing Human and Language Models Sentence Processing Difficulties on Complex Structures
di: Amouyal, Samuel Joseph, et al.
Pubblicazione: (2025)
di: Amouyal, Samuel Joseph, et al.
Pubblicazione: (2025)
CiteME: Can Language Models Accurately Cite Scientific Claims?
di: Press, Ori, et al.
Pubblicazione: (2024)
di: Press, Ori, et al.
Pubblicazione: (2024)
Preventing Rogue Agents Improves Multi-Agent Collaboration
di: Barbi, Ohav, et al.
Pubblicazione: (2025)
di: Barbi, Ohav, et al.
Pubblicazione: (2025)
DOLOMITES: Domain-Specific Long-Form Methodical Tasks
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2024)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2024)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
di: Wadhwa, Manya, et al.
Pubblicazione: (2025)
di: Wadhwa, Manya, et al.
Pubblicazione: (2025)
A Simple Joint Model for Improved Contextual Neural Lemmatization
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2019)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2019)
VideoGameBench: Can Vision-Language Models complete popular video games?
di: Zhang, Alex L., et al.
Pubblicazione: (2025)
di: Zhang, Alex L., et al.
Pubblicazione: (2025)
What if you said that differently?: How Explanation Formats Affect Human Feedback Efficacy and User Perception
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
di: Bharadwaj, Anirudh, et al.
Pubblicazione: (2025)
di: Bharadwaj, Anirudh, et al.
Pubblicazione: (2025)
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
di: Deng, Xiang, et al.
Pubblicazione: (2025)
di: Deng, Xiang, et al.
Pubblicazione: (2025)
EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments
di: Liu, Zefang, et al.
Pubblicazione: (2025)
di: Liu, Zefang, et al.
Pubblicazione: (2025)
Retrieval-Pretrained Transformer: Long-range Language Modeling with Self-retrieval
di: Rubin, Ohad, et al.
Pubblicazione: (2023)
di: Rubin, Ohad, et al.
Pubblicazione: (2023)
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
di: Yifei, Li S., et al.
Pubblicazione: (2025)
di: Yifei, Li S., et al.
Pubblicazione: (2025)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
di: Jang, Lawrence Keunho, et al.
Pubblicazione: (2026)
di: Jang, Lawrence Keunho, et al.
Pubblicazione: (2026)
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2025)
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2025)
Leveraging In-Context Learning for Language Model Agents
di: Gupta, Shivanshu, et al.
Pubblicazione: (2025)
di: Gupta, Shivanshu, et al.
Pubblicazione: (2025)
Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2024)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2024)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
di: Koh, Jing Yu, et al.
Pubblicazione: (2024)
di: Koh, Jing Yu, et al.
Pubblicazione: (2024)
SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
di: Bogin, Ben, et al.
Pubblicazione: (2024)
di: Bogin, Ben, et al.
Pubblicazione: (2024)
Leveraging Code to Improve In-context Learning for Semantic Parsing
di: Bogin, Ben, et al.
Pubblicazione: (2023)
di: Bogin, Ben, et al.
Pubblicazione: (2023)
The KoLMogorov Test: Compression by Code Generation
di: Yoran, Ori, et al.
Pubblicazione: (2025)
di: Yoran, Ori, et al.
Pubblicazione: (2025)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
di: Long, Xiang, et al.
Pubblicazione: (2026)
di: Long, Xiang, et al.
Pubblicazione: (2026)
GeoBenchX: Benchmarking LLMs in Agent Solving Multistep Geospatial Tasks
di: Krechetova, Varvara, et al.
Pubblicazione: (2025)
di: Krechetova, Varvara, et al.
Pubblicazione: (2025)
ClawBench: Can AI Agents Complete Everyday Online Tasks?
di: Zhang, Yuxuan, et al.
Pubblicazione: (2026)
di: Zhang, Yuxuan, et al.
Pubblicazione: (2026)
Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven Priors
di: Amos, Ido, et al.
Pubblicazione: (2023)
di: Amos, Ido, et al.
Pubblicazione: (2023)
Weakly Supervised Text-to-SQL Parsing through Question Decomposition
di: Wolfson, Tomer, et al.
Pubblicazione: (2021)
di: Wolfson, Tomer, et al.
Pubblicazione: (2021)
ExpertQA: Expert-Curated Questions and Attributed Answers
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
di: Wang, Qiyao, et al.
Pubblicazione: (2026)
di: Wang, Qiyao, et al.
Pubblicazione: (2026)
WebArena: A Realistic Web Environment for Building Autonomous Agents
di: Zhou, Shuyan, et al.
Pubblicazione: (2023)
di: Zhou, Shuyan, et al.
Pubblicazione: (2023)
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
di: Jimenez, Carlos E., et al.
Pubblicazione: (2023)
di: Jimenez, Carlos E., et al.
Pubblicazione: (2023)
Large Language Models Can Self-Improve At Web Agent Tasks
di: Patel, Ajay, et al.
Pubblicazione: (2024)
di: Patel, Ajay, et al.
Pubblicazione: (2024)
On Reference (In-)Determinacy in Natural Language Inference
di: Chen, Sihao, et al.
Pubblicazione: (2025)
di: Chen, Sihao, et al.
Pubblicazione: (2025)
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
di: Jain, Dhruv, et al.
Pubblicazione: (2025)
di: Jain, Dhruv, et al.
Pubblicazione: (2025)
Self-Execution Simulation Improves Coding Models
di: Maimon, Gallil, et al.
Pubblicazione: (2026)
di: Maimon, Gallil, et al.
Pubblicazione: (2026)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
di: Lee, Gyubok, et al.
Pubblicazione: (2025)
di: Lee, Gyubok, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Answering Questions by Meta-Reasoning over Multiple Chains of Thought
di: Yoran, Ori, et al.
Pubblicazione: (2023) -
Making Retrieval-Augmented Language Models Robust to Irrelevant Context
di: Yoran, Ori, et al.
Pubblicazione: (2023) -
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
di: Ivgi, Maor, et al.
Pubblicazione: (2024) -
Large Language Models for Psycholinguistic Plausibility Pretesting
di: Amouyal, Samuel Joseph, et al.
Pubblicazione: (2024) -
When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models
di: Amouyal, Samuel Joseph, et al.
Pubblicazione: (2025)