WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Fuente:
arXiv
Saved in:
| Main Authors: | Drouin, Alexandre, Gasse, Maxime, Caccia, Massimo, Laradji, Issam H., Del Verme, Manuel, Marty, Tom, Boisvert, Léo, Thakkar, Megh, Cappart, Quentin, Vazquez, David, Chapados, Nicolas, Lacoste, Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
by: Boisvert, Léo, et al.
Published: (2024)
by: Boisvert, Léo, et al.
Published: (2024)
The BrowserGym Ecosystem for Web Agent Research
by: De Chezelles, Thibault Le Sellier, et al.
Published: (2024)
by: De Chezelles, Thibault Le Sellier, et al.
Published: (2024)
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
by: Boisvert, Léo, et al.
Published: (2025)
by: Boisvert, Léo, et al.
Published: (2025)
FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents
by: Kerboua, Imene, et al.
Published: (2025)
by: Kerboua, Imene, et al.
Published: (2025)
DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
by: Boisvert, Leo, et al.
Published: (2025)
by: Boisvert, Leo, et al.
Published: (2025)
Towards a Generic Representation of Combinatorial Problems for Learning-Based Approaches
by: Boisvert, Léo, et al.
Published: (2024)
by: Boisvert, Léo, et al.
Published: (2024)
LineRetriever: Planning-Aware Observation Reduction for Web Agents
by: Kerboua, Imene, et al.
Published: (2025)
by: Kerboua, Imene, et al.
Published: (2025)
How to Train Your LLM Web Agent: A Statistical Diagnosis
by: Vattikonda, Dheeraj, et al.
Published: (2025)
by: Vattikonda, Dheeraj, et al.
Published: (2025)
JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
by: Nekoei, Hadi, et al.
Published: (2025)
by: Nekoei, Hadi, et al.
Published: (2025)
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
by: Gurung, Alexander, et al.
Published: (2026)
by: Gurung, Alexander, et al.
Published: (2026)
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
Too Big to Fool: Resisting Deception in Language Models
by: Samsami, Mohammad Reza, et al.
Published: (2024)
by: Samsami, Mohammad Reza, et al.
Published: (2024)
TACTiS-2: Better, Faster, Simpler Attentional Copulas for Multivariate Time Series
by: Ashok, Arjun, et al.
Published: (2023)
by: Ashok, Arjun, et al.
Published: (2023)
Privileged Information Distillation for Language Models
by: Penaloza, Emiliano, et al.
Published: (2026)
by: Penaloza, Emiliano, et al.
Published: (2026)
The Role of Data and Metrics in Measuring Inequality Worldwide. A Tribute to Giovanni Andrea Cornia's Lifelong Work on the World Ginis
by: Ceriani, Lidia, et al.
Published: (2026)
by: Ceriani, Lidia, et al.
Published: (2026)
Context is Key: A Benchmark for Forecasting with Essential Textual Information
by: Williams, Andrew Robert, et al.
Published: (2024)
by: Williams, Andrew Robert, et al.
Published: (2024)
A Guide To Effectively Leveraging LLMs for Low-Resource Text Summarization: Data Augmentation and Semi-supervised Approaches
by: Sahu, Gaurav, et al.
Published: (2024)
by: Sahu, Gaurav, et al.
Published: (2024)
Prompt-based Pseudo-labeling Strategy for Sample-Efficient Semi-Supervised Extractive Summarization
by: Sahu, Gaurav, et al.
Published: (2023)
by: Sahu, Gaurav, et al.
Published: (2023)
Dr-CiK: A Testbed for Foresight-Driven Agents
by: Tang, Yihong, et al.
Published: (2026)
by: Tang, Yihong, et al.
Published: (2026)
DRBench: A Realistic Benchmark for Enterprise Deep Research
by: Abaskohi, Amirhossein, et al.
Published: (2025)
by: Abaskohi, Amirhossein, et al.
Published: (2025)
SSLR: A Semi-Supervised Learning Method for Isolated Sign Language Recognition
by: Algafri, Hasan, et al.
Published: (2025)
by: Algafri, Hasan, et al.
Published: (2025)
An Exact Framework for Solving the Space-Time Dependent TSP
by: Rudich, Isaac, et al.
Published: (2023)
by: Rudich, Isaac, et al.
Published: (2023)
TIC-TAC: A Framework for Improved Covariance Estimation in Deep Heteroscedastic Regression
by: Shukla, Megh, et al.
Published: (2023)
by: Shukla, Megh, et al.
Published: (2023)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
by: Nejadgholi, Isar, et al.
Published: (2025)
by: Nejadgholi, Isar, et al.
Published: (2025)
Learning a Generic Value-Selection Heuristic Inside a Constraint Programming Solver
by: Marty, Tom, et al.
Published: (2023)
by: Marty, Tom, et al.
Published: (2023)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
BlockLLM: Memory-Efficient Adaptation of LLMs by Selecting and Optimizing the Right Coordinate Blocks
by: Ramesh, Amrutha Varshini, et al.
Published: (2024)
by: Ramesh, Amrutha Varshini, et al.
Published: (2024)
Global Rewards in Multi-Agent Deep Reinforcement Learning for Autonomous Mobility on Demand Systems
by: Hoppe, Heiko, et al.
Published: (2023)
by: Hoppe, Heiko, et al.
Published: (2023)
Deep Learning for Data-Driven Districting-and-Routing
by: Ferraz, Arthur, et al.
Published: (2024)
by: Ferraz, Arthur, et al.
Published: (2024)
Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs
by: Ashok, Arjun, et al.
Published: (2025)
by: Ashok, Arjun, et al.
Published: (2025)
Bounded Creativity: Understanding the Restrictions on Creative Work in Advertising Agencies
by: Alexandre Anderson Romeiro
Published: (2015)
by: Alexandre Anderson Romeiro
Published: (2015)
Predicting Poverty
by: Verme, Paolo
Published: (2025)
by: Verme, Paolo
Published: (2025)
Constraints to growth and job creation in low-income commonwealth of independent states countries / Paolo Verme
by: Verme, Paolo
Published: (2006)
by: Verme, Paolo
Published: (2006)
Reseña de "Strong imagination: madness, creativity and human nature" de D. Nettle
by: Giacomo Verme
Published: (2006)
by: Giacomo Verme
Published: (2006)
Reseña de "Brain and culture. Neurobiology, ideology and social change" de Wexler, B.
by: Giacomo Verme
Published: (2007)
by: Giacomo Verme
Published: (2007)
ESTRUCTURA SOCIAL DEL DELFÍN NARÍZ DE BOTELLA Tursiops truncatus (CETACEA: DELPHINIDAE) EN LA COSTA SUROESTE DE LA ISLA DE TENERIFE (ISLAS CANARIAS), ESPAÑA
by: Valeria Verme
Published: (2012)
by: Valeria Verme
Published: (2012)
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
by: Takahashi, Jun, et al.
Published: (2025)
by: Takahashi, Jun, et al.
Published: (2025)
Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs
by: Thakkar, Megh, et al.
Published: (2024)
by: Thakkar, Megh, et al.
Published: (2024)
Fast Convergence of Softmax Policy Mirror Ascent
by: Asad, Reza, et al.
Published: (2024)
by: Asad, Reza, et al.
Published: (2024)
Similar Items
-
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
by: Boisvert, Léo, et al.
Published: (2024) -
The BrowserGym Ecosystem for Web Agent Research
by: De Chezelles, Thibault Le Sellier, et al.
Published: (2024) -
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
by: Boisvert, Léo, et al.
Published: (2025) -
FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents
by: Kerboua, Imene, et al.
Published: (2025) -
DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
by: Boisvert, Leo, et al.
Published: (2025)