The BrowserGym Ecosystem for Web Agent Research
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | De Chezelles, Thibault Le Sellier, Gasse, Maxime, Drouin, Alexandre, Caccia, Massimo, Boisvert, Léo, Thakkar, Megh, Marty, Tom, Assouel, Rim, Shayegan, Sahar Omidi, Jang, Lawrence Keunho, Lù, Xing Han, Yoran, Ori, Kong, Dehan, Xu, Frank F., Reddy, Siva, Cappart, Quentin, Neubig, Graham, Salakhutdinov, Ruslan, Chapados, Nicolas, Lacoste, Alexandre |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
von: Boisvert, Léo, et al.
Veröffentlicht: (2024)
von: Boisvert, Léo, et al.
Veröffentlicht: (2024)
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
von: Drouin, Alexandre, et al.
Veröffentlicht: (2024)
von: Drouin, Alexandre, et al.
Veröffentlicht: (2024)
FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents
von: Kerboua, Imene, et al.
Veröffentlicht: (2025)
von: Kerboua, Imene, et al.
Veröffentlicht: (2025)
How to Train Your LLM Web Agent: A Statistical Diagnosis
von: Vattikonda, Dheeraj, et al.
Veröffentlicht: (2025)
von: Vattikonda, Dheeraj, et al.
Veröffentlicht: (2025)
LineRetriever: Planning-Aware Observation Reduction for Web Agents
von: Kerboua, Imene, et al.
Veröffentlicht: (2025)
von: Kerboua, Imene, et al.
Veröffentlicht: (2025)
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
von: Boisvert, Léo, et al.
Veröffentlicht: (2025)
von: Boisvert, Léo, et al.
Veröffentlicht: (2025)
Towards a Generic Representation of Combinatorial Problems for Learning-Based Approaches
von: Boisvert, Léo, et al.
Veröffentlicht: (2024)
von: Boisvert, Léo, et al.
Veröffentlicht: (2024)
DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
von: Boisvert, Leo, et al.
Veröffentlicht: (2025)
von: Boisvert, Leo, et al.
Veröffentlicht: (2025)
TACTiS-2: Better, Faster, Simpler Attentional Copulas for Multivariate Time Series
von: Ashok, Arjun, et al.
Veröffentlicht: (2023)
von: Ashok, Arjun, et al.
Veröffentlicht: (2023)
Too Big to Fool: Resisting Deception in Language Models
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
von: Nekoei, Hadi, et al.
Veröffentlicht: (2025)
von: Nekoei, Hadi, et al.
Veröffentlicht: (2025)
Making Retrieval-Augmented Language Models Robust to Irrelevant Context
von: Yoran, Ori, et al.
Veröffentlicht: (2023)
von: Yoran, Ori, et al.
Veröffentlicht: (2023)
Context is Key: A Benchmark for Forecasting with Essential Textual Information
von: Williams, Andrew Robert, et al.
Veröffentlicht: (2024)
von: Williams, Andrew Robert, et al.
Veröffentlicht: (2024)
Preventing Rogue Agents Improves Multi-Agent Collaboration
von: Barbi, Ohav, et al.
Veröffentlicht: (2025)
von: Barbi, Ohav, et al.
Veröffentlicht: (2025)
Visual symbolic mechanisms: Emergent symbol processing in vision language models
von: Assouel, Rim, et al.
Veröffentlicht: (2025)
von: Assouel, Rim, et al.
Veröffentlicht: (2025)
Privileged Information Distillation for Language Models
von: Penaloza, Emiliano, et al.
Veröffentlicht: (2026)
von: Penaloza, Emiliano, et al.
Veröffentlicht: (2026)
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
von: Assouel, Rim, et al.
Veröffentlicht: (2026)
von: Assouel, Rim, et al.
Veröffentlicht: (2026)
Gym-Anything: Turn any Software into an Agent Environment
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2026)
TIC-TAC: A Framework for Improved Covariance Estimation in Deep Heteroscedastic Regression
von: Shukla, Megh, et al.
Veröffentlicht: (2023)
von: Shukla, Megh, et al.
Veröffentlicht: (2023)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2025)
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2025)
From Loops to Oops: Fallback Behaviors of Language Models Under Uncertainty
von: Ivgi, Maor, et al.
Veröffentlicht: (2024)
von: Ivgi, Maor, et al.
Veröffentlicht: (2024)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
von: Jang, Lawrence Keunho, et al.
Veröffentlicht: (2026)
von: Jang, Lawrence Keunho, et al.
Veröffentlicht: (2026)
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
von: Sahu, Gaurav, et al.
Veröffentlicht: (2024)
von: Sahu, Gaurav, et al.
Veröffentlicht: (2024)
Object-centric Binding in Contrastive Language-Image Pretraining
von: Assouel, Rim, et al.
Veröffentlicht: (2025)
von: Assouel, Rim, et al.
Veröffentlicht: (2025)
Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs
von: Ashok, Arjun, et al.
Veröffentlicht: (2025)
von: Ashok, Arjun, et al.
Veröffentlicht: (2025)
Towards Self-Supervised Covariance Estimation in Deep Heteroscedastic Regression
von: Shukla, Megh, et al.
Veröffentlicht: (2025)
von: Shukla, Megh, et al.
Veröffentlicht: (2025)
Binding Visual Features Point by Point
von: Haputhanthri, Udith, et al.
Veröffentlicht: (2026)
von: Haputhanthri, Udith, et al.
Veröffentlicht: (2026)
SimGym: Traffic-Grounded Browser Agents for Offline A/B Testing in E-Commerce
von: Castelo, Alberto, et al.
Veröffentlicht: (2026)
von: Castelo, Alberto, et al.
Veröffentlicht: (2026)
Training Software Engineering Agents and Verifiers with SWE-Gym
von: Pan, Jiayi, et al.
Veröffentlicht: (2024)
von: Pan, Jiayi, et al.
Veröffentlicht: (2024)
An Evaluation of "Choice" as a Selection Tool in the Field of Western History.
von: Boisvert, Marianne
Veröffentlicht: (1977)
von: Boisvert, Marianne
Veröffentlicht: (1977)
Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems
von: Madhusudhan, Nishanth, et al.
Veröffentlicht: (2026)
von: Madhusudhan, Nishanth, et al.
Veröffentlicht: (2026)
Atlas de los pueblos del Asia Meridional y Oriental / Jean Sellier ; cartografia, Bertrand de Brun, Anne Le Fur ; traduccion de Jordi Terre
von: Sellier, Jean
von: Sellier, Jean
Atlas de los pueblos de Africa / Jean Sellier ; cartografia: Bertrand de Brun, Anne Le Fur ; traducción de Alberto Roca Alvarez
von: Sellier, Jean
von: Sellier, Jean
Answering Questions by Meta-Reasoning over Multiple Chains of Thought
von: Yoran, Ori, et al.
Veröffentlicht: (2023)
von: Yoran, Ori, et al.
Veröffentlicht: (2023)
The KoLMogorov Test: Compression by Code Generation
von: Yoran, Ori, et al.
Veröffentlicht: (2025)
von: Yoran, Ori, et al.
Veröffentlicht: (2025)
Mar de Provas no Sahel: interrogar a pedagogia universitária a distância
von: Stéphanie Gasse
Veröffentlicht: (2017)
von: Stéphanie Gasse
Veröffentlicht: (2017)
PKDB163655
von: Gasse, Françoise
Veröffentlicht: (2000)
von: Gasse, Françoise
Veröffentlicht: (2000)
Apriel-1.5-OpenReasoner: RL Post-Training for General-Purpose and Efficient Reasoning
von: Pardinas, Rafael, et al.
Veröffentlicht: (2026)
von: Pardinas, Rafael, et al.
Veröffentlicht: (2026)
MotionMap: Representing Multimodality in Human Pose Forecasting
von: Hosseininejad, Reyhaneh, et al.
Veröffentlicht: (2024)
von: Hosseininejad, Reyhaneh, et al.
Veröffentlicht: (2024)
ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
von: Boisvert, Léo, et al.
Veröffentlicht: (2024) -
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
von: Drouin, Alexandre, et al.
Veröffentlicht: (2024) -
FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents
von: Kerboua, Imene, et al.
Veröffentlicht: (2025) -
How to Train Your LLM Web Agent: A Statistical Diagnosis
von: Vattikonda, Dheeraj, et al.
Veröffentlicht: (2025) -
LineRetriever: Planning-Aware Observation Reduction for Web Agents
von: Kerboua, Imene, et al.
Veröffentlicht: (2025)