WALT: Web Agents that Learn Tools
Fuente:
arXiv
Saved in:
| Main Authors: | Prabhu, Viraj, Dai, Yutong, Fernandez, Matthew, Gu, Jing, Ramakrishnan, Krithika, Luo, Yanqi, Savarese, Silvio, Xiong, Caiming, Li, Junnan, Chen, Zeyuan, Xu, Ran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCUBA: Salesforce Computer Use Benchmark
by: Dai, Yutong, et al.
Published: (2025)
by: Dai, Yutong, et al.
Published: (2025)
CoAct-1: Computer-using Multi-Agent System with Coding Actions
by: Song, Linxin, et al.
Published: (2025)
by: Song, Linxin, et al.
Published: (2025)
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
by: Yang, Luyu, et al.
Published: (2026)
by: Yang, Luyu, et al.
Published: (2026)
Trust but Verify: Programmatic VLM Evaluation in the Wild
by: Prabhu, Viraj, et al.
Published: (2024)
by: Prabhu, Viraj, et al.
Published: (2024)
GTA1: GUI Test-time Scaling Agent
by: Yang, Yan, et al.
Published: (2025)
by: Yang, Yan, et al.
Published: (2025)
Asynchronous Tool Usage for Real-Time Agents
by: Ginart, Antonio A., et al.
Published: (2024)
by: Ginart, Antonio A., et al.
Published: (2024)
W&D:Scaling Parallel Tool Calling for Efficient Deep Research Agents
by: Lin, Xiaoqiang, et al.
Published: (2026)
by: Lin, Xiaoqiang, et al.
Published: (2026)
Enabling High Data Throughput Reinforcement Learning on GPUs: A Domain Agnostic Framework for Data-Driven Scientific Research
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Shared Imagination: LLMs Hallucinate Alike
by: Zhou, Yilun, et al.
Published: (2024)
by: Zhou, Yilun, et al.
Published: (2024)
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
by: Li, Jierui, et al.
Published: (2024)
by: Li, Jierui, et al.
Published: (2024)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
by: Nguyen, Xuan-Phi, et al.
Published: (2025)
by: Nguyen, Xuan-Phi, et al.
Published: (2025)
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
by: Panagopoulou, Artemis, et al.
Published: (2023)
by: Panagopoulou, Artemis, et al.
Published: (2023)
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
by: Pang, Bo, et al.
Published: (2025)
by: Pang, Bo, et al.
Published: (2025)
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
BLIP3o-NEXT: Next Frontier of Native Image Generation
by: Chen, Jiuhai, et al.
Published: (2025)
by: Chen, Jiuhai, et al.
Published: (2025)
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
by: Pang, Bo, et al.
Published: (2025)
by: Pang, Bo, et al.
Published: (2025)
MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
Empowering Time Series Analysis with Synthetic Data: A Survey and Outlook in the Era of Foundation Models
by: Liu, Xu, et al.
Published: (2025)
by: Liu, Xu, et al.
Published: (2025)
Moirai 2.0: When Less Is More for Time Series Forecasting
by: Liu, Chenghao, et al.
Published: (2025)
by: Liu, Chenghao, et al.
Published: (2025)
ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding
by: Xue, Le, et al.
Published: (2023)
by: Xue, Le, et al.
Published: (2023)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
by: Wang, Ziyang, et al.
Published: (2025)
by: Wang, Ziyang, et al.
Published: (2025)
Unified Training of Universal Time Series Forecasting Transformers
by: Woo, Gerald, et al.
Published: (2024)
by: Woo, Gerald, et al.
Published: (2024)
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
by: Prabhakar, Akshara, et al.
Published: (2025)
by: Prabhakar, Akshara, et al.
Published: (2025)
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
by: Qin, Can, et al.
Published: (2024)
by: Qin, Can, et al.
Published: (2024)
REX: Rapid Exploration and eXploitation for AI Agents
by: Murthy, Rithesh, et al.
Published: (2023)
by: Murthy, Rithesh, et al.
Published: (2023)
Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization
by: Yao, Weiran, et al.
Published: (2023)
by: Yao, Weiran, et al.
Published: (2023)
Text2Data: Low-Resource Data Generation with Textual Control
by: Wang, Shiyu, et al.
Published: (2024)
by: Wang, Shiyu, et al.
Published: (2024)
ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning
by: Yue, Murong, et al.
Published: (2025)
by: Yue, Murong, et al.
Published: (2025)
HIVE: Harnessing Human Feedback for Instructional Visual Editing
by: Zhang, Shu, et al.
Published: (2023)
by: Zhang, Shu, et al.
Published: (2023)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
by: Peng, Yun, et al.
Published: (2024)
by: Peng, Yun, et al.
Published: (2024)
Test-Time Adaptation for LLM Agents via Environment Interaction
by: Chen, Arthur, et al.
Published: (2025)
by: Chen, Arthur, et al.
Published: (2025)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
by: Peng, Xiangyu, et al.
Published: (2025)
by: Peng, Xiangyu, et al.
Published: (2025)
Reward Models Identify Consistency, Not Causality
by: Xu, Yuhui, et al.
Published: (2025)
by: Xu, Yuhui, et al.
Published: (2025)
LZ Penalty: An information-theoretic repetition penalty for autoregressive language models
by: Ginart, Antonio A., et al.
Published: (2025)
by: Ginart, Antonio A., et al.
Published: (2025)
CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval
by: Liu, Ye, et al.
Published: (2024)
by: Liu, Ye, et al.
Published: (2024)
GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation
by: Aksu, Taha, et al.
Published: (2024)
by: Aksu, Taha, et al.
Published: (2024)
Entropy-Based Block Pruning for Efficient Large Language Models
by: Yang, Liangwei, et al.
Published: (2025)
by: Yang, Liangwei, et al.
Published: (2025)
Similar Items
-
SCUBA: Salesforce Computer Use Benchmark
by: Dai, Yutong, et al.
Published: (2025) -
CoAct-1: Computer-using Multi-Agent System with Coding Actions
by: Song, Linxin, et al.
Published: (2025) -
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
by: Yang, Luyu, et al.
Published: (2026) -
Trust but Verify: Programmatic VLM Evaluation in the Wild
by: Prabhu, Viraj, et al.
Published: (2024) -
GTA1: GUI Test-time Scaling Agent
by: Yang, Yan, et al.
Published: (2025)