Build the web for agents, not agents for the web
Fuente:
arXiv
Saved in:
| Main Authors: | Lù, Xing Han, Kamath, Gaurav, Mosbach, Marius, Reddy, Siva |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Value Drifts: Tracing Value Alignment During LLM Post-Training
by: Bhatia, Mehar, et al.
Published: (2025)
by: Bhatia, Mehar, et al.
Published: (2025)
Forecasting Downstream Performance of LLMs With Proxy Metrics
by: Patel, Arkil, et al.
Published: (2026)
by: Patel, Arkil, et al.
Published: (2026)
BEARCUBS: A benchmark for computer-using web agents
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
Not All Data Are Unlearned Equally
by: Krishnan, Aravind, et al.
Published: (2025)
by: Krishnan, Aravind, et al.
Published: (2025)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
by: Lù, Xing Han, et al.
Published: (2024)
by: Lù, Xing Han, et al.
Published: (2024)
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
by: Rizvi-Martel, Michael, et al.
Published: (2026)
by: Rizvi-Martel, Michael, et al.
Published: (2026)
Do Generalisation Results Generalise?
by: Boglioni, Matteo, et al.
Published: (2025)
by: Boglioni, Matteo, et al.
Published: (2025)
Faithfulness Measurable Masked Language Models
by: Madsen, Andreas, et al.
Published: (2023)
by: Madsen, Andreas, et al.
Published: (2023)
Language Models Largely Exhibit Human-like Constituent Ordering Preferences
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
Understanding the Influence of Synthetic Data for Text Embedders
by: Springer, Jacob Mitchell, et al.
Published: (2025)
by: Springer, Jacob Mitchell, et al.
Published: (2025)
Scope Ambiguities in Large Language Models
by: Kamath, Gaurav, et al.
Published: (2024)
by: Kamath, Gaurav, et al.
Published: (2024)
Are self-explanations from Large Language Models faithful?
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Structured Distillation of Web Agent Capabilities Enables Generalization
by: Lù, Xing Han, et al.
Published: (2026)
by: Lù, Xing Han, et al.
Published: (2026)
What explains the success of cross-modal fine-tuning with ORCA?
by: García-de-Herreros, Paloma, et al.
Published: (2024)
by: García-de-Herreros, Paloma, et al.
Published: (2024)
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
CARD: Towards Conditional Design of Multi-agent Topological Structures
by: Wu, Tongtong, et al.
Published: (2026)
by: Wu, Tongtong, et al.
Published: (2026)
Operationalising the Superficial Alignment Hypothesis via Task Complexity
by: Vergara-Browne, Tomás, et al.
Published: (2026)
by: Vergara-Browne, Tomás, et al.
Published: (2026)
Essential-Web v1.0: 24T tokens of organized web data
by: AI, Essential, et al.
Published: (2025)
by: AI, Essential, et al.
Published: (2025)
Interpretability Needs a New Paradigm
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
LLM2Vec-Gen: Generative Embeddings from Large Language Models
by: BehnamGhader, Parishad, et al.
Published: (2026)
by: BehnamGhader, Parishad, et al.
Published: (2026)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
by: BehnamGhader, Parishad, et al.
Published: (2024)
by: BehnamGhader, Parishad, et al.
Published: (2024)
The emergence of numerical representations in communicating artificial agents
by: Mihai, Daniela, et al.
Published: (2026)
by: Mihai, Daniela, et al.
Published: (2026)
Understanding the planning of LLM agents: A survey
by: Huang, Xu, et al.
Published: (2024)
by: Huang, Xu, et al.
Published: (2024)
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2024)
by: Hameed, Marawan Gamal Abdel, et al.
Published: (2024)
VinePPO: Refining Credit Assignment in RL Training of LLMs
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
Multi-agent Architecture Search via Agentic Supernet
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
Aviary: training language agents on challenging scientific tasks
by: Narayanan, Siddharth, et al.
Published: (2024)
by: Narayanan, Siddharth, et al.
Published: (2024)
BRIDGE: Predicting Human Task Completion Time From Model Performance
by: Liu, Fengyuan, et al.
Published: (2026)
by: Liu, Fengyuan, et al.
Published: (2026)
Lucy: edgerunning agentic web search on mobile with machine generated task vectors
by: Dao, Alan, et al.
Published: (2025)
by: Dao, Alan, et al.
Published: (2025)
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
by: Jain, Dhruv, et al.
Published: (2025)
by: Jain, Dhruv, et al.
Published: (2025)
Safe Multi-agent Reinforcement Learning with Natural Language Constraints
by: Wang, Ziyan, et al.
Published: (2024)
by: Wang, Ziyan, et al.
Published: (2024)
Agentic-imodels: Evolving agentic interpretability tools via autoresearch
by: Singh, Chandan, et al.
Published: (2026)
by: Singh, Chandan, et al.
Published: (2026)
Humans and LLMs Diverge on Probabilistic Inferences
by: Kamath, Gaurav, et al.
Published: (2026)
by: Kamath, Gaurav, et al.
Published: (2026)
The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents
by: Lu, Xing Han, et al.
Published: (2023)
by: Lu, Xing Han, et al.
Published: (2023)
Automated test generation to evaluate tool-augmented LLMs as conversational AI agents
by: Arcadinho, Samuel, et al.
Published: (2024)
by: Arcadinho, Samuel, et al.
Published: (2024)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
by: Marjanović, Sara Vera, et al.
Published: (2025)
by: Marjanović, Sara Vera, et al.
Published: (2025)
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
by: Aghajohari, Milad, et al.
Published: (2025)
by: Aghajohari, Milad, et al.
Published: (2025)
The Denario project: Deep knowledge AI agents for scientific discovery
by: Villaescusa-Navarro, Francisco, et al.
Published: (2025)
by: Villaescusa-Navarro, Francisco, et al.
Published: (2025)
Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification
by: Maheshwari, Gaurav, et al.
Published: (2026)
by: Maheshwari, Gaurav, et al.
Published: (2026)
Similar Items
-
Value Drifts: Tracing Value Alignment During LLM Post-Training
by: Bhatia, Mehar, et al.
Published: (2025) -
Forecasting Downstream Performance of LLMs With Proxy Metrics
by: Patel, Arkil, et al.
Published: (2026) -
BEARCUBS: A benchmark for computer-using web agents
by: Song, Yixiao, et al.
Published: (2025) -
Not All Data Are Unlearned Equally
by: Krishnan, Aravind, et al.
Published: (2025) -
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
by: Lù, Xing Han, et al.
Published: (2024)