Build the web for agents, not agents for the web
Fuente:
arXiv
Salvato in:
| Autori principali: | Lù, Xing Han, Kamath, Gaurav, Mosbach, Marius, Reddy, Siva |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Value Drifts: Tracing Value Alignment During LLM Post-Training
di: Bhatia, Mehar, et al.
Pubblicazione: (2025)
di: Bhatia, Mehar, et al.
Pubblicazione: (2025)
Forecasting Downstream Performance of LLMs With Proxy Metrics
di: Patel, Arkil, et al.
Pubblicazione: (2026)
di: Patel, Arkil, et al.
Pubblicazione: (2026)
BEARCUBS: A benchmark for computer-using web agents
di: Song, Yixiao, et al.
Pubblicazione: (2025)
di: Song, Yixiao, et al.
Pubblicazione: (2025)
Not All Data Are Unlearned Equally
di: Krishnan, Aravind, et al.
Pubblicazione: (2025)
di: Krishnan, Aravind, et al.
Pubblicazione: (2025)
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
di: Lù, Xing Han, et al.
Pubblicazione: (2024)
di: Lù, Xing Han, et al.
Pubblicazione: (2024)
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
di: Rizvi-Martel, Michael, et al.
Pubblicazione: (2026)
di: Rizvi-Martel, Michael, et al.
Pubblicazione: (2026)
Do Generalisation Results Generalise?
di: Boglioni, Matteo, et al.
Pubblicazione: (2025)
di: Boglioni, Matteo, et al.
Pubblicazione: (2025)
Faithfulness Measurable Masked Language Models
di: Madsen, Andreas, et al.
Pubblicazione: (2023)
di: Madsen, Andreas, et al.
Pubblicazione: (2023)
Language Models Largely Exhibit Human-like Constituent Ordering Preferences
di: Tur, Ada Defne, et al.
Pubblicazione: (2025)
di: Tur, Ada Defne, et al.
Pubblicazione: (2025)
Understanding the Influence of Synthetic Data for Text Embedders
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2025)
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2025)
Scope Ambiguities in Large Language Models
di: Kamath, Gaurav, et al.
Pubblicazione: (2024)
di: Kamath, Gaurav, et al.
Pubblicazione: (2024)
Are self-explanations from Large Language Models faithful?
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
Structured Distillation of Web Agent Capabilities Enables Generalization
di: Lù, Xing Han, et al.
Pubblicazione: (2026)
di: Lù, Xing Han, et al.
Pubblicazione: (2026)
What explains the success of cross-modal fine-tuning with ORCA?
di: García-de-Herreros, Paloma, et al.
Pubblicazione: (2024)
di: García-de-Herreros, Paloma, et al.
Pubblicazione: (2024)
SafeArena: Evaluating the Safety of Autonomous Web Agents
di: Tur, Ada Defne, et al.
Pubblicazione: (2025)
di: Tur, Ada Defne, et al.
Pubblicazione: (2025)
CARD: Towards Conditional Design of Multi-agent Topological Structures
di: Wu, Tongtong, et al.
Pubblicazione: (2026)
di: Wu, Tongtong, et al.
Pubblicazione: (2026)
Operationalising the Superficial Alignment Hypothesis via Task Complexity
di: Vergara-Browne, Tomás, et al.
Pubblicazione: (2026)
di: Vergara-Browne, Tomás, et al.
Pubblicazione: (2026)
Essential-Web v1.0: 24T tokens of organized web data
di: AI, Essential, et al.
Pubblicazione: (2025)
di: AI, Essential, et al.
Pubblicazione: (2025)
Interpretability Needs a New Paradigm
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
LLM2Vec-Gen: Generative Embeddings from Large Language Models
di: BehnamGhader, Parishad, et al.
Pubblicazione: (2026)
di: BehnamGhader, Parishad, et al.
Pubblicazione: (2026)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
di: Lù, Xing Han, et al.
Pubblicazione: (2025)
di: Lù, Xing Han, et al.
Pubblicazione: (2025)
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
di: BehnamGhader, Parishad, et al.
Pubblicazione: (2024)
di: BehnamGhader, Parishad, et al.
Pubblicazione: (2024)
The emergence of numerical representations in communicating artificial agents
di: Mihai, Daniela, et al.
Pubblicazione: (2026)
di: Mihai, Daniela, et al.
Pubblicazione: (2026)
Understanding the planning of LLM agents: A survey
di: Huang, Xu, et al.
Pubblicazione: (2024)
di: Huang, Xu, et al.
Pubblicazione: (2024)
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
di: Hameed, Marawan Gamal Abdel, et al.
Pubblicazione: (2024)
di: Hameed, Marawan Gamal Abdel, et al.
Pubblicazione: (2024)
VinePPO: Refining Credit Assignment in RL Training of LLMs
di: Kazemnejad, Amirhossein, et al.
Pubblicazione: (2024)
di: Kazemnejad, Amirhossein, et al.
Pubblicazione: (2024)
Multi-agent Architecture Search via Agentic Supernet
di: Zhang, Guibin, et al.
Pubblicazione: (2025)
di: Zhang, Guibin, et al.
Pubblicazione: (2025)
Aviary: training language agents on challenging scientific tasks
di: Narayanan, Siddharth, et al.
Pubblicazione: (2024)
di: Narayanan, Siddharth, et al.
Pubblicazione: (2024)
BRIDGE: Predicting Human Task Completion Time From Model Performance
di: Liu, Fengyuan, et al.
Pubblicazione: (2026)
di: Liu, Fengyuan, et al.
Pubblicazione: (2026)
Lucy: edgerunning agentic web search on mobile with machine generated task vectors
di: Dao, Alan, et al.
Pubblicazione: (2025)
di: Dao, Alan, et al.
Pubblicazione: (2025)
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
di: Jain, Dhruv, et al.
Pubblicazione: (2025)
di: Jain, Dhruv, et al.
Pubblicazione: (2025)
Safe Multi-agent Reinforcement Learning with Natural Language Constraints
di: Wang, Ziyan, et al.
Pubblicazione: (2024)
di: Wang, Ziyan, et al.
Pubblicazione: (2024)
Agentic-imodels: Evolving agentic interpretability tools via autoresearch
di: Singh, Chandan, et al.
Pubblicazione: (2026)
di: Singh, Chandan, et al.
Pubblicazione: (2026)
Humans and LLMs Diverge on Probabilistic Inferences
di: Kamath, Gaurav, et al.
Pubblicazione: (2026)
di: Kamath, Gaurav, et al.
Pubblicazione: (2026)
The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents
di: Lu, Xing Han, et al.
Pubblicazione: (2023)
di: Lu, Xing Han, et al.
Pubblicazione: (2023)
Automated test generation to evaluate tool-augmented LLMs as conversational AI agents
di: Arcadinho, Samuel, et al.
Pubblicazione: (2024)
di: Arcadinho, Samuel, et al.
Pubblicazione: (2024)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2025)
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2025)
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
di: Aghajohari, Milad, et al.
Pubblicazione: (2025)
di: Aghajohari, Milad, et al.
Pubblicazione: (2025)
The Denario project: Deep knowledge AI agents for scientific discovery
di: Villaescusa-Navarro, Francisco, et al.
Pubblicazione: (2025)
di: Villaescusa-Navarro, Francisco, et al.
Pubblicazione: (2025)
Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification
di: Maheshwari, Gaurav, et al.
Pubblicazione: (2026)
di: Maheshwari, Gaurav, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Value Drifts: Tracing Value Alignment During LLM Post-Training
di: Bhatia, Mehar, et al.
Pubblicazione: (2025) -
Forecasting Downstream Performance of LLMs With Proxy Metrics
di: Patel, Arkil, et al.
Pubblicazione: (2026) -
BEARCUBS: A benchmark for computer-using web agents
di: Song, Yixiao, et al.
Pubblicazione: (2025) -
Not All Data Are Unlearned Equally
di: Krishnan, Aravind, et al.
Pubblicazione: (2025) -
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
di: Lù, Xing Han, et al.
Pubblicazione: (2024)