MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report)
Fuente:
arXiv
Saved in:
| Main Authors: | Steiner, Aaron, Peeters, Ralph, Bizer, Christian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents
by: Peeters, Ralph, et al.
Published: (2025)
by: Peeters, Ralph, et al.
Published: (2025)
Entity Matching using Large Language Models
by: Peeters, Ralph, et al.
Published: (2023)
by: Peeters, Ralph, et al.
Published: (2023)
Fine-tuning Large Language Models for Entity Matching
by: Steiner, Aaron, et al.
Published: (2024)
by: Steiner, Aaron, et al.
Published: (2024)
Automatic End-to-End Data Integration using Large Language Models
by: Steiner, Aaron, et al.
Published: (2026)
by: Steiner, Aaron, et al.
Published: (2026)
Long Context vs. RAG for LLMs: An Evaluation and Revisits
by: Li, Xinze, et al.
Published: (2024)
by: Li, Xinze, et al.
Published: (2024)
Agri-Query: A Case Study on RAG vs. Long-Context LLMs for Cross-Lingual Technical Question Answering
by: Gun, Julius, et al.
Published: (2025)
by: Gun, Julius, et al.
Published: (2025)
Vector RAG vs LLM-Compiled Wiki: A Preregistered Comparison on a Small Multi-Domain Research
by: Cochran, Theodore O.
Published: (2026)
by: Cochran, Theodore O.
Published: (2026)
Autoregressive vs. Masked Diffusion Language Models: A Controlled Comparison
by: Vicentino, Caio
Published: (2026)
by: Vicentino, Caio
Published: (2026)
A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems
by: Cuconasu, Florin, et al.
Published: (2024)
by: Cuconasu, Florin, et al.
Published: (2024)
Self-Refinement Strategies for LLM-based Product Attribute Value Extraction
by: Brinkmann, Alexander, et al.
Published: (2025)
by: Brinkmann, Alexander, et al.
Published: (2025)
Evaluating Knowledge Generation and Self-Refinement Strategies for LLM-based Column Type Annotation
by: Korini, Keti, et al.
Published: (2025)
by: Korini, Keti, et al.
Published: (2025)
A Systematic Evaluation of LLM Strategies for Mental Health Text Analysis: Fine-tuning vs. Prompt Engineering vs. RAG
by: Kermani, Arshia, et al.
Published: (2025)
by: Kermani, Arshia, et al.
Published: (2025)
RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture
by: Balaguer, Angels, et al.
Published: (2024)
by: Balaguer, Angels, et al.
Published: (2024)
Fine-Tuning vs. RAG for Multi-Hop Question Answering with Novel Knowledge
by: Yang, Zhuoyi, et al.
Published: (2026)
by: Yang, Zhuoyi, et al.
Published: (2026)
A Status Quo Investigation of Large Language Models towards Cost-Effective CFD Automation with OpenFOAMGPT: ChatGPT vs. Qwen vs. Deepseek
by: Wang, Wenkang, et al.
Published: (2025)
by: Wang, Wenkang, et al.
Published: (2025)
Route Before Retrieve: Activating Latent Routing Abilities of LLMs for RAG vs. Long-Context Selection
by: Chen, Yiwen, et al.
Published: (2026)
by: Chen, Yiwen, et al.
Published: (2026)
ChatGPT vs Gemini vs LLaMA on Multilingual Sentiment Analysis
by: Buscemi, Alessio, et al.
Published: (2024)
by: Buscemi, Alessio, et al.
Published: (2024)
Role of Dependency Distance in Text Simplification: A Human vs ChatGPT Simplification Comparison
by: Lee, Sumi, et al.
Published: (2024)
by: Lee, Sumi, et al.
Published: (2024)
Agent Bain vs. Agent McKinsey: A New Text-to-SQL Benchmark for the Business Domain
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG
by: Klila, Jaafer, et al.
Published: (2026)
by: Klila, Jaafer, et al.
Published: (2026)
Using LLMs for the Extraction and Normalization of Product Attribute Values
by: Brinkmann, Alexander, et al.
Published: (2024)
by: Brinkmann, Alexander, et al.
Published: (2024)
ExtractGPT: Exploring the Potential of Large Language Models for Product Attribute Value Extraction
by: Brinkmann, Alexander, et al.
Published: (2023)
by: Brinkmann, Alexander, et al.
Published: (2023)
Dense vs Sparse Pretraining at Tiny Scale: Active-Parameter vs Total-Parameter Matching
by: Wael, Abdalrahman
Published: (2026)
by: Wael, Abdalrahman
Published: (2026)
Can AI-Generated Persuasion Be Detected? Persuaficial Benchmark and AI vs. Human Linguistic Differences
by: Modzelewski, Arkadiusz, et al.
Published: (2026)
by: Modzelewski, Arkadiusz, et al.
Published: (2026)
On the Generalization vs Fidelity Paradox in Knowledge Distillation
by: Ramesh, Suhas Kamasetty, et al.
Published: (2025)
by: Ramesh, Suhas Kamasetty, et al.
Published: (2025)
Text Understanding in GPT-4 vs Humans
by: Shultz, Thomas R., et al.
Published: (2024)
by: Shultz, Thomas R., et al.
Published: (2024)
Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling
by: Li, Margaret, et al.
Published: (2024)
by: Li, Margaret, et al.
Published: (2024)
Scaling Public Health Text Annotation: Zero-Shot Learning vs. Crowdsourcing for Improved Efficiency and Labeling Accuracy
by: Kazari, Kamyar, et al.
Published: (2025)
by: Kazari, Kamyar, et al.
Published: (2025)
Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations
by: Lamparth, Max, et al.
Published: (2024)
by: Lamparth, Max, et al.
Published: (2024)
Reasoning Over Recall: Evaluating the Efficacy of Generalist Architectures vs. Specialized Fine-Tunes in RAG-Based Mental Health Dialogue Systems
by: Kafi, Md Abdullah Al, et al.
Published: (2026)
by: Kafi, Md Abdullah Al, et al.
Published: (2026)
The Data-Dollars Tradeoff: Privacy Harms vs. Economic Risk in Personalized AI Adoption
by: Erlei, Alexander, et al.
Published: (2026)
by: Erlei, Alexander, et al.
Published: (2026)
A Comprehensive Dataset for Human vs. AI Generated Text Detection
by: Roy, Rajarshi, et al.
Published: (2025)
by: Roy, Rajarshi, et al.
Published: (2025)
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
by: Zhang, Siyue, et al.
Published: (2025)
by: Zhang, Siyue, et al.
Published: (2025)
Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs
by: Yang, Zhe, et al.
Published: (2024)
by: Yang, Zhe, et al.
Published: (2024)
BERT vs GPT for financial engineering
by: Sharkey, Edward, et al.
Published: (2024)
by: Sharkey, Edward, et al.
Published: (2024)
Arabizi vs LLMs: Can the Genie Understand the Language of Aladdin?
by: Almaoui, Perla Al, et al.
Published: (2025)
by: Almaoui, Perla Al, et al.
Published: (2025)
Sustained Vowels for Pre- vs Post-Treatment COPD Classification
by: Triantafyllopoulos, Andreas, et al.
Published: (2024)
by: Triantafyllopoulos, Andreas, et al.
Published: (2024)
Generalists vs. Specialists: Evaluating Large Language Models for Urdu
by: Arif, Samee, et al.
Published: (2024)
by: Arif, Samee, et al.
Published: (2024)
Right vs. Right: Can LLMs Make Tough Choices?
by: Yuan, Jiaqing, et al.
Published: (2024)
by: Yuan, Jiaqing, et al.
Published: (2024)
Assessing RAG and HyDE on 1B vs. 4B-Parameter Gemma LLMs for Personal Assistants Integretion
by: Sorstkins, Andrejs
Published: (2025)
by: Sorstkins, Andrejs
Published: (2025)
Similar Items
-
WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents
by: Peeters, Ralph, et al.
Published: (2025) -
Entity Matching using Large Language Models
by: Peeters, Ralph, et al.
Published: (2023) -
Fine-tuning Large Language Models for Entity Matching
by: Steiner, Aaron, et al.
Published: (2024) -
Automatic End-to-End Data Integration using Large Language Models
by: Steiner, Aaron, et al.
Published: (2026) -
Long Context vs. RAG for LLMs: An Evaluation and Revisits
by: Li, Xinze, et al.
Published: (2024)