Revisiting Observation Reduction for Web Agents: Comprehensive Evaluation with a Lightweight Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Enomoto, Masafumi, Obara, Ryoma, Zhang, Haochen, Oyamada, Masafumi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Read More, Think More: Revisiting Observation Reduction for Web Agents
von: Enomoto, Masafumi, et al.
Veröffentlicht: (2026)
von: Enomoto, Masafumi, et al.
Veröffentlicht: (2026)
LightPAL: Lightweight Passage Retrieval for Open Domain Multi-Document Summarization
von: Enomoto, Masafumi, et al.
Veröffentlicht: (2024)
von: Enomoto, Masafumi, et al.
Veröffentlicht: (2024)
Can a Crow Hatch a Falcon? Lineage Matters in Predicting Large Language Model Performance
von: Tamura, Takuya, et al.
Veröffentlicht: (2025)
von: Tamura, Takuya, et al.
Veröffentlicht: (2025)
LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agents
von: Yano, Taro, et al.
Veröffentlicht: (2025)
von: Yano, Taro, et al.
Veröffentlicht: (2025)
cotomi Act: Learning to Automate Work by Watching You
von: Oyamada, Masafumi, et al.
Veröffentlicht: (2026)
von: Oyamada, Masafumi, et al.
Veröffentlicht: (2026)
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
von: Yamauchi, Yusuke, et al.
Veröffentlicht: (2025)
von: Yamauchi, Yusuke, et al.
Veröffentlicht: (2025)
Jellyfish: A Large Language Model for Data Preprocessing
von: Zhang, Haochen, et al.
Veröffentlicht: (2023)
von: Zhang, Haochen, et al.
Veröffentlicht: (2023)
Effective Harness Engineering for Algorithm Discovery with Coding Agents
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2026)
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2026)
Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2025)
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2025)
Can Large Language Models Invent Algorithms to Improve Themselves?: Algorithm Discovery for Recursive Self-Improvement through Reinforcement Learning
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2024)
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2024)
$M^3$ Scaling Law: Optimizing Multi-Epoch, Multi-Lingual, and Multi-Stage Training for Low-Resource Language Models
von: Akimoto, Kosuke, et al.
Veröffentlicht: (2024)
von: Akimoto, Kosuke, et al.
Veröffentlicht: (2024)
Context Quality Matters in Training Fusion-in-Decoder for Extractive Open-Domain Question Answering
von: Akimoto, Kosuke, et al.
Veröffentlicht: (2024)
von: Akimoto, Kosuke, et al.
Veröffentlicht: (2024)
Understanding the Impact of Confidence in Retrieval Augmented Generation: A Case Study in the Medical Domain
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2024)
von: Ozaki, Shintaro, et al.
Veröffentlicht: (2024)
Large Language Models as Data Preprocessors
von: Zhang, Haochen, et al.
Veröffentlicht: (2023)
von: Zhang, Haochen, et al.
Veröffentlicht: (2023)
LineRetriever: Planning-Aware Observation Reduction for Web Agents
von: Kerboua, Imene, et al.
Veröffentlicht: (2025)
von: Kerboua, Imene, et al.
Veröffentlicht: (2025)
On Synthesizing Data for Context Attribution in Question Answering
von: Radevski, Gorjan, et al.
Veröffentlicht: (2025)
von: Radevski, Gorjan, et al.
Veröffentlicht: (2025)
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
TimeWarp: Evaluating Web Agents by Revisiting the Past
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2026)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2026)
DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation
von: Zhang, Enze, et al.
Veröffentlicht: (2025)
von: Zhang, Enze, et al.
Veröffentlicht: (2025)
Best-of-$\infty$ -- Asymptotic Performance of Test-Time LLM Ensembling
von: Komiyama, Junpei, et al.
Veröffentlicht: (2025)
von: Komiyama, Junpei, et al.
Veröffentlicht: (2025)
UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs
von: He, Chaoqun, et al.
Veröffentlicht: (2024)
von: He, Chaoqun, et al.
Veröffentlicht: (2024)
Region4Web: Rethinking Observation Space Granularity for Web Agents
von: Kwon, Donguk, et al.
Veröffentlicht: (2026)
von: Kwon, Donguk, et al.
Veröffentlicht: (2026)
Evaluating Structural Generalization in Neural Machine Translation
von: Kumon, Ryoma, et al.
Veröffentlicht: (2024)
von: Kumon, Ryoma, et al.
Veröffentlicht: (2024)
WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents
von: Peeters, Ralph, et al.
Veröffentlicht: (2025)
von: Peeters, Ralph, et al.
Veröffentlicht: (2025)
Constructive Approach to Bidirectional Influence between Qualia Structure and Language Emergence
von: Taniguchi, Tadahiro, et al.
Veröffentlicht: (2024)
von: Taniguchi, Tadahiro, et al.
Veröffentlicht: (2024)
WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code
von: Lin, Zhiyu, et al.
Veröffentlicht: (2025)
von: Lin, Zhiyu, et al.
Veröffentlicht: (2025)
Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset
von: Yamamoto, Taisei, et al.
Veröffentlicht: (2025)
von: Yamamoto, Taisei, et al.
Veröffentlicht: (2025)
Evaluating Cultural and Social Awareness of LLM Web Agents
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
Comprehensive and Efficient Distillation for Lightweight Sentiment Analysis Models
von: Xie, Guangyu, et al.
Veröffentlicht: (2025)
von: Xie, Guangyu, et al.
Veröffentlicht: (2025)
TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research Agents
von: Chen, Yanyu, et al.
Veröffentlicht: (2026)
von: Chen, Yanyu, et al.
Veröffentlicht: (2026)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
von: Wadhwa, Manya, et al.
Veröffentlicht: (2025)
von: Wadhwa, Manya, et al.
Veröffentlicht: (2025)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
von: Wang, Peng, et al.
Veröffentlicht: (2025)
von: Wang, Peng, et al.
Veröffentlicht: (2025)
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation
von: Zou, Henry Peng, et al.
Veröffentlicht: (2026)
von: Zou, Henry Peng, et al.
Veröffentlicht: (2026)
Direct Quantized Training of Language Models with Stochastic Rounding
von: Zhao, Kaiyan, et al.
Veröffentlicht: (2024)
von: Zhao, Kaiyan, et al.
Veröffentlicht: (2024)
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
von: Fang, Tianqing, et al.
Veröffentlicht: (2025)
Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models
von: Kumon, Ryoma, et al.
Veröffentlicht: (2026)
von: Kumon, Ryoma, et al.
Veröffentlicht: (2026)
Analyzing the Inner Workings of Transformers in Compositional Generalization
von: Kumon, Ryoma, et al.
Veröffentlicht: (2025)
von: Kumon, Ryoma, et al.
Veröffentlicht: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
AI Planning Framework for LLM-Based Web Agents
von: Shahnovsky, Orit, et al.
Veröffentlicht: (2026)
von: Shahnovsky, Orit, et al.
Veröffentlicht: (2026)
Infogent: An Agent-Based Framework for Web Information Aggregation
von: Reddy, Revanth Gangi, et al.
Veröffentlicht: (2024)
von: Reddy, Revanth Gangi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Read More, Think More: Revisiting Observation Reduction for Web Agents
von: Enomoto, Masafumi, et al.
Veröffentlicht: (2026) -
LightPAL: Lightweight Passage Retrieval for Open Domain Multi-Document Summarization
von: Enomoto, Masafumi, et al.
Veröffentlicht: (2024) -
Can a Crow Hatch a Falcon? Lineage Matters in Predicting Large Language Model Performance
von: Tamura, Takuya, et al.
Veröffentlicht: (2025) -
LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agents
von: Yano, Taro, et al.
Veröffentlicht: (2025) -
cotomi Act: Learning to Automate Work by Watching You
von: Oyamada, Masafumi, et al.
Veröffentlicht: (2026)