Saved in:
| Main Authors: | Zhang, Xianren, Prasad, Shreyas, Wang, Di, Zeng, Qiuhai, Wang, Suhang, Yan, Wenbo, Hans, Mat |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.15832 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
by: Zeng, Qiuhai, et al.
Published: (2025)
by: Zeng, Qiuhai, et al.
Published: (2025)
Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents
by: Zeng, Jingying, et al.
Published: (2025)
by: Zeng, Jingying, et al.
Published: (2025)
A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction
by: Pan, Ruihao, et al.
Published: (2026)
by: Pan, Ruihao, et al.
Published: (2026)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
DocTER: Evaluating Document-based Knowledge Editing
by: Wu, Suhang, et al.
Published: (2023)
by: Wu, Suhang, et al.
Published: (2023)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
Decoding Time Series with LLMs: A Multi-Agent Framework for Cross-Domain Annotation
by: Lin, Minhua, et al.
Published: (2024)
by: Lin, Minhua, et al.
Published: (2024)
GraphSkill: Documentation-Guided Hierarchical Retrieval-Augmented Coding for Complex Graph Reasoning
by: Wang, Fali, et al.
Published: (2026)
by: Wang, Fali, et al.
Published: (2026)
Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts
by: Li, Aiden Yiliu, et al.
Published: (2026)
by: Li, Aiden Yiliu, et al.
Published: (2026)
MindFlow: Revolutionizing E-commerce Customer Support with Multimodal LLM Agents
by: Gong, Ming, et al.
Published: (2025)
by: Gong, Ming, et al.
Published: (2025)
RISK: A Framework for GUI Agents in E-commerce Risk Management
by: Chen, Renqi, et al.
Published: (2025)
by: Chen, Renqi, et al.
Published: (2025)
ADORE: Autonomous Domain-Oriented Relevance Engine for E-commerce
by: Fang, Zheng, et al.
Published: (2025)
by: Fang, Zheng, et al.
Published: (2025)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
by: Yu, Tao, et al.
Published: (2025)
by: Yu, Tao, et al.
Published: (2025)
WebWalker: Benchmarking LLMs in Web Traversal
by: Wu, Jialong, et al.
Published: (2025)
by: Wu, Jialong, et al.
Published: (2025)
WebSailor: Navigating Super-human Reasoning for Web Agent
by: Li, Kuan, et al.
Published: (2025)
by: Li, Kuan, et al.
Published: (2025)
WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
by: Zhuang, Yuchen, et al.
Published: (2025)
by: Zhuang, Yuchen, et al.
Published: (2025)
Mango: Multi-Agent Web Navigation via Global-View Optimization
by: Tong, Weixi, et al.
Published: (2026)
by: Tong, Weixi, et al.
Published: (2026)
Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
by: Liu, Yinuo, et al.
Published: (2025)
by: Liu, Yinuo, et al.
Published: (2025)
Compass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce Embeddings
by: Ueareeworakul, Pakorn, et al.
Published: (2025)
by: Ueareeworakul, Pakorn, et al.
Published: (2025)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
by: Lei, Fangyu, et al.
Published: (2025)
by: Lei, Fangyu, et al.
Published: (2025)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
by: Kim, Serin, et al.
Published: (2026)
by: Kim, Serin, et al.
Published: (2026)
WebXSkill: Skill Learning for Autonomous Web Agents
by: Wang, Zhaoyang, et al.
Published: (2026)
by: Wang, Zhaoyang, et al.
Published: (2026)
RiskWebWorld: A Realistic Interactive Benchmark for GUI Agents in E-commerce Risk Management
by: Chen, Renqi, et al.
Published: (2026)
by: Chen, Renqi, et al.
Published: (2026)
GroundAct: Can LLM Agents Ground Actions in Environmental States?
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
by: Zhang, Zhiwei, et al.
Published: (2025)
by: Zhang, Zhiwei, et al.
Published: (2025)
Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
by: Wang, Yumeng, et al.
Published: (2025)
by: Wang, Yumeng, et al.
Published: (2025)
MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents
by: Wang, Shouju, et al.
Published: (2026)
by: Wang, Shouju, et al.
Published: (2026)
QueryNER: Segmentation of E-commerce Queries
by: Palen-Michel, Chester, et al.
Published: (2024)
by: Palen-Michel, Chester, et al.
Published: (2024)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
by: Deng, Shihan, et al.
Published: (2024)
by: Deng, Shihan, et al.
Published: (2024)
DynaWeb: Model-Based Reinforcement Learning of Web Agents
by: Ding, Hang, et al.
Published: (2026)
by: Ding, Hang, et al.
Published: (2026)
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
by: Yan, Jianhao, et al.
Published: (2024)
by: Yan, Jianhao, et al.
Published: (2024)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
by: Zhang, Ziyun, et al.
Published: (2026)
by: Zhang, Ziyun, et al.
Published: (2026)
ANGO: A Next-Level Evaluation Benchmark For Generation-Oriented Language Models In Chinese Domain
by: Wang, Bingchao
Published: (2024)
by: Wang, Bingchao
Published: (2024)
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
by: Kong, Fanheng, et al.
Published: (2026)
by: Kong, Fanheng, et al.
Published: (2026)
SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints
by: Li, Zekun, et al.
Published: (2025)
by: Li, Zekun, et al.
Published: (2025)
General2Specialized LLMs Translation for E-commerce
by: Chen, Kaidi, et al.
Published: (2024)
by: Chen, Kaidi, et al.
Published: (2024)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
Similar Items
-
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
by: Zeng, Qiuhai, et al.
Published: (2025) -
Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents
by: Zeng, Jingying, et al.
Published: (2025) -
A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction
by: Pan, Ruihao, et al.
Published: (2026) -
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026) -
DocTER: Evaluating Document-based Knowledge Editing
by: Wu, Suhang, et al.
Published: (2023)