AgentA/B: Automated and Scalable Web A/BTesting with Interactive LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Yuxuan, Hsu, Ting-Yao, Gu, Hansu, Cui, Limeng, Xie, Yaochen, Headden, William, Yao, Bingsheng, Veeragouni, Akash, Liu, Jiapeng, Nag, Sreyashi, Wang, Jessie, Wang, Dakuo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025)
by: Sun, Lu, et al.
Published: (2025)
DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
by: Yao, Bingsheng, et al.
Published: (2025)
by: Yao, Bingsheng, et al.
Published: (2025)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
by: Chen, Jiaju, et al.
Published: (2025)
by: Chen, Jiaju, et al.
Published: (2025)
Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
by: Yao, Bingsheng, et al.
Published: (2025)
by: Yao, Bingsheng, et al.
Published: (2025)
Reasoning with Graphs: Structuring Implicit Knowledge to Enhance LLMs Reasoning
by: Han, Haoyu, et al.
Published: (2025)
by: Han, Haoyu, et al.
Published: (2025)
From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents
by: Yao, Bingsheng, et al.
Published: (2026)
by: Yao, Bingsheng, et al.
Published: (2026)
"I Like Sunnie More Than I Expected!": Exploring User Expectation and Perception of an Anthropomorphic LLM-based Conversational Agent for Well-Being Support
by: Wu, Siyi, et al.
Published: (2024)
by: Wu, Siyi, et al.
Published: (2024)
Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
User Interaction Patterns and Breakdowns in Conversing with LLM-Powered Voice Assistants
by: Mahmood, Amama, et al.
Published: (2023)
by: Mahmood, Amama, et al.
Published: (2023)
Secret Use of Large Language Model (LLM)
by: Zhang, Zhiping, et al.
Published: (2024)
by: Zhang, Zhiping, et al.
Published: (2024)
Bridging Knowledge Gaps in Clinical AI: An Activity Theory Perspective on Interdisciplinary Data Work for Telehealth
by: Yao, Bingsheng, et al.
Published: (2024)
by: Yao, Bingsheng, et al.
Published: (2024)
Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Striking a Balance: Evaluating How Aggregations of Multiple Forecasts Impact Judgment Under Uncertainty
by: Zou, Ruishi, et al.
Published: (2024)
by: Zou, Ruishi, et al.
Published: (2024)
Exploring Parent's Needs for Children-Centered AI to Support Preschoolers' Interactive Storytelling and Reading Activities
by: Sun, Yuling, et al.
Published: (2024)
by: Sun, Yuling, et al.
Published: (2024)
"It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents
by: Zhang, Zhiping, et al.
Published: (2023)
by: Zhang, Zhiping, et al.
Published: (2023)
WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
by: Tang, Jingyu, et al.
Published: (2025)
by: Tang, Jingyu, et al.
Published: (2025)
EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association
by: Wang, Weiqi, et al.
Published: (2025)
by: Wang, Weiqi, et al.
Published: (2025)
StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children's Story-Based Learning
by: Chen, Jiaju, et al.
Published: (2023)
by: Chen, Jiaju, et al.
Published: (2023)
"I Wish There Were an AI": Challenges and AI Potential in Cancer Patient-Provider Communication
by: Yang, Ziqi, et al.
Published: (2024)
by: Yang, Ziqi, et al.
Published: (2024)
Human and LLM-Based Voice Assistant Interaction: An Analytical Framework for User Verbal and Nonverbal Behaviors
by: Chan, Szeyi, et al.
Published: (2024)
by: Chan, Szeyi, et al.
Published: (2024)
Who Changed the Destiny of Rural Students, and How?: Unpacking ICT-Mediated Remote Education in Rural China
by: Sun, Yuling, et al.
Published: (2024)
by: Sun, Yuling, et al.
Published: (2024)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
by: Zhang, Ziyun, et al.
Published: (2026)
by: Zhang, Ziyun, et al.
Published: (2026)
Human-Centered Privacy Research in the Age of Large Language Models
by: Li, Tianshi, et al.
Published: (2024)
by: Li, Tianshi, et al.
Published: (2024)
More Samples or More Prompts? Exploring Effective In-Context Sampling for LLM Few-Shot Prompt Engineering
by: Yao, Bingsheng, et al.
Published: (2023)
by: Yao, Bingsheng, et al.
Published: (2023)
Learning with Less: Knowledge Distillation from Large Language Models via Unlabeled Data
by: Li, Juanhui, et al.
Published: (2024)
by: Li, Juanhui, et al.
Published: (2024)
WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark
by: Yuan, Peng, et al.
Published: (2026)
by: Yuan, Peng, et al.
Published: (2026)
Characterizing LLM-Empowered Personalized Story-Reading and Interaction for Children: Insights from Multi-Stakeholder Perspectives
by: Chen, Jiaju, et al.
Published: (2025)
by: Chen, Jiaju, et al.
Published: (2025)
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework
by: Pisano, Matthew, et al.
Published: (2023)
by: Pisano, Matthew, et al.
Published: (2023)
Toward Metaphor-Fluid Conversation Design for Voice User Interfaces
by: Desai, Smit, et al.
Published: (2025)
by: Desai, Smit, et al.
Published: (2025)
WebXSkill: Skill Learning for Autonomous Web Agents
by: Wang, Zhaoyang, et al.
Published: (2026)
by: Wang, Zhaoyang, et al.
Published: (2026)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
by: Yu, Tao, et al.
Published: (2025)
by: Yu, Tao, et al.
Published: (2025)
AutoWebGLM: A Large Language Model-based Web Navigating Agent
by: Lai, Hanyu, et al.
Published: (2024)
by: Lai, Hanyu, et al.
Published: (2024)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
SynthAgent: Adapting Web Agents with Synthetic Supervision
by: Wang, Zhaoyang, et al.
Published: (2025)
by: Wang, Zhaoyang, et al.
Published: (2025)
WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
by: Wei, Zhepei, et al.
Published: (2025)
by: Wei, Zhepei, et al.
Published: (2025)
Large Language Models with Temporal Reasoning for Longitudinal Clinical Summarization and Prediction
by: Kruse, Maya, et al.
Published: (2025)
by: Kruse, Maya, et al.
Published: (2025)
Similar Items
-
UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design
by: Lu, Yuxuan, et al.
Published: (2025) -
UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
by: Lu, Yuxuan, et al.
Published: (2025) -
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025) -
DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
by: Yao, Bingsheng, et al.
Published: (2025) -
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
by: Chen, Jiaju, et al.
Published: (2025)