Efficient Agent Evaluation via Diversity-Guided User Simulation
Fuente:
arXiv
Saved in:
| Main Authors: | Nakash, Itay, Kour, George, Anaby-Tavor, Ateret |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Breaking ReAct Agents: Foot-in-the-Door Attack Will Get You In
by: Nakash, Itay, et al.
Published: (2024)
by: Nakash, Itay, et al.
Published: (2024)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
by: Kour, George, et al.
Published: (2025)
by: Kour, George, et al.
Published: (2025)
Effective Red-Teaming of Policy-Adherent Agents
by: Nakash, Itay, et al.
Published: (2025)
by: Nakash, Itay, et al.
Published: (2025)
Exploring Straightforward Conversational Red-Teaming
by: Kour, George, et al.
Published: (2024)
by: Kour, George, et al.
Published: (2024)
On the Robustness of Agentic Function Calling
by: Rabinovich, Ella, et al.
Published: (2025)
by: Rabinovich, Ella, et al.
Published: (2025)
From Zero to Hero: Cold-Start Anomaly Detection
by: Reiss, Tal, et al.
Published: (2024)
by: Reiss, Tal, et al.
Published: (2024)
CRISP: Complex Reasoning with Interpretable Step-based Plans
by: Vetzler, Matan, et al.
Published: (2025)
by: Vetzler, Matan, et al.
Published: (2025)
What's the Plan? Evaluating and Developing Planning-Aware Techniques for Language Models
by: Hirsch, Eran, et al.
Published: (2024)
by: Hirsch, Eran, et al.
Published: (2024)
Near-Miss: Latent Policy Failure Detection in Agentic Workflows
by: Rabinovich, Ella, et al.
Published: (2026)
by: Rabinovich, Ella, et al.
Published: (2026)
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios
by: Ackerman, Samuel, et al.
Published: (2024)
by: Ackerman, Samuel, et al.
Published: (2024)
OASBuilder: Generating OpenAPI Specifications from Online API Documentation with Large Language Models
by: Lazar, Koren, et al.
Published: (2025)
by: Lazar, Koren, et al.
Published: (2025)
Cognitive-Mental-LLM: Evaluating Reasoning in Large Language Models for Mental Health Prediction via Online Text
by: Patil, Avinash, et al.
Published: (2025)
by: Patil, Avinash, et al.
Published: (2025)
Towards Enforcing Company Policy Adherence in Agentic Workflows
by: Zwerdling, Naama, et al.
Published: (2025)
by: Zwerdling, Naama, et al.
Published: (2025)
RecUserSim: A Realistic and Diverse User Simulator for Evaluating Conversational Recommender Systems
by: Chen, Luyu, et al.
Published: (2025)
by: Chen, Luyu, et al.
Published: (2025)
Decision-aware User Simulation Agent for Evaluating Conversational Recommender Systems
by: Li, Yuan-Chi, et al.
Published: (2026)
by: Li, Yuan-Chi, et al.
Published: (2026)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
by: Burdisso, Sergio, et al.
Published: (2025)
by: Burdisso, Sergio, et al.
Published: (2025)
SpeCrawler: Generating OpenAPI Specifications from API Documentation Using Large Language Models
by: Lazar, Koren, et al.
Published: (2024)
by: Lazar, Koren, et al.
Published: (2024)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
by: Chopra, Harshita, et al.
Published: (2026)
by: Chopra, Harshita, et al.
Published: (2026)
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
by: Zhu, Ming, et al.
Published: (2026)
by: Zhu, Ming, et al.
Published: (2026)
AgentSME for Simulating Diverse Communication Modes in Smart Education
by: Yang, Wen-Xi, et al.
Published: (2025)
by: Yang, Wen-Xi, et al.
Published: (2025)
Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants
by: Nathani, Deepak, et al.
Published: (2026)
by: Nathani, Deepak, et al.
Published: (2026)
Efficient Generation of Diverse Cooperative Agents with World Models
by: Loo, Yi, et al.
Published: (2025)
by: Loo, Yi, et al.
Published: (2025)
AgentInit: Initializing LLM-based Multi-Agent Systems via Diversity and Expertise Orchestration for Effective and Efficient Collaboration
by: Tian, Chunhao, et al.
Published: (2025)
by: Tian, Chunhao, et al.
Published: (2025)
Simulating User Agents for Embodied Conversational-AI
by: Philipov, Daniel, et al.
Published: (2024)
by: Philipov, Daniel, et al.
Published: (2024)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
by: Baharav, Tavor Z., et al.
Published: (2025)
by: Baharav, Tavor Z., et al.
Published: (2025)
An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
by: Pantazopoulos, Georgios, et al.
Published: (2025)
by: Pantazopoulos, Georgios, et al.
Published: (2025)
Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis
by: Chong, Penny, et al.
Published: (2026)
by: Chong, Penny, et al.
Published: (2026)
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
by: Sirdeshmukh, Ved, et al.
Published: (2026)
by: Sirdeshmukh, Ved, et al.
Published: (2026)
USimAgent: Large Language Models for Simulating Search Users
by: Zhang, Erhan, et al.
Published: (2024)
by: Zhang, Erhan, et al.
Published: (2024)
User Behavior Simulation with Large Language Model based Agents
by: Wang, Lei, et al.
Published: (2023)
by: Wang, Lei, et al.
Published: (2023)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
by: Itzhak, Itay, et al.
Published: (2026)
by: Itzhak, Itay, et al.
Published: (2026)
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
by: Fandina, Ora Nova, et al.
Published: (2024)
by: Fandina, Ora Nova, et al.
Published: (2024)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
Large Language Model Agent for User-friendly Chemical Process Simulations
by: Liang, Jingkang, et al.
Published: (2026)
by: Liang, Jingkang, et al.
Published: (2026)
Meituan Merchant Business Diagnosis via Policy-Guided Dual-Process User Simulation
by: Chen, Ziyang, et al.
Published: (2026)
by: Chen, Ziyang, et al.
Published: (2026)
Utility-Guided Agent Orchestration for Efficient LLM Tool Use
by: Liu, Boyan, et al.
Published: (2026)
by: Liu, Boyan, et al.
Published: (2026)
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
by: He, Muyu, et al.
Published: (2025)
by: He, Muyu, et al.
Published: (2025)
Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation
by: Panagopoulos, George
Published: (2026)
by: Panagopoulos, George
Published: (2026)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
by: Burdisso, Sergio, et al.
Published: (2025)
by: Burdisso, Sergio, et al.
Published: (2025)
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Similar Items
-
Breaking ReAct Agents: Foot-in-the-Door Attack Will Get You In
by: Nakash, Itay, et al.
Published: (2024) -
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
by: Kour, George, et al.
Published: (2025) -
Effective Red-Teaming of Policy-Adherent Agents
by: Nakash, Itay, et al.
Published: (2025) -
Exploring Straightforward Conversational Red-Teaming
by: Kour, George, et al.
Published: (2024) -
On the Robustness of Agentic Function Calling
by: Rabinovich, Ella, et al.
Published: (2025)