Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Muyu, Kumar, Anand, Mackey, Tsach, Rajeev, Meghana, Zou, James, Rajani, Nazneen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
von: He, Muyu, et al.
Veröffentlicht: (2025)
von: He, Muyu, et al.
Veröffentlicht: (2025)
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
von: He, Muyu, et al.
Veröffentlicht: (2026)
von: He, Muyu, et al.
Veröffentlicht: (2026)
VERITAS: A Unified Approach to Reliability Evaluation
von: Ramamurthy, Rajkumar, et al.
Veröffentlicht: (2024)
von: Ramamurthy, Rajkumar, et al.
Veröffentlicht: (2024)
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
von: Rajeev, Meghana, et al.
Veröffentlicht: (2025)
von: Rajeev, Meghana, et al.
Veröffentlicht: (2025)
Stage-wise Fine-tuning for Graph-to-Text Generation
von: Wang, Qingyun, et al.
Veröffentlicht: (2021)
von: Wang, Qingyun, et al.
Veröffentlicht: (2021)
Simulating User Agents for Embodied Conversational-AI
von: Philipov, Daniel, et al.
Veröffentlicht: (2024)
von: Philipov, Daniel, et al.
Veröffentlicht: (2024)
Self-rationalization improves LLM as a fine-grained judge
von: Trivedi, Prapti, et al.
Veröffentlicht: (2024)
von: Trivedi, Prapti, et al.
Veröffentlicht: (2024)
ReviewRobot: Explainable Paper Review Generation based on Knowledge Synthesis
von: Wang, Qingyun, et al.
Veröffentlicht: (2020)
von: Wang, Qingyun, et al.
Veröffentlicht: (2020)
Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents
von: Komoravolu, Sameer, et al.
Veröffentlicht: (2025)
von: Komoravolu, Sameer, et al.
Veröffentlicht: (2025)
HumanLM: Simulating Users with State Alignment Beats Response Imitation
von: Wu, Shirley, et al.
Veröffentlicht: (2026)
von: Wu, Shirley, et al.
Veröffentlicht: (2026)
ScamAgents: How AI Agents Can Simulate Human-Level Scam Calls
von: Badhe, Sanket
Veröffentlicht: (2025)
von: Badhe, Sanket
Veröffentlicht: (2025)
Paper2Agent: Reimagining Research Papers As Interactive and Reliable AI Agents
von: Miao, Jiacheng, et al.
Veröffentlicht: (2025)
von: Miao, Jiacheng, et al.
Veröffentlicht: (2025)
OpenDeception: Learning Deception and Trust in Human-AI Interaction via Multi-Agent Simulation
von: Wu, Yichen, et al.
Veröffentlicht: (2025)
von: Wu, Yichen, et al.
Veröffentlicht: (2025)
Can LLMs Infer Conversational Agent Users' Personality Traits from Chat History?
von: Cögendez, Derya, et al.
Veröffentlicht: (2026)
von: Cögendez, Derya, et al.
Veröffentlicht: (2026)
Testing and Understanding Erroneous Planning in LLM Agents through Synthesized User Inputs
von: Ji, Zhenlan, et al.
Veröffentlicht: (2024)
von: Ji, Zhenlan, et al.
Veröffentlicht: (2024)
BiomechAgent: AI-Assisted Biomechanical Analysis Through Code-Generating Agents
von: Cotton, R. James, et al.
Veröffentlicht: (2026)
von: Cotton, R. James, et al.
Veröffentlicht: (2026)
From Control to Foresight: Simulation as a New Paradigm for Human-Agent Collaboration
von: He, Gaole, et al.
Veröffentlicht: (2026)
von: He, Gaole, et al.
Veröffentlicht: (2026)
Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation
von: Niu, Cheng, et al.
Veröffentlicht: (2024)
von: Niu, Cheng, et al.
Veröffentlicht: (2024)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
von: Chopra, Harshita, et al.
Veröffentlicht: (2026)
von: Chopra, Harshita, et al.
Veröffentlicht: (2026)
Can Large Language Model Agents Simulate Human Trust Behavior?
von: Xie, Chengxing, et al.
Veröffentlicht: (2024)
von: Xie, Chengxing, et al.
Veröffentlicht: (2024)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
What's documented in AI? Systematic Analysis of 32K AI Model Cards
von: Liang, Weixin, et al.
Veröffentlicht: (2024)
von: Liang, Weixin, et al.
Veröffentlicht: (2024)
Persona Dynamics: Unveiling the Impact of Personality Traits on Agents in Text-Based Games
von: Lim, Seungwon, et al.
Veröffentlicht: (2025)
von: Lim, Seungwon, et al.
Veröffentlicht: (2025)
AgentCourt: Simulating Court with Adversarial Evolvable Lawyer Agents
von: Chen, Guhong, et al.
Veröffentlicht: (2024)
von: Chen, Guhong, et al.
Veröffentlicht: (2024)
UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind
von: Qian, Cheng, et al.
Veröffentlicht: (2026)
von: Qian, Cheng, et al.
Veröffentlicht: (2026)
AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation
von: He, Zhitao, et al.
Veröffentlicht: (2024)
von: He, Zhitao, et al.
Veröffentlicht: (2024)
Recursive Multi-Agent Systems
von: Yang, Xiyuan, et al.
Veröffentlicht: (2026)
von: Yang, Xiyuan, et al.
Veröffentlicht: (2026)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
von: Zhang, Xichen, et al.
Veröffentlicht: (2026)
von: Zhang, Xichen, et al.
Veröffentlicht: (2026)
LM Agents for Coordinating Multi-User Information Gathering
von: Jhamtani, Harsh, et al.
Veröffentlicht: (2025)
von: Jhamtani, Harsh, et al.
Veröffentlicht: (2025)
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
von: Cohen, Myke C., et al.
Veröffentlicht: (2026)
von: Cohen, Myke C., et al.
Veröffentlicht: (2026)
UserBench: An Interactive Gym Environment for User-Centric Agents
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
Latent Collaboration in Multi-Agent Systems
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
A Proactive Multi-Agent Dialogue Framework for Assessing Social Language Disorder Traits in Autism
von: Hu, Chuanbo, et al.
Veröffentlicht: (2026)
von: Hu, Chuanbo, et al.
Veröffentlicht: (2026)
Memory in the Age of AI Agents
von: Hu, Yuyang, et al.
Veröffentlicht: (2025)
von: Hu, Yuyang, et al.
Veröffentlicht: (2025)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
von: Yu, Tao, et al.
Veröffentlicht: (2025)
von: Yu, Tao, et al.
Veröffentlicht: (2025)
Strategic Persuasion with Trait-Conditioned Multi-Agent Systems for Iterative Legal Argumentation
von: Siedler, Philipp D.
Veröffentlicht: (2026)
von: Siedler, Philipp D.
Veröffentlicht: (2026)
Agentic Test-Time Scaling for WebAgents
von: Lee, Nicholas, et al.
Veröffentlicht: (2026)
von: Lee, Nicholas, et al.
Veröffentlicht: (2026)
Goal Alignment in LLM-Based User Simulators for Conversational AI
von: Mehri, Shuhaib, et al.
Veröffentlicht: (2025)
von: Mehri, Shuhaib, et al.
Veröffentlicht: (2025)
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
von: Burdisso, Sergio, et al.
Veröffentlicht: (2025)
von: Burdisso, Sergio, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
von: He, Muyu, et al.
Veröffentlicht: (2025) -
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
von: He, Muyu, et al.
Veröffentlicht: (2026) -
VERITAS: A Unified Approach to Reliability Evaluation
von: Ramamurthy, Rajkumar, et al.
Veröffentlicht: (2024) -
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
von: Rajeev, Meghana, et al.
Veröffentlicht: (2025) -
Stage-wise Fine-tuning for Graph-to-Text Generation
von: Wang, Qingyun, et al.
Veröffentlicht: (2021)