RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Ming, Tan, Juntao, Murthy, Rithesh, Qiu, Jielin, Yang, Liangwei, Zhao, Wenting, Savarese, Silvio, Heinecke, Shelby, Wang, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
by: Murthy, Rithesh, et al.
Published: (2025)
by: Murthy, Rithesh, et al.
Published: (2025)
VoiceAgentRAG: Solving the RAG Latency Bottleneck in Real-Time Voice Agents Using Dual-Agent Architectures
by: Qiu, Jielin, et al.
Published: (2026)
by: Qiu, Jielin, et al.
Published: (2026)
Building Enterprise Realtime Voice Agents from Scratch: A Technical Tutorial
by: Qiu, Jielin, et al.
Published: (2026)
by: Qiu, Jielin, et al.
Published: (2026)
Enterprise Sales Copilot: Enabling Real-Time AI Support with Automatic Information Retrieval in Live Sales Calls
by: Qiu, Jielin, et al.
Published: (2026)
by: Qiu, Jielin, et al.
Published: (2026)
Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training
by: Qiu, Jielin, et al.
Published: (2026)
by: Qiu, Jielin, et al.
Published: (2026)
Prompt Optimization Via Diffusion Language Models
by: Wang, Shiyu, et al.
Published: (2026)
by: Wang, Shiyu, et al.
Published: (2026)
AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech
by: Qiu, Jielin, et al.
Published: (2026)
by: Qiu, Jielin, et al.
Published: (2026)
Position: Vector Prompt Interfaces Should Be Exposed to Enable Customization of Large Language Models
by: Yang, Liangwei, et al.
Published: (2026)
by: Yang, Liangwei, et al.
Published: (2026)
PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data
by: Tan, Juntao, et al.
Published: (2025)
by: Tan, Juntao, et al.
Published: (2025)
LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering
by: Qiu, Jielin, et al.
Published: (2025)
by: Qiu, Jielin, et al.
Published: (2025)
UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
by: Qiu, Jielin, et al.
Published: (2025)
by: Qiu, Jielin, et al.
Published: (2025)
Entropy-Based Block Pruning for Efficient Large Language Models
by: Yang, Liangwei, et al.
Published: (2025)
by: Yang, Liangwei, et al.
Published: (2025)
PRACT: Optimizing Principled Reasoning and Acting of LLM Agent
by: Liu, Zhiwei, et al.
Published: (2024)
by: Liu, Zhiwei, et al.
Published: (2024)
UserBench: An Interactive Gym Environment for User-Centric Agents
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
Personalized Multi-task Training for Recommender System
by: Yang, Liangwei, et al.
Published: (2024)
by: Yang, Liangwei, et al.
Published: (2024)
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
by: Murthy, Rithesh, et al.
Published: (2024)
by: Murthy, Rithesh, et al.
Published: (2024)
AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning
by: Zhang, Jianguo, et al.
Published: (2024)
by: Zhang, Jianguo, et al.
Published: (2024)
GeoGNN: Quantifying and Mitigating Semantic Drift in Text-Attributed Graphs
by: Yang, Liangwei, et al.
Published: (2025)
by: Yang, Liangwei, et al.
Published: (2025)
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
by: Kokane, Shirley, et al.
Published: (2024)
by: Kokane, Shirley, et al.
Published: (2024)
APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
by: Liu, Zuxin, et al.
Published: (2024)
by: Liu, Zuxin, et al.
Published: (2024)
AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System
by: Liu, Zhiwei, et al.
Published: (2024)
by: Liu, Zhiwei, et al.
Published: (2024)
MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models
by: Liu, Zhiwei, et al.
Published: (2025)
by: Liu, Zhiwei, et al.
Published: (2025)
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
by: Zhou, Xuhui, et al.
Published: (2026)
by: Zhou, Xuhui, et al.
Published: (2026)
REX: Rapid Exploration and eXploitation for AI Agents
by: Murthy, Rithesh, et al.
Published: (2023)
by: Murthy, Rithesh, et al.
Published: (2023)
Retroformer: Retrospective Large Language Agents with Policy Gradient Optimization
by: Yao, Weiran, et al.
Published: (2023)
by: Yao, Weiran, et al.
Published: (2023)
Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents
by: Zhang, Kexun, et al.
Published: (2024)
by: Zhang, Kexun, et al.
Published: (2024)
xLAM: A Family of Large Action Models to Empower AI Agent Systems
by: Zhang, Jianguo, et al.
Published: (2024)
by: Zhang, Jianguo, et al.
Published: (2024)
Towards More Robust and Accurate Sequential Recommendation with Cascade-guided Adversarial Training
by: Tan, Juntao, et al.
Published: (2023)
by: Tan, Juntao, et al.
Published: (2023)
ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning
by: Yue, Murong, et al.
Published: (2025)
by: Yue, Murong, et al.
Published: (2025)
xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
Causal Layering via Conditional Entropy
by: Feigenbaum, Itai, et al.
Published: (2024)
by: Feigenbaum, Itai, et al.
Published: (2024)
Editing Arbitrary Propositions in LLMs without Subject Labels
by: Feigenbaum, Itai, et al.
Published: (2024)
by: Feigenbaum, Itai, et al.
Published: (2024)
LATTE: Learning to Think with Vision Specialists
by: Ma, Zixian, et al.
Published: (2024)
by: Ma, Zixian, et al.
Published: (2024)
Can Context Bridge the Reality Gap? Sim-to-Real Transfer of Context-Aware Policies
by: Iannotta, Marco, et al.
Published: (2025)
by: Iannotta, Marco, et al.
Published: (2025)
Test-Time Adaptation for LLM Agents via Environment Interaction
by: Chen, Arthur, et al.
Published: (2025)
by: Chen, Arthur, et al.
Published: (2025)
APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
by: Prabhakar, Akshara, et al.
Published: (2025)
by: Prabhakar, Akshara, et al.
Published: (2025)
Real-is-Sim: Bridging the Sim-to-Real Gap with a Dynamic Digital Twin
by: Abou-Chakra, Jad, et al.
Published: (2025)
by: Abou-Chakra, Jad, et al.
Published: (2025)
PolySim: Bridging the Sim-to-Real Gap for Humanoid Control via Multi-Simulator Dynamics Randomization
by: Lei, Zixing, et al.
Published: (2025)
by: Lei, Zixing, et al.
Published: (2025)
Bridging the Sim-to-Real Gap with Bayesian Inference
by: Rothfuss, Jonas, et al.
Published: (2024)
by: Rothfuss, Jonas, et al.
Published: (2024)
Similar Items
-
Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
by: Murthy, Rithesh, et al.
Published: (2025) -
VoiceAgentRAG: Solving the RAG Latency Bottleneck in Real-Time Voice Agents Using Dual-Agent Architectures
by: Qiu, Jielin, et al.
Published: (2026) -
Building Enterprise Realtime Voice Agents from Scratch: A Technical Tutorial
by: Qiu, Jielin, et al.
Published: (2026) -
Enterprise Sales Copilot: Enabling Real-Time AI Support with Automatic Information Retrieval in Live Sales Calls
by: Qiu, Jielin, et al.
Published: (2026) -
Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training
by: Qiu, Jielin, et al.
Published: (2026)