MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness
Fuente:
arXiv
Saved in:
| Main Authors: | Hathidara, Ashutosh, Yu, Julien, Senthil, Vaishali, Schreiber, Sebastian, Ankisettipalli, Anil Babu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
by: Senthil, Vaishali, et al.
Published: (2026)
by: Senthil, Vaishali, et al.
Published: (2026)
Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky
by: Hathidara, Ashutosh, et al.
Published: (2025)
by: Hathidara, Ashutosh, et al.
Published: (2025)
MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
by: Guo, Shengyu, et al.
Published: (2026)
by: Guo, Shengyu, et al.
Published: (2026)
Implementing An Artificial Quantum Perceptron
by: Hathidara, Ashutosh, et al.
Published: (2024)
by: Hathidara, Ashutosh, et al.
Published: (2024)
Mining Tweets to Predict Future Bitcoin Price
by: Hathidara, Ashutosh, et al.
Published: (2024)
by: Hathidara, Ashutosh, et al.
Published: (2024)
UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
by: Wang, Chao, et al.
Published: (2024)
by: Wang, Chao, et al.
Published: (2024)
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
by: Seshadri, Preethi, et al.
Published: (2026)
by: Seshadri, Preethi, et al.
Published: (2026)
AgentBench: Evaluating LLMs as Agents
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
UserBench: An Interactive Gym Environment for User-Centric Agents
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery
by: Kleczek, Darek, et al.
Published: (2026)
by: Kleczek, Darek, et al.
Published: (2026)
FedMABench: Benchmarking Mobile Agents on Decentralized Heterogeneous User Data
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
by: Weng, Muyan, et al.
Published: (2026)
by: Weng, Muyan, et al.
Published: (2026)
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
by: Anurin, Andrey, et al.
Published: (2024)
by: Anurin, Andrey, et al.
Published: (2024)
Bench to the Future: A Pastcasting Benchmark for Forecasting Agents
by: FutureSearch, et al.
Published: (2025)
by: FutureSearch, et al.
Published: (2025)
BenchMARL: Benchmarking Multi-Agent Reinforcement Learning
by: Bettini, Matteo, et al.
Published: (2023)
by: Bettini, Matteo, et al.
Published: (2023)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
ExtractBench: A Benchmark and Evaluation Methodology for Complex Structured Extraction
by: Ferguson, Nick, et al.
Published: (2026)
by: Ferguson, Nick, et al.
Published: (2026)
MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
by: Huang, Qian, et al.
Published: (2023)
by: Huang, Qian, et al.
Published: (2023)
OptiProxy-NAS: Optimization Proxy based End-to-End Neural Architecture Search
by: Lyu, Bo, et al.
Published: (2025)
by: Lyu, Bo, et al.
Published: (2025)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
by: Shen, Yuanzhe, et al.
Published: (2026)
by: Shen, Yuanzhe, et al.
Published: (2026)
DataSciBench: An LLM Agent Benchmark for Data Science
by: Zhang, Dan, et al.
Published: (2025)
by: Zhang, Dan, et al.
Published: (2025)
Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
by: Guo, Jian-Ting, et al.
Published: (2025)
by: Guo, Jian-Ting, et al.
Published: (2025)
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
by: Bogavelli, Tara, et al.
Published: (2026)
by: Bogavelli, Tara, et al.
Published: (2026)
PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization
by: Zhu, Jiachen, et al.
Published: (2026)
by: Zhu, Jiachen, et al.
Published: (2026)
AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models
by: Zbeeb, Mohammad, et al.
Published: (2025)
by: Zbeeb, Mohammad, et al.
Published: (2025)
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
by: Wang, Lu, et al.
Published: (2025)
by: Wang, Lu, et al.
Published: (2025)
AgentEval: Generative Agents as Reliable Proxies for Human Evaluation of AI-Generated Content
by: Vu, Thanh, et al.
Published: (2025)
by: Vu, Thanh, et al.
Published: (2025)
EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective
by: Wang, Yuyao, et al.
Published: (2026)
by: Wang, Yuyao, et al.
Published: (2026)
CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics
by: Wang, Weida, et al.
Published: (2025)
by: Wang, Weida, et al.
Published: (2025)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
by: Moore, Robert J., et al.
Published: (2026)
by: Moore, Robert J., et al.
Published: (2026)
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents
by: De Brouwer, Edward, et al.
Published: (2026)
by: De Brouwer, Edward, et al.
Published: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
Evaluation and Benchmarking of LLM Agents: A Survey
by: Mohammadi, Mahmoud, et al.
Published: (2025)
by: Mohammadi, Mahmoud, et al.
Published: (2025)
The Era of Real-World Human Interaction: RL from User Conversations
by: Jin, Chuanyang, et al.
Published: (2025)
by: Jin, Chuanyang, et al.
Published: (2025)
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
by: Jiang, Yixing, et al.
Published: (2025)
by: Jiang, Yixing, et al.
Published: (2025)
PepSpecBench: A Unified Evaluation Benchmark for Peptide Tandem Mass Spectrometry Prediction
by: Yang, Zhiwen, et al.
Published: (2026)
by: Yang, Zhiwen, et al.
Published: (2026)
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
by: Ma, Yinghao, et al.
Published: (2025)
by: Ma, Yinghao, et al.
Published: (2025)
IGL-Bench: Establishing the Comprehensive Benchmark for Imbalanced Graph Learning
by: Qin, Jiawen, et al.
Published: (2024)
by: Qin, Jiawen, et al.
Published: (2024)
Similar Items
-
CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
by: Senthil, Vaishali, et al.
Published: (2026) -
Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky
by: Hathidara, Ashutosh, et al.
Published: (2025) -
MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
by: Guo, Shengyu, et al.
Published: (2026) -
Implementing An Artificial Quantum Perceptron
by: Hathidara, Ashutosh, et al.
Published: (2024) -
Mining Tweets to Predict Future Bitcoin Price
by: Hathidara, Ashutosh, et al.
Published: (2024)