Saved in:
| Main Authors: | Gao, Shanshan, Zhou, Liyi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.10448 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents
by: Sheng, Strick, et al.
Published: (2026)
by: Sheng, Strick, et al.
Published: (2026)
AI Agent Smart Contract Exploit Generation
by: Gervais, Arthur, et al.
Published: (2025)
by: Gervais, Arthur, et al.
Published: (2025)
MoPHES:Leveraging on-device LLMs as Agent for Mobile Psychological Health Evaluation and Support
by: Wei, Xun, et al.
Published: (2025)
by: Wei, Xun, et al.
Published: (2025)
LATTICE: Evaluating Decision Support Utility of Crypto Agents
by: Chan, Aaron, et al.
Published: (2026)
by: Chan, Aaron, et al.
Published: (2026)
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
by: Sui, Xingyu, et al.
Published: (2026)
by: Sui, Xingyu, et al.
Published: (2026)
ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support
by: Chen, Tiantian, et al.
Published: (2026)
by: Chen, Tiantian, et al.
Published: (2026)
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents
by: Unlu, Eren
Published: (2026)
by: Unlu, Eren
Published: (2026)
ChoiceMates: Supporting Unfamiliar Online Decision-Making with Multi-Agent Conversational Interactions
by: Park, Jeongeon, et al.
Published: (2023)
by: Park, Jeongeon, et al.
Published: (2023)
DeepPsy-Agent: A Stage-Aware and Deep-Thinking Emotional Support Agent System
by: Chen, Kai, et al.
Published: (2025)
by: Chen, Kai, et al.
Published: (2025)
2-Step Agent: A Framework for the Interaction of a Decision Maker with AI Decision Support
by: Nyberg, Otto, et al.
Published: (2026)
by: Nyberg, Otto, et al.
Published: (2026)
Evaluating Collaborative and Autonomous Agents in Data-Stream-Supported Coordination of Mobile Crowdsourcing
by: Bruns, Ralf, et al.
Published: (2024)
by: Bruns, Ralf, et al.
Published: (2024)
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
by: Dong, Haonan, et al.
Published: (2026)
by: Dong, Haonan, et al.
Published: (2026)
SAGE: A Service Agent Graph-guided Evaluation Benchmark
by: Shi, Ling, et al.
Published: (2026)
by: Shi, Ling, et al.
Published: (2026)
ANX: Protocol-First Design for AI Agent Interaction with a Supporting 3EX Decoupled Architecture
by: Mingze, Xu
Published: (2026)
by: Mingze, Xu
Published: (2026)
OnlineMate: An LLM-Based Multi-Agent Companion System for Cognitive Support in Online Learning
by: Gao, Xian, et al.
Published: (2025)
by: Gao, Xian, et al.
Published: (2025)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases
by: Li, Dubai, et al.
Published: (2026)
by: Li, Dubai, et al.
Published: (2026)
Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents
by: Li, Jinyang, et al.
Published: (2024)
by: Li, Jinyang, et al.
Published: (2024)
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
by: Atinafu, Yonas, et al.
Published: (2026)
by: Atinafu, Yonas, et al.
Published: (2026)
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning
by: Wu, Yuyang, et al.
Published: (2026)
by: Wu, Yuyang, et al.
Published: (2026)
Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First
by: Liu, Shu, et al.
Published: (2025)
by: Liu, Shu, et al.
Published: (2025)
ArguAgent: AI-Supported Real-Time Grouping for Productive Argumentation in STEM Classrooms
by: Kleiman, Jennifer, et al.
Published: (2026)
by: Kleiman, Jennifer, et al.
Published: (2026)
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
by: Chen, Rongqian, et al.
Published: (2026)
by: Chen, Rongqian, et al.
Published: (2026)
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
by: Priyanshu, Aman, et al.
Published: (2026)
by: Priyanshu, Aman, et al.
Published: (2026)
Do Agents Know What They Can't Do? Evaluating Feasibility Awareness in Tool-Using Agents
by: Cheng, Liang, et al.
Published: (2026)
by: Cheng, Liang, et al.
Published: (2026)
SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation
by: Chen, Jingxuan, et al.
Published: (2024)
by: Chen, Jingxuan, et al.
Published: (2024)
Toward User Comprehension Supports for LLM Agent Skill Specifications
by: Wen, Zikai Alex
Published: (2026)
by: Wen, Zikai Alex
Published: (2026)
macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
by: Yang, Pei, et al.
Published: (2025)
by: Yang, Pei, et al.
Published: (2025)
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
by: Levy, Ido, et al.
Published: (2024)
by: Levy, Ido, et al.
Published: (2024)
Can AI Master Econometrics? Evidence from Econometrics AI Agent on Expert-Level Tasks
by: Chen, Qiang, et al.
Published: (2025)
by: Chen, Qiang, et al.
Published: (2025)
MemoCoder: Automated Function Synthesis using LLM-Supported Agents
by: Jia, Yiping, et al.
Published: (2025)
by: Jia, Yiping, et al.
Published: (2025)
Can Agents Fix Agent Issues?
by: Rahardja, Alfin Wijaya, et al.
Published: (2025)
by: Rahardja, Alfin Wijaya, et al.
Published: (2025)
A Novel Task-Driven Method with Evolvable Interactive Agents Using Event Trees for Enhanced Emergency Decision Support
by: Xiao, Xingyu, et al.
Published: (2024)
by: Xiao, Xingyu, et al.
Published: (2024)
Can LLMs Support Medical Knowledge Imputation? An Evaluation-Based Perspective
by: Yao, Xinyu, et al.
Published: (2025)
by: Yao, Xinyu, et al.
Published: (2025)
Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning
by: Li, Sijia, et al.
Published: (2026)
by: Li, Sijia, et al.
Published: (2026)
Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-based Customer Support
by: Zhao, Cen Mia, et al.
Published: (2025)
by: Zhao, Cen Mia, et al.
Published: (2025)
Evaluation and Benchmarking of LLM Agents: A Survey
by: Mohammadi, Mahmoud, et al.
Published: (2025)
by: Mohammadi, Mahmoud, et al.
Published: (2025)
Evaluating Cognitive Age Alignment in Interactive AI Agents
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions
by: Song, Maojia, et al.
Published: (2025)
by: Song, Maojia, et al.
Published: (2025)
Similar Items
-
When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents
by: Sheng, Strick, et al.
Published: (2026) -
AI Agent Smart Contract Exploit Generation
by: Gervais, Arthur, et al.
Published: (2025) -
MoPHES:Leveraging on-device LLMs as Agent for Mobile Psychological Health Evaluation and Support
by: Wei, Xun, et al.
Published: (2025) -
LATTICE: Evaluating Decision Support Utility of Crypto Agents
by: Chan, Aaron, et al.
Published: (2026) -
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
by: Sui, Xingyu, et al.
Published: (2026)