VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
Fuente:
arXiv
Saved in:
| Main Authors: | Sung, Yoo Yeon, Kim, Hannah, Zhang, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
by: Lee, Jungjae, et al.
Published: (2025)
by: Lee, Jungjae, et al.
Published: (2025)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
by: Gor, Maharshi, et al.
Published: (2026)
by: Gor, Maharshi, et al.
Published: (2026)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
by: Yoon, Sion, et al.
Published: (2024)
by: Yoon, Sion, et al.
Published: (2024)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
by: Shao, Yijia, et al.
Published: (2024)
by: Shao, Yijia, et al.
Published: (2024)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
by: Zhou, Kaitlyn, et al.
Published: (2024)
by: Zhou, Kaitlyn, et al.
Published: (2024)
Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures
by: Shen, Hua, et al.
Published: (2025)
by: Shen, Hua, et al.
Published: (2025)
Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
by: Kim, Taesoo, et al.
Published: (2025)
by: Kim, Taesoo, et al.
Published: (2025)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
by: Kim, Jane Paik
Published: (2026)
by: Kim, Jane Paik
Published: (2026)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
by: Huq, Faria, et al.
Published: (2025)
by: Huq, Faria, et al.
Published: (2025)
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
by: Liu, Tianjian, et al.
Published: (2025)
by: Liu, Tianjian, et al.
Published: (2025)
Human-Centered AI in Multidisciplinary Medical Discussions: Evaluating the Feasibility of a Chat-Based Approach to Case Assessment
by: Sawano, Shinnosuke, et al.
Published: (2025)
by: Sawano, Shinnosuke, et al.
Published: (2025)
Automated Interpretability and Feature Discovery in Language Models with Agents
by: Marin-Llobet, Arnau, et al.
Published: (2026)
by: Marin-Llobet, Arnau, et al.
Published: (2026)
LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
by: Sun, Chongyan, et al.
Published: (2024)
by: Sun, Chongyan, et al.
Published: (2024)
Game Development as Human-LLM Interaction
by: Hong, Jiale, et al.
Published: (2024)
by: Hong, Jiale, et al.
Published: (2024)
Humans overrely on overconfident language models, across languages
by: Rathi, Neil, et al.
Published: (2025)
by: Rathi, Neil, et al.
Published: (2025)
Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
by: Yao, Bingsheng, et al.
Published: (2025)
by: Yao, Bingsheng, et al.
Published: (2025)
ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
by: Yang, Bufang, et al.
Published: (2025)
by: Yang, Bufang, et al.
Published: (2025)
ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wild
by: Yang, Bufang, et al.
Published: (2025)
by: Yang, Bufang, et al.
Published: (2025)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
by: Zhu, Andrew, et al.
Published: (2025)
by: Zhu, Andrew, et al.
Published: (2025)
Human-Centered LLM-Agent User Interface: A Position Paper
by: Chin, Daniel, et al.
Published: (2024)
by: Chin, Daniel, et al.
Published: (2024)
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
by: Zhao, Runcong, et al.
Published: (2025)
by: Zhao, Runcong, et al.
Published: (2025)
Do Proactive Agents Really Need an LLM to Decide When to Wake and What to Anchor?
by: Liu, Xiaoze, et al.
Published: (2026)
by: Liu, Xiaoze, et al.
Published: (2026)
Creativity in LLM-based Multi-Agent Systems: A Survey
by: Lin, Yi-Cheng, et al.
Published: (2025)
by: Lin, Yi-Cheng, et al.
Published: (2025)
BADGE: BADminton report Generation and Evaluation with LLM
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
Evaluating LLMs as Human Surrogates in Controlled Experiments
by: Hoq, Adnan, et al.
Published: (2026)
by: Hoq, Adnan, et al.
Published: (2026)
Can Unconfident LLM Annotations Be Used for Confident Conclusions?
by: Gligorić, Kristina, et al.
Published: (2024)
by: Gligorić, Kristina, et al.
Published: (2024)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
by: Li, Weiyue, et al.
Published: (2026)
by: Li, Weiyue, et al.
Published: (2026)
PAIR-SAFE: A Paired-Agent Approach for Runtime Auditing and Refining AI-Mediated Mental Health Support
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
by: Palta, Shramay, et al.
Published: (2025)
by: Palta, Shramay, et al.
Published: (2025)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
by: Wang, Zora Zhiruo, et al.
Published: (2025)
by: Wang, Zora Zhiruo, et al.
Published: (2025)
Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents
by: Schiepanski, Thassilo M., et al.
Published: (2025)
by: Schiepanski, Thassilo M., et al.
Published: (2025)
Can Large Language Model Agents Simulate Human Trust Behavior?
by: Xie, Chengxing, et al.
Published: (2024)
by: Xie, Chengxing, et al.
Published: (2024)
EditScribe: Non-Visual Image Editing with Natural Language Verification Loops
by: Chang, Ruei-Che, et al.
Published: (2024)
by: Chang, Ruei-Che, et al.
Published: (2024)
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
by: Yao, Xintong
Published: (2026)
by: Yao, Xintong
Published: (2026)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
From Control to Foresight: Simulation as a New Paradigm for Human-Agent Collaboration
by: He, Gaole, et al.
Published: (2026)
by: He, Gaole, et al.
Published: (2026)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
Large Language Models Can Solve Real-World Planning Rigorously with Formal Verification Tools
by: Hao, Yilun, et al.
Published: (2024)
by: Hao, Yilun, et al.
Published: (2024)
Similar Items
-
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification
by: Lee, Jungjae, et al.
Published: (2025) -
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
by: Gor, Maharshi, et al.
Published: (2026) -
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
by: Yoon, Sion, et al.
Published: (2024) -
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
by: Shao, Yijia, et al.
Published: (2024) -
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)