Benchmarking LLM Tool-Use in the Wild
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Peijie, Liu, Wei, Yang, Yifan, Li, Jinjian, Zhang, Zelong, Feng, Xiao, Zhang, Feng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
IDRBench: Interactive Deep Research Benchmark
von: Feng, Yingchaojie, et al.
Veröffentlicht: (2026)
von: Feng, Yingchaojie, et al.
Veröffentlicht: (2026)
ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wild
von: Yang, Bufang, et al.
Veröffentlicht: (2025)
von: Yang, Bufang, et al.
Veröffentlicht: (2025)
LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces
von: Kirgis, Peter, et al.
Veröffentlicht: (2026)
von: Kirgis, Peter, et al.
Veröffentlicht: (2026)
OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents
von: Sun, Qiang, et al.
Veröffentlicht: (2024)
von: Sun, Qiang, et al.
Veröffentlicht: (2024)
SSKG Hub: An Expert-Guided Platform for LLM-Empowered Sustainability Standards Knowledge Graphs
von: He, Chaoyue, et al.
Veröffentlicht: (2026)
von: He, Chaoyue, et al.
Veröffentlicht: (2026)
Large Language Models Can Solve Real-World Planning Rigorously with Formal Verification Tools
von: Hao, Yilun, et al.
Veröffentlicht: (2024)
von: Hao, Yilun, et al.
Veröffentlicht: (2024)
The Persuasion Paradox: When LLM Explanations Fail to Improve Human-AI Team Performance
von: Cohen, Ruth, et al.
Veröffentlicht: (2026)
von: Cohen, Ruth, et al.
Veröffentlicht: (2026)
LLM4AD: Large Language Models for Autonomous Driving -- Concept, Review, Benchmark, Experiments, and Future Trends
von: Cui, Can, et al.
Veröffentlicht: (2024)
von: Cui, Can, et al.
Veröffentlicht: (2024)
ShareChat: A Dataset of Chatbot Conversations in the Wild
von: Yan, Yueru, et al.
Veröffentlicht: (2025)
von: Yan, Yueru, et al.
Veröffentlicht: (2025)
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
von: Liu, Shuyu, et al.
Veröffentlicht: (2025)
von: Liu, Shuyu, et al.
Veröffentlicht: (2025)
REALM: A Dataset of Real-World LLM Use Cases
von: Cheng, Jingwen, et al.
Veröffentlicht: (2025)
von: Cheng, Jingwen, et al.
Veröffentlicht: (2025)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
von: Li, Weiyue, et al.
Veröffentlicht: (2026)
von: Li, Weiyue, et al.
Veröffentlicht: (2026)
MAIC-UI: Making Interactive Courseware with Generative UI
von: Tu, Shangqing, et al.
Veröffentlicht: (2026)
von: Tu, Shangqing, et al.
Veröffentlicht: (2026)
SensorPersona: An LLM-Empowered System for Continual Persona Extraction from Longitudinal Mobile Sensor Streams
von: Yang, Bufang, et al.
Veröffentlicht: (2026)
von: Yang, Bufang, et al.
Veröffentlicht: (2026)
LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models
von: Liu, Minqian, et al.
Veröffentlicht: (2025)
von: Liu, Minqian, et al.
Veröffentlicht: (2025)
Generative Echo Chamber? Effects of LLM-Powered Search Systems on Diverse Information Seeking
von: Sharma, Nikhil, et al.
Veröffentlicht: (2024)
von: Sharma, Nikhil, et al.
Veröffentlicht: (2024)
Do Proactive Agents Really Need an LLM to Decide When to Wake and What to Anchor?
von: Liu, Xiaoze, et al.
Veröffentlicht: (2026)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2026)
Creating General User Models from Computer Use
von: Shaikh, Omar, et al.
Veröffentlicht: (2025)
von: Shaikh, Omar, et al.
Veröffentlicht: (2025)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
von: Sung, Yoo Yeon, et al.
Veröffentlicht: (2025)
von: Sung, Yoo Yeon, et al.
Veröffentlicht: (2025)
Augmenting Research Ideation with Data: An Empirical Investigation in Social Science
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
BADGE: BADminton report Generation and Evaluation with LLM
von: Chiang, Shang-Hsuan, et al.
Veröffentlicht: (2024)
von: Chiang, Shang-Hsuan, et al.
Veröffentlicht: (2024)
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits
von: Jiang, Hang, et al.
Veröffentlicht: (2023)
von: Jiang, Hang, et al.
Veröffentlicht: (2023)
User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios
von: Wu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Creativity in LLM-based Multi-Agent Systems: A Survey
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
Understand User Opinions of Large Language Models via LLM-Powered In-the-Moment User Experience Interviews
von: Liu, Mengqiao, et al.
Veröffentlicht: (2025)
von: Liu, Mengqiao, et al.
Veröffentlicht: (2025)
Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
von: Zhang, Xingjian, et al.
Veröffentlicht: (2025)
von: Zhang, Xingjian, et al.
Veröffentlicht: (2025)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
Awaking the Slides: A Tuning-free and Knowledge-regulated AI Tutoring System via Language Model Coordination
von: Zhang-Li, Daniel, et al.
Veröffentlicht: (2024)
von: Zhang-Li, Daniel, et al.
Veröffentlicht: (2024)
Large Language Model-Brained GUI Agents: A Survey
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
PKG API: A Tool for Personal Knowledge Graph Management
von: Bernard, Nolwenn, et al.
Veröffentlicht: (2024)
von: Bernard, Nolwenn, et al.
Veröffentlicht: (2024)
Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
von: Lam, Michelle S., et al.
Veröffentlicht: (2024)
von: Lam, Michelle S., et al.
Veröffentlicht: (2024)
ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
von: Yang, Bufang, et al.
Veröffentlicht: (2025)
von: Yang, Bufang, et al.
Veröffentlicht: (2025)
Understanding How Paper Writers Use AI-Generated Captions in Figure Caption Writing
von: Yin, Ho, et al.
Veröffentlicht: (2025)
von: Yin, Ho, et al.
Veröffentlicht: (2025)
Game Plot Design with an LLM-powered Assistant: An Empirical Study with Game Designers
von: Alavi, Seyed Hossein, et al.
Veröffentlicht: (2024)
von: Alavi, Seyed Hossein, et al.
Veröffentlicht: (2024)
AI as Teammate or Tool? A Review of Human-AI Interaction in Decision Support
von: Samu, Most. Sharmin Sultana, et al.
Veröffentlicht: (2026)
von: Samu, Most. Sharmin Sultana, et al.
Veröffentlicht: (2026)
Evaluating Large Language Models in Analysing Classroom Dialogue
von: Long, Yun, et al.
Veröffentlicht: (2024)
von: Long, Yun, et al.
Veröffentlicht: (2024)
MindGuard: Towards Accessible and Sitgma-free Mental Health First Aid via Edge LLM
von: Ji, Sijie, et al.
Veröffentlicht: (2024)
von: Ji, Sijie, et al.
Veröffentlicht: (2024)
UFO: A UI-Focused Agent for Windows OS Interaction
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
Psychometric Comparability of LLM-Based Digital Twins
von: Zhang, Yufei, et al.
Veröffentlicht: (2025)
von: Zhang, Yufei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
IDRBench: Interactive Deep Research Benchmark
von: Feng, Yingchaojie, et al.
Veröffentlicht: (2026) -
ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wild
von: Yang, Bufang, et al.
Veröffentlicht: (2025) -
LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces
von: Kirgis, Peter, et al.
Veröffentlicht: (2026) -
OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents
von: Sun, Qiang, et al.
Veröffentlicht: (2024) -
SSKG Hub: An Expert-Guided Platform for LLM-Empowered Sustainability Standards Knowledge Graphs
von: He, Chaoyue, et al.
Veröffentlicht: (2026)