The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xuan, Weihao, Zeng, Qingcheng, Qi, Heli, Xiao, Yunze, Wang, Junjue, Yokoya, Naoto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
von: Xuan, Weihao, et al.
Veröffentlicht: (2025)
von: Xuan, Weihao, et al.
Veröffentlicht: (2025)
Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery
von: Tsujimoto, Mai, et al.
Veröffentlicht: (2025)
von: Tsujimoto, Mai, et al.
Veröffentlicht: (2025)
DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding
von: Xuan, Weihao, et al.
Veröffentlicht: (2025)
von: Xuan, Weihao, et al.
Veröffentlicht: (2025)
Thinking Out Loud: Do Reasoning Models Know When They're Right?
von: Zeng, Qingcheng, et al.
Veröffentlicht: (2025)
von: Zeng, Qingcheng, et al.
Veröffentlicht: (2025)
Direction-aware 3D Large Multimodal Models
von: Liu, Quan, et al.
Veröffentlicht: (2026)
von: Liu, Quan, et al.
Veröffentlicht: (2026)
Segment Anything with Multiple Modalities
von: Xiao, Aoran, et al.
Veröffentlicht: (2024)
von: Xiao, Aoran, et al.
Veröffentlicht: (2024)
Experience-Driven Multi-Agent Systems Are Training-free Context-aware Earth Observers
von: Dai, Pengyu, et al.
Veröffentlicht: (2026)
von: Dai, Pengyu, et al.
Veröffentlicht: (2026)
Foundation Models for Remote Sensing and Earth Observation: A Survey
von: Xiao, Aoran, et al.
Veröffentlicht: (2024)
von: Xiao, Aoran, et al.
Veröffentlicht: (2024)
Sentipolis: Emotion-Aware Agents for Social Simulations
von: Fu, Chiyuan, et al.
Veröffentlicht: (2026)
von: Fu, Chiyuan, et al.
Veröffentlicht: (2026)
Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations
von: Wang, Junjue, et al.
Veröffentlicht: (2026)
von: Wang, Junjue, et al.
Veröffentlicht: (2026)
The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models
von: Yu, Kefan, et al.
Veröffentlicht: (2025)
von: Yu, Kefan, et al.
Veröffentlicht: (2025)
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
von: Wang, Junjue, et al.
Veröffentlicht: (2025)
von: Wang, Junjue, et al.
Veröffentlicht: (2025)
OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation
von: Zhao, Sijie, et al.
Veröffentlicht: (2026)
von: Zhao, Sijie, et al.
Veröffentlicht: (2026)
Code-Switching Information Retrieval: Benchmarks, Analysis, and the Limits of Current Retrievers
von: Zeng, Qingcheng, et al.
Veröffentlicht: (2026)
von: Zeng, Qingcheng, et al.
Veröffentlicht: (2026)
Taming Object Hallucinations with Verified Atomic Confidence Estimation
von: Liu, Jiarui, et al.
Veröffentlicht: (2025)
von: Liu, Jiarui, et al.
Veröffentlicht: (2025)
The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models
von: Xiao, Yunze, et al.
Veröffentlicht: (2026)
von: Xiao, Yunze, et al.
Veröffentlicht: (2026)
Extreme Miscalibration and the Illusion of Adversarial Robustness
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Towards Valid Student Simulation with Large Language Models
von: Yuan, Zhihao, et al.
Veröffentlicht: (2026)
von: Yuan, Zhihao, et al.
Veröffentlicht: (2026)
Say Something Else: Rethinking Contextual Privacy as Information Sufficiency
von: Xiao, Yunze, et al.
Veröffentlicht: (2026)
von: Xiao, Yunze, et al.
Veröffentlicht: (2026)
Hallucination, Monofacts, and Miscalibration: An Empirical Investigation
von: Miao, Miranda Muqing, et al.
Veröffentlicht: (2025)
von: Miao, Miranda Muqing, et al.
Veröffentlicht: (2025)
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
von: Xuan, Weihao, et al.
Veröffentlicht: (2025)
von: Xuan, Weihao, et al.
Veröffentlicht: (2025)
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search
von: Wu, Fang, et al.
Veröffentlicht: (2025)
von: Wu, Fang, et al.
Veröffentlicht: (2025)
Verified Critical Step Optimization for LLM Agents
von: Li, Mukai, et al.
Veröffentlicht: (2026)
von: Li, Mukai, et al.
Veröffentlicht: (2026)
SynRS3D: A Synthetic Dataset for Global 3D Semantic Understanding from Monocular Remote Sensing Imagery
von: Song, Jian, et al.
Veröffentlicht: (2024)
von: Song, Jian, et al.
Veröffentlicht: (2024)
NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
von: Liu, Jiayu, et al.
Veröffentlicht: (2026)
von: Liu, Jiayu, et al.
Veröffentlicht: (2026)
The Tool Illusion: Rethinking Tool Use in Web Agents
von: Lou, Renze, et al.
Veröffentlicht: (2026)
von: Lou, Renze, et al.
Veröffentlicht: (2026)
Incongruent Positivity: When Miscalibrated Positivity Undermines Online Supportive Conversations
von: Almajed, Leen, et al.
Veröffentlicht: (2025)
von: Almajed, Leen, et al.
Veröffentlicht: (2025)
Large Language Models are Miscalibrated In-Context Learners
von: Li, Chengzu, et al.
Veröffentlicht: (2023)
von: Li, Chengzu, et al.
Veröffentlicht: (2023)
Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs
von: Jin, Mingyu, et al.
Veröffentlicht: (2026)
von: Jin, Mingyu, et al.
Veröffentlicht: (2026)
Saving the legacy of Hero Ibash: Evaluating Four Language Models for Aminoacian
von: Xiao, Yunze, et al.
Veröffentlicht: (2024)
von: Xiao, Yunze, et al.
Veröffentlicht: (2024)
When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors
von: Yang, Chenghao, et al.
Veröffentlicht: (2026)
von: Yang, Chenghao, et al.
Veröffentlicht: (2026)
Learning to Use Tools via Cooperative and Interactive Agents
von: Shi, Zhengliang, et al.
Veröffentlicht: (2024)
von: Shi, Zhengliang, et al.
Veröffentlicht: (2024)
Leveraging Human Production-Interpretation Asymmetries to Test LLM Cognitive Plausibility
von: Lam, Suet-Ying, et al.
Veröffentlicht: (2025)
von: Lam, Suet-Ying, et al.
Veröffentlicht: (2025)
ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
von: Xiao, Yunzhong, et al.
Veröffentlicht: (2025)
von: Xiao, Yunzhong, et al.
Veröffentlicht: (2025)
Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents
von: Qi, Xuan
Veröffentlicht: (2026)
von: Qi, Xuan
Veröffentlicht: (2026)
Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models
von: Cox, Kyle, et al.
Veröffentlicht: (2025)
von: Cox, Kyle, et al.
Veröffentlicht: (2025)
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
von: Xu, Minrui, et al.
Veröffentlicht: (2026)
von: Xu, Minrui, et al.
Veröffentlicht: (2026)
Mitigating the Impact of False Negatives in Dense Retrieval with Contrastive Confidence Regularization
von: Wang, Shiqi, et al.
Veröffentlicht: (2023)
von: Wang, Shiqi, et al.
Veröffentlicht: (2023)
MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
von: Li, Binxu, et al.
Veröffentlicht: (2024)
von: Li, Binxu, et al.
Veröffentlicht: (2024)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
von: Xu, Haoyuan, et al.
Veröffentlicht: (2026)
von: Xu, Haoyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
von: Xuan, Weihao, et al.
Veröffentlicht: (2025) -
Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery
von: Tsujimoto, Mai, et al.
Veröffentlicht: (2025) -
DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding
von: Xuan, Weihao, et al.
Veröffentlicht: (2025) -
Thinking Out Loud: Do Reasoning Models Know When They're Right?
von: Zeng, Qingcheng, et al.
Veröffentlicht: (2025) -
Direction-aware 3D Large Multimodal Models
von: Liu, Quan, et al.
Veröffentlicht: (2026)