MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Subramani, Nishant, Eisner, Jason, Svegliato, Justin, Van Durme, Benjamin, Su, Yu, Thomson, Sam
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910921556754432
author Subramani, Nishant
Eisner, Jason
Svegliato, Justin
Van Durme, Benjamin
Su, Yu
Thomson, Sam
author_facet Subramani, Nishant
Eisner, Jason
Svegliato, Justin
Van Durme, Benjamin
Su, Yu
Thomson, Sam
contents Tool-using agents that act in the world need to be both useful and safe. Well-calibrated model confidences can be used to weigh the risk versus reward of potential actions, but prior work shows that many models are poorly calibrated. Inspired by interpretability literature exploring the internals of models, we propose a novel class of model-internal confidence estimators (MICE) to better assess confidence when calling tools. MICE first decodes from each intermediate layer of the language model using logitLens and then computes similarity scores between each layer's generation and the final output. These features are fed into a learned probabilistic classifier to assess confidence in the decoded output. On the simulated trial and error (STE) tool-calling dataset using Llama3 models, we find that MICE beats or matches the baselines on smoothed expected calibration error. Using MICE confidences to determine whether to call a tool significantly improves over strong baselines on a new metric, expected tool-calling utility. Further experiments show that MICE is sample-efficient, can generalize zero-shot to unseen APIs, and results in higher tool-calling utility in scenarios with varying risk levels. Our code is open source, available at https://github.com/microsoft/mice_for_cats.
format Preprint
id arxiv_https___arxiv_org_abs_2504_20168
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
Subramani, Nishant
Eisner, Jason
Svegliato, Justin
Van Durme, Benjamin
Su, Yu
Thomson, Sam
Computation and Language
Artificial Intelligence
Machine Learning
Tool-using agents that act in the world need to be both useful and safe. Well-calibrated model confidences can be used to weigh the risk versus reward of potential actions, but prior work shows that many models are poorly calibrated. Inspired by interpretability literature exploring the internals of models, we propose a novel class of model-internal confidence estimators (MICE) to better assess confidence when calling tools. MICE first decodes from each intermediate layer of the language model using logitLens and then computes similarity scores between each layer's generation and the final output. These features are fed into a learned probabilistic classifier to assess confidence in the decoded output. On the simulated trial and error (STE) tool-calling dataset using Llama3 models, we find that MICE beats or matches the baselines on smoothed expected calibration error. Using MICE confidences to determine whether to call a tool significantly improves over strong baselines on a new metric, expected tool-calling utility. Further experiments show that MICE is sample-efficient, can generalize zero-shot to unseen APIs, and results in higher tool-calling utility in scenarios with varying risk levels. Our code is open source, available at https://github.com/microsoft/mice_for_cats.
title MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.20168