TRUST: Test-time Resource Utilization for Superior Trustworthiness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Harikumar, Haripriya, Rana, Santu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908396391759872
author Harikumar, Haripriya
Rana, Santu
author_facet Harikumar, Haripriya
Rana, Santu
contents Standard uncertainty estimation techniques, such as dropout, often struggle to clearly distinguish reliable predictions from unreliable ones. We attribute this limitation to noisy classifier weights, which, while not impairing overall class-level predictions, render finer-level statistics less informative. To address this, we propose a novel test-time optimization method that accounts for the impact of such noise to produce more reliable confidence estimates. This score defines a monotonic subset-selection function, where population accuracy consistently increases as samples with lower scores are removed, and it demonstrates superior performance in standard risk-based metrics such as AUSE and AURC. Additionally, our method effectively identifies discrepancies between training and test distributions, reliably differentiates in-distribution from out-of-distribution samples, and elucidates key differences between CNN and ViT classifiers across various vision datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06048
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TRUST: Test-time Resource Utilization for Superior Trustworthiness
Harikumar, Haripriya
Rana, Santu
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Standard uncertainty estimation techniques, such as dropout, often struggle to clearly distinguish reliable predictions from unreliable ones. We attribute this limitation to noisy classifier weights, which, while not impairing overall class-level predictions, render finer-level statistics less informative. To address this, we propose a novel test-time optimization method that accounts for the impact of such noise to produce more reliable confidence estimates. This score defines a monotonic subset-selection function, where population accuracy consistently increases as samples with lower scores are removed, and it demonstrates superior performance in standard risk-based metrics such as AUSE and AURC. Additionally, our method effectively identifies discrepancies between training and test distributions, reliably differentiates in-distribution from out-of-distribution samples, and elucidates key differences between CNN and ViT classifiers across various vision datasets.
title TRUST: Test-time Resource Utilization for Superior Trustworthiness
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.06048