O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sharifdeen, Ashshak, Munir, Muhammad Akhtar, Baliah, Sanoojan, Khan, Salman, Khan, Muhammad Haris
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929761563967488
author Sharifdeen, Ashshak
Munir, Muhammad Akhtar
Baliah, Sanoojan
Khan, Salman
Khan, Muhammad Haris
author_facet Sharifdeen, Ashshak
Munir, Muhammad Akhtar
Baliah, Sanoojan
Khan, Salman
Khan, Muhammad Haris
contents Test-time prompt tuning for vision-language models (VLMs) is getting attention because of their ability to learn with unlabeled data without fine-tuning. Although test-time prompt tuning methods for VLMs can boost accuracy, the resulting models tend to demonstrate poor calibration, which casts doubts on the reliability and trustworthiness of these models. Notably, more attention needs to be devoted to calibrating the test-time prompt tuning in vision-language models. To this end, we propose a new approach, called O-TPT that introduces orthogonality constraints on the textual features corresponding to the learnable prompts for calibrating test-time prompt tuning in VLMs. Towards introducing orthogonality constraints, we make the following contributions. First, we uncover new insights behind the suboptimal calibration performance of existing methods relying on textual feature dispersion. Second, we show that imposing a simple orthogonalization of textual features is a more effective approach towards obtaining textual dispersion. We conduct extensive experiments on various datasets with different backbones and baselines. The results indicate that our method consistently outperforms the prior state of the art in significantly reducing the overall average calibration error. Also, our method surpasses the zero-shot calibration performance on fine-grained classification tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12096
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models
Sharifdeen, Ashshak
Munir, Muhammad Akhtar
Baliah, Sanoojan
Khan, Salman
Khan, Muhammad Haris
Computer Vision and Pattern Recognition
Test-time prompt tuning for vision-language models (VLMs) is getting attention because of their ability to learn with unlabeled data without fine-tuning. Although test-time prompt tuning methods for VLMs can boost accuracy, the resulting models tend to demonstrate poor calibration, which casts doubts on the reliability and trustworthiness of these models. Notably, more attention needs to be devoted to calibrating the test-time prompt tuning in vision-language models. To this end, we propose a new approach, called O-TPT that introduces orthogonality constraints on the textual features corresponding to the learnable prompts for calibrating test-time prompt tuning in VLMs. Towards introducing orthogonality constraints, we make the following contributions. First, we uncover new insights behind the suboptimal calibration performance of existing methods relying on textual feature dispersion. Second, we show that imposing a simple orthogonalization of textual features is a more effective approach towards obtaining textual dispersion. We conduct extensive experiments on various datasets with different backbones and baselines. The results indicate that our method consistently outperforms the prior state of the art in significantly reducing the overall average calibration error. Also, our method surpasses the zero-shot calibration performance on fine-grained classification tasks.
title O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.12096