Representation-Level Counterfactual Calibration for Debiased Zero-Shot Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Peng, Pei, Xie, MingKun, Hao, Hang, Jin, Tong, Huang, ShengJun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917054259396608
author Peng, Pei
Xie, MingKun
Hao, Hang
Jin, Tong
Huang, ShengJun
author_facet Peng, Pei
Xie, MingKun
Hao, Hang
Jin, Tong
Huang, ShengJun
contents Object-context shortcuts remain a persistent challenge in vision-language models, undermining zero-shot reliability when test-time scenes differ from familiar training co-occurrences. We recast this issue as a causal inference problem and ask: Would the prediction remain if the object appeared in a different environment? To answer this at inference time, we estimate object and background expectations within CLIP's representation space, and synthesize counterfactual embeddings by recombining object features with diverse alternative contexts sampled from external datasets, batch neighbors, or text-derived descriptions. By estimating the Total Direct Effect and simulating intervention, we further subtract background-only activation, preserving beneficial object-context interactions while mitigating hallucinated scores. Without retraining or prompt design, our method substantially improves both worst-group and average accuracy on context-sensitive benchmarks, establishing a new zero-shot state of the art. Beyond performance, our framework provides a lightweight representation-level counterfactual approach, offering a practical causal avenue for debiased and reliable multimodal reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26466
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Representation-Level Counterfactual Calibration for Debiased Zero-Shot Recognition
Peng, Pei
Xie, MingKun
Hao, Hang
Jin, Tong
Huang, ShengJun
Computer Vision and Pattern Recognition
Machine Learning
Object-context shortcuts remain a persistent challenge in vision-language models, undermining zero-shot reliability when test-time scenes differ from familiar training co-occurrences. We recast this issue as a causal inference problem and ask: Would the prediction remain if the object appeared in a different environment? To answer this at inference time, we estimate object and background expectations within CLIP's representation space, and synthesize counterfactual embeddings by recombining object features with diverse alternative contexts sampled from external datasets, batch neighbors, or text-derived descriptions. By estimating the Total Direct Effect and simulating intervention, we further subtract background-only activation, preserving beneficial object-context interactions while mitigating hallucinated scores. Without retraining or prompt design, our method substantially improves both worst-group and average accuracy on context-sensitive benchmarks, establishing a new zero-shot state of the art. Beyond performance, our framework provides a lightweight representation-level counterfactual approach, offering a practical causal avenue for debiased and reliable multimodal reasoning.
title Representation-Level Counterfactual Calibration for Debiased Zero-Shot Recognition
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2510.26466