ZeroSiam: An Efficient Asymmetry for Test-Time Entropy Optimization without Collapse

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Guohao, Niu, Shuaicheng, Chen, Deyu, Yang, Jiahao, Zhang, Zitian, Tan, Mingkui, Wu, Pengcheng, Shen, Zhiqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911693368459264
author Chen, Guohao
Niu, Shuaicheng
Chen, Deyu
Yang, Jiahao
Zhang, Zitian
Tan, Mingkui
Wu, Pengcheng
Shen, Zhiqi
author_facet Chen, Guohao
Niu, Shuaicheng
Chen, Deyu
Yang, Jiahao
Zhang, Zitian
Tan, Mingkui
Wu, Pengcheng
Shen, Zhiqi
contents Test-time entropy minimization helps adapt a model to novel environments and incentivize its reasoning capability, unleashing the model's potential during inference by allowing it to evolve and improve in real-time using its own predictions, achieving promising performance. However, pure entropy minimization can favor non-generalizable shortcuts, such as inflating the logit norm and driving all predictions to a dominant class to reduce entropy, risking collapsed solutions (e.g., constant one-hot outputs) that trivially minimize the objective without meaningful learning. In this paper, we reveal asymmetry as a key mechanism for collapse prevention and introduce ZeroSiam--an efficient asymmetric Siamese architecture tailored for test-time entropy minimization. ZeroSiam prevents collapse through asymmetric divergence alignment, efficiently achieved by a learnable predictor and a stop-gradient operator before the classifier. We provide empirical and theoretical evidence that ZeroSiam not only prevents collapse, but also regularizes biased learning signals, enhancing performance even when no collapse occurs. Despite its simplicity, extensive results show that ZeroSiam performs more stably over prior methods using negligible overhead, demonstrating efficacy on both vision adaptation and large language model reasoning tasks across challenging test scenarios and diverse models, including particularly collapse-prone tiny models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23183
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ZeroSiam: An Efficient Asymmetry for Test-Time Entropy Optimization without Collapse
Chen, Guohao
Niu, Shuaicheng
Chen, Deyu
Yang, Jiahao
Zhang, Zitian
Tan, Mingkui
Wu, Pengcheng
Shen, Zhiqi
Machine Learning
Networking and Internet Architecture
Test-time entropy minimization helps adapt a model to novel environments and incentivize its reasoning capability, unleashing the model's potential during inference by allowing it to evolve and improve in real-time using its own predictions, achieving promising performance. However, pure entropy minimization can favor non-generalizable shortcuts, such as inflating the logit norm and driving all predictions to a dominant class to reduce entropy, risking collapsed solutions (e.g., constant one-hot outputs) that trivially minimize the objective without meaningful learning. In this paper, we reveal asymmetry as a key mechanism for collapse prevention and introduce ZeroSiam--an efficient asymmetric Siamese architecture tailored for test-time entropy minimization. ZeroSiam prevents collapse through asymmetric divergence alignment, efficiently achieved by a learnable predictor and a stop-gradient operator before the classifier. We provide empirical and theoretical evidence that ZeroSiam not only prevents collapse, but also regularizes biased learning signals, enhancing performance even when no collapse occurs. Despite its simplicity, extensive results show that ZeroSiam performs more stably over prior methods using negligible overhead, demonstrating efficacy on both vision adaptation and large language model reasoning tasks across challenging test scenarios and diverse models, including particularly collapse-prone tiny models.
title ZeroSiam: An Efficient Asymmetry for Test-Time Entropy Optimization without Collapse
topic Machine Learning
Networking and Internet Architecture
url https://arxiv.org/abs/2509.23183