Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rahman, Md Ashiqur, Yang, Chiao-An, Cheng, Michael N., Hao, Lim Jun, Jiang, Jeremiah, Lim, Teck-Yian, Yeh, Raymond A.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918127430795264
author Rahman, Md Ashiqur
Yang, Chiao-An
Cheng, Michael N.
Hao, Lim Jun
Jiang, Jeremiah
Lim, Teck-Yian
Yeh, Raymond A.
author_facet Rahman, Md Ashiqur
Yang, Chiao-An
Cheng, Michael N.
Hao, Lim Jun
Jiang, Jeremiah
Lim, Teck-Yian
Yeh, Raymond A.
contents Scale variation is a fundamental challenge in computer vision. Objects of the same class can have different sizes, and their perceived size is further affected by the distance from the camera. These variations are local to the objects, i.e., different object sizes may change differently within the same image. To effectively handle scale variations, we present a deep equilibrium canonicalizer (DEC) to improve the local scale equivariance of a model. DEC can be easily incorporated into existing network architectures and can be adapted to a pre-trained model. Notably, we show that on the competitive ImageNet benchmark, DEC improves both model performance and local scale consistency across four popular pre-trained deep-nets, e.g., ViT, DeiT, Swin, and BEiT. Our code is available at https://github.com/ashiq24/local-scale-equivariance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_14187
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer
Rahman, Md Ashiqur
Yang, Chiao-An
Cheng, Michael N.
Hao, Lim Jun
Jiang, Jeremiah
Lim, Teck-Yian
Yeh, Raymond A.
Computer Vision and Pattern Recognition
Graphics
Machine Learning
Scale variation is a fundamental challenge in computer vision. Objects of the same class can have different sizes, and their perceived size is further affected by the distance from the camera. These variations are local to the objects, i.e., different object sizes may change differently within the same image. To effectively handle scale variations, we present a deep equilibrium canonicalizer (DEC) to improve the local scale equivariance of a model. DEC can be easily incorporated into existing network architectures and can be adapted to a pre-trained model. Notably, we show that on the competitive ImageNet benchmark, DEC improves both model performance and local scale consistency across four popular pre-trained deep-nets, e.g., ViT, DeiT, Swin, and BEiT. Our code is available at https://github.com/ashiq24/local-scale-equivariance.
title Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2508.14187