REOrdering Patches Improves Vision Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kutscher, Declan, Chan, David M., Bai, Yutong, Darrell, Trevor, Gupta, Ritwik
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911227263844352
author Kutscher, Declan
Chan, David M.
Bai, Yutong
Darrell, Trevor
Gupta, Ritwik
author_facet Kutscher, Declan
Chan, David M.
Bai, Yutong
Darrell, Trevor
Gupta, Ritwik
contents Sequence models such as transformers require inputs to be represented as one-dimensional sequences. In vision, this typically involves flattening images using a fixed row-major (raster-scan) order. While full self-attention is permutation-equivariant, modern long-sequence transformers increasingly rely on architectural approximations that break this invariance and introduce sensitivity to patch ordering. We show that patch order significantly affects model performance in such settings, with simple alternatives like column-major or Hilbert curves yielding notable accuracy shifts. Motivated by this, we propose REOrder, a two-stage framework for discovering task-optimal patch orderings. First, we derive an information-theoretic prior by evaluating the compressibility of various patch sequences. Then, we learn a policy over permutations by optimizing a Plackett-Luce policy using REINFORCE. This approach enables efficient learning in a combinatorial permutation space. REOrder improves top-1 accuracy over row-major ordering on ImageNet-1K by up to 3.01% and Functional Map of the World by 13.35%.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23751
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle REOrdering Patches Improves Vision Models
Kutscher, Declan
Chan, David M.
Bai, Yutong
Darrell, Trevor
Gupta, Ritwik
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Sequence models such as transformers require inputs to be represented as one-dimensional sequences. In vision, this typically involves flattening images using a fixed row-major (raster-scan) order. While full self-attention is permutation-equivariant, modern long-sequence transformers increasingly rely on architectural approximations that break this invariance and introduce sensitivity to patch ordering. We show that patch order significantly affects model performance in such settings, with simple alternatives like column-major or Hilbert curves yielding notable accuracy shifts. Motivated by this, we propose REOrder, a two-stage framework for discovering task-optimal patch orderings. First, we derive an information-theoretic prior by evaluating the compressibility of various patch sequences. Then, we learn a policy over permutations by optimizing a Plackett-Luce policy using REINFORCE. This approach enables efficient learning in a combinatorial permutation space. REOrder improves top-1 accuracy over row-major ordering on ImageNet-1K by up to 3.01% and Functional Map of the World by 13.35%.
title REOrdering Patches Improves Vision Models
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.23751