AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jajal, Purvish, Eliopoulos, Nick John, Chou, Benjamin Shiue-Hal, Thiruvathukal, George K., Lu, Yung-Hsiang, Davis, James C.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909917806329856
author Jajal, Purvish
Eliopoulos, Nick John
Chou, Benjamin Shiue-Hal
Thiruvathukal, George K.
Lu, Yung-Hsiang
Davis, James C.
author_facet Jajal, Purvish
Eliopoulos, Nick John
Chou, Benjamin Shiue-Hal
Thiruvathukal, George K.
Lu, Yung-Hsiang
Davis, James C.
contents Modern transformer architectures achieve remarkable performance across tasks and domains but remain rigid in how they allocate computation at inference time. Real-world deployment often requires models to adapt to diverse hardware and latency constraints, yet most approaches to dynamic computation focus on a single axis -- such as reducing the number of tokens. We present a novel capability: AdaPerceiver, the first transformer architecture with unified adaptivity across depth, width, and tokens within a single model. We propose an architecture that supports adaptivity along these axes. We couple this with an efficient joint training regime that ensures the model maintains performance across its various configurations. We evaluate AdaPerceiver on image classification, semantic segmentation, and depth estimation tasks. On image classification, AdaPerceiver expands the accuracy-throughput Pareto front. It achieves 85.4% accuracy while yielding 36% higher throughput than FlexiViT-L. On dense prediction, AdaPerceiver matches ViT-H/14 while having $\sim$26x fewer encoder FLOPs (floating-point operations) on semantic segmentation and depth estimation. Finally, we show how AdaPerceiver equipped with a policy can maintain ImageNet1K accuracy ($\pm0.1$ percentage points) while reducing FLOPs by $24-33$%.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18105
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
Jajal, Purvish
Eliopoulos, Nick John
Chou, Benjamin Shiue-Hal
Thiruvathukal, George K.
Lu, Yung-Hsiang
Davis, James C.
Computer Vision and Pattern Recognition
Machine Learning
Modern transformer architectures achieve remarkable performance across tasks and domains but remain rigid in how they allocate computation at inference time. Real-world deployment often requires models to adapt to diverse hardware and latency constraints, yet most approaches to dynamic computation focus on a single axis -- such as reducing the number of tokens. We present a novel capability: AdaPerceiver, the first transformer architecture with unified adaptivity across depth, width, and tokens within a single model. We propose an architecture that supports adaptivity along these axes. We couple this with an efficient joint training regime that ensures the model maintains performance across its various configurations. We evaluate AdaPerceiver on image classification, semantic segmentation, and depth estimation tasks. On image classification, AdaPerceiver expands the accuracy-throughput Pareto front. It achieves 85.4% accuracy while yielding 36% higher throughput than FlexiViT-L. On dense prediction, AdaPerceiver matches ViT-H/14 while having $\sim$26x fewer encoder FLOPs (floating-point operations) on semantic segmentation and depth estimation. Finally, we show how AdaPerceiver equipped with a policy can maintain ImageNet1K accuracy ($\pm0.1$ percentage points) while reducing FLOPs by $24-33$%.
title AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2511.18105