Transformer^-1: Input-Adaptive Computation for Resource-Constrained Deployment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: AI, Lumen, School, Tengzhou No. 1 Middle, Ji, Shihao, Song, Zihui, Zhong, Fucheng, Jia, Jisen, Wu, Zhaobo, Cao, Zheyi, Tianhao, Xu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909468591128576
author AI, Lumen
School, Tengzhou No. 1 Middle
Ji, Shihao
Song, Zihui
Zhong, Fucheng
Jia, Jisen
Wu, Zhaobo
Cao, Zheyi
Tianhao, Xu
author_facet AI, Lumen
School, Tengzhou No. 1 Middle
Ji, Shihao
Song, Zihui
Zhong, Fucheng
Jia, Jisen
Wu, Zhaobo
Cao, Zheyi
Tianhao, Xu
contents Addressing the resource waste caused by fixed computation paradigms in deep learning models under dynamic scenarios, this paper proposes a Transformer$^{-1}$ architecture based on the principle of deep adaptivity. This architecture achieves dynamic matching between input features and computational resources by establishing a joint optimization model for complexity and computation. Our core contributions include: (1) designing a two-layer control mechanism, composed of a complexity predictor and a reinforcement learning policy network, enabling end-to-end optimization of computation paths; (2) deriving a lower bound theory for dynamic computation, proving the system's theoretical reach to optimal efficiency; and (3) proposing a layer folding technique and a CUDA Graph pre-compilation scheme, overcoming the engineering bottlenecks of dynamic architectures. In the ImageNet-1K benchmark test, our method reduces FLOPs by 42.7\% and peak memory usage by 34.1\% compared to the standard Transformer, while maintaining comparable accuracy ($\pm$0.3\%). Furthermore, we conducted practical deployment on the Jetson AGX Xavier platform, verifying the effectiveness and practical value of this method in resource-constrained environments. To further validate the generality of the method, we also conducted experiments on several natural language processing tasks and achieved significant improvements in resource efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2501_16394
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transformer^-1: Input-Adaptive Computation for Resource-Constrained Deployment
AI, Lumen
School, Tengzhou No. 1 Middle
Ji, Shihao
Song, Zihui
Zhong, Fucheng
Jia, Jisen
Wu, Zhaobo
Cao, Zheyi
Tianhao, Xu
Machine Learning
Addressing the resource waste caused by fixed computation paradigms in deep learning models under dynamic scenarios, this paper proposes a Transformer$^{-1}$ architecture based on the principle of deep adaptivity. This architecture achieves dynamic matching between input features and computational resources by establishing a joint optimization model for complexity and computation. Our core contributions include: (1) designing a two-layer control mechanism, composed of a complexity predictor and a reinforcement learning policy network, enabling end-to-end optimization of computation paths; (2) deriving a lower bound theory for dynamic computation, proving the system's theoretical reach to optimal efficiency; and (3) proposing a layer folding technique and a CUDA Graph pre-compilation scheme, overcoming the engineering bottlenecks of dynamic architectures. In the ImageNet-1K benchmark test, our method reduces FLOPs by 42.7\% and peak memory usage by 34.1\% compared to the standard Transformer, while maintaining comparable accuracy ($\pm$0.3\%). Furthermore, we conducted practical deployment on the Jetson AGX Xavier platform, verifying the effectiveness and practical value of this method in resource-constrained environments. To further validate the generality of the method, we also conducted experiments on several natural language processing tasks and achieved significant improvements in resource efficiency.
title Transformer^-1: Input-Adaptive Computation for Resource-Constrained Deployment
topic Machine Learning
url https://arxiv.org/abs/2501.16394