ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kaixin, Yang, Ruiqing, Zhang, Yuan, You, Shan, Huang, Tao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915621728419840
author Zhang, Kaixin
Yang, Ruiqing
Zhang, Yuan
You, Shan
Huang, Tao
author_facet Zhang, Kaixin
Yang, Ruiqing
Zhang, Yuan
You, Shan
Huang, Tao
contents Visual Autoregressive (VAR) models enable efficient image generation via next-scale prediction but face escalating computational costs as sequence length grows. Existing static pruning methods degrade performance by permanently removing weights or tokens, disrupting pretrained dependencies. To address this, we propose ActVAR, a dynamic activation framework that introduces dual sparsity across model weights and token sequences to enhance efficiency without sacrificing capacity. ActVAR decomposes feedforward networks (FFNs) into lightweight expert sub-networks and employs a learnable router to dynamically select token-specific expert subsets based on content. Simultaneously, a gated token selector identifies high-update-potential tokens for computation while reconstructing unselected tokens to preserve global context and sequence alignment. Training employs a two-stage knowledge distillation strategy, where the original VAR model supervises the learning of routing and gating policies to align with pretrained knowledge. Experiments on the ImageNet $256\times 256$ benchmark demonstrate that ActVAR achieves up to $21.2\%$ FLOPs reduction with minimal performance degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12893
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
Zhang, Kaixin
Yang, Ruiqing
Zhang, Yuan
You, Shan
Huang, Tao
Computer Vision and Pattern Recognition
Visual Autoregressive (VAR) models enable efficient image generation via next-scale prediction but face escalating computational costs as sequence length grows. Existing static pruning methods degrade performance by permanently removing weights or tokens, disrupting pretrained dependencies. To address this, we propose ActVAR, a dynamic activation framework that introduces dual sparsity across model weights and token sequences to enhance efficiency without sacrificing capacity. ActVAR decomposes feedforward networks (FFNs) into lightweight expert sub-networks and employs a learnable router to dynamically select token-specific expert subsets based on content. Simultaneously, a gated token selector identifies high-update-potential tokens for computation while reconstructing unselected tokens to preserve global context and sequence alignment. Training employs a two-stage knowledge distillation strategy, where the original VAR model supervises the learning of routing and gating policies to align with pretrained knowledge. Experiments on the ImageNet $256\times 256$ benchmark demonstrate that ActVAR achieves up to $21.2\%$ FLOPs reduction with minimal performance degradation.
title ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.12893