Memory Efficient Matting with Adaptive Token Routing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lin, Yiheng, Hu, Yihan, Zhang, Chenyi, Liu, Ting, Qu, Xiaochao, Liu, Luoqi, Zhao, Yao, Wei, Yunchao
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910748882501632
author Lin, Yiheng
Hu, Yihan
Zhang, Chenyi
Liu, Ting
Qu, Xiaochao
Liu, Luoqi
Zhao, Yao
Wei, Yunchao
author_facet Lin, Yiheng
Hu, Yihan
Zhang, Chenyi
Liu, Ting
Qu, Xiaochao
Liu, Luoqi
Zhao, Yao
Wei, Yunchao
contents Transformer-based models have recently achieved outstanding performance in image matting. However, their application to high-resolution images remains challenging due to the quadratic complexity of global self-attention. To address this issue, we propose MEMatte, a \textbf{m}emory-\textbf{e}fficient \textbf{m}atting framework for processing high-resolution images. MEMatte incorporates a router before each global attention block, directing informative tokens to the global attention while routing other tokens to a Lightweight Token Refinement Module (LTRM). Specifically, the router employs a local-global strategy to predict the routing probability of each token, and the LTRM utilizes efficient modules to simulate global attention. Additionally, we introduce a Batch-constrained Adaptive Token Routing (BATR) mechanism, which allows each router to dynamically route tokens based on image content and the stages of attention block in the network. Furthermore, we construct an ultra high-resolution image matting dataset, UHR-395, comprising 35,500 training images and 1,000 test images, with an average resolution of $4872\times6017$. This dataset is created by compositing 395 different alpha mattes across 11 categories onto various backgrounds, all with high-quality manual annotation. Extensive experiments demonstrate that MEMatte outperforms existing methods on both high-resolution and real-world datasets, significantly reducing memory usage by approximately 88% and latency by 50% on the Composition-1K benchmark. Our code is available at https://github.com/linyiheng123/MEMatte.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10702
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Memory Efficient Matting with Adaptive Token Routing
Lin, Yiheng
Hu, Yihan
Zhang, Chenyi
Liu, Ting
Qu, Xiaochao
Liu, Luoqi
Zhao, Yao
Wei, Yunchao
Computer Vision and Pattern Recognition
Transformer-based models have recently achieved outstanding performance in image matting. However, their application to high-resolution images remains challenging due to the quadratic complexity of global self-attention. To address this issue, we propose MEMatte, a \textbf{m}emory-\textbf{e}fficient \textbf{m}atting framework for processing high-resolution images. MEMatte incorporates a router before each global attention block, directing informative tokens to the global attention while routing other tokens to a Lightweight Token Refinement Module (LTRM). Specifically, the router employs a local-global strategy to predict the routing probability of each token, and the LTRM utilizes efficient modules to simulate global attention. Additionally, we introduce a Batch-constrained Adaptive Token Routing (BATR) mechanism, which allows each router to dynamically route tokens based on image content and the stages of attention block in the network. Furthermore, we construct an ultra high-resolution image matting dataset, UHR-395, comprising 35,500 training images and 1,000 test images, with an average resolution of $4872\times6017$. This dataset is created by compositing 395 different alpha mattes across 11 categories onto various backgrounds, all with high-quality manual annotation. Extensive experiments demonstrate that MEMatte outperforms existing methods on both high-resolution and real-world datasets, significantly reducing memory usage by approximately 88% and latency by 50% on the Composition-1K benchmark. Our code is available at https://github.com/linyiheng123/MEMatte.
title Memory Efficient Matting with Adaptive Token Routing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.10702