Rotary Position Embedding for Vision Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heo, Byeongho, Park, Song, Han, Dongyoon, Yun, Sangdoo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911956491829248
author Heo, Byeongho
Park, Song
Han, Dongyoon
Yun, Sangdoo
author_facet Heo, Byeongho
Park, Song
Han, Dongyoon
Yun, Sangdoo
contents Rotary Position Embedding (RoPE) performs remarkably on language models, especially for length extrapolation of Transformers. However, the impacts of RoPE on computer vision domains have been underexplored, even though RoPE appears capable of enhancing Vision Transformer (ViT) performance in a way similar to the language domain. This study provides a comprehensive analysis of RoPE when applied to ViTs, utilizing practical implementations of RoPE for 2D vision data. The analysis reveals that RoPE demonstrates impressive extrapolation performance, i.e., maintaining precision while increasing image resolution at inference. It eventually leads to performance improvement for ImageNet-1k, COCO detection, and ADE-20k segmentation. We believe this study provides thorough guidelines to apply RoPE into ViT, promising improved backbone performance with minimal extra computational overhead. Our code and pre-trained models are available at https://github.com/naver-ai/rope-vit
format Preprint
id arxiv_https___arxiv_org_abs_2403_13298
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rotary Position Embedding for Vision Transformer
Heo, Byeongho
Park, Song
Han, Dongyoon
Yun, Sangdoo
Computer Vision and Pattern Recognition
Machine Learning
Rotary Position Embedding (RoPE) performs remarkably on language models, especially for length extrapolation of Transformers. However, the impacts of RoPE on computer vision domains have been underexplored, even though RoPE appears capable of enhancing Vision Transformer (ViT) performance in a way similar to the language domain. This study provides a comprehensive analysis of RoPE when applied to ViTs, utilizing practical implementations of RoPE for 2D vision data. The analysis reveals that RoPE demonstrates impressive extrapolation performance, i.e., maintaining precision while increasing image resolution at inference. It eventually leads to performance improvement for ImageNet-1k, COCO detection, and ADE-20k segmentation. We believe this study provides thorough guidelines to apply RoPE into ViT, promising improved backbone performance with minimal extra computational overhead. Our code and pre-trained models are available at https://github.com/naver-ai/rope-vit
title Rotary Position Embedding for Vision Transformer
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2403.13298