EL-VIT: Probing Vision Transformer with Interactive Visualization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhou, Hong, Zhang, Rui, Lai, Peifeng, Guo, Chaoran, Wang, Yong, Sun, Zhida, Li, Junjie
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916103355105280
author Zhou, Hong
Zhang, Rui
Lai, Peifeng
Guo, Chaoran
Wang, Yong
Sun, Zhida
Li, Junjie
author_facet Zhou, Hong
Zhang, Rui
Lai, Peifeng
Guo, Chaoran
Wang, Yong
Sun, Zhida
Li, Junjie
contents Nowadays, Vision Transformer (ViT) is widely utilized in various computer vision tasks, owing to its unique self-attention mechanism. However, the model architecture of ViT is complex and often challenging to comprehend, leading to a steep learning curve. ViT developers and users frequently encounter difficulties in interpreting its inner workings. Therefore, a visualization system is needed to assist ViT users in understanding its functionality. This paper introduces EL-VIT, an interactive visual analytics system designed to probe the Vision Transformer and facilitate a better understanding of its operations. The system consists of four layers of visualization views. The first three layers include model overview, knowledge background graph, and model detail view. These three layers elucidate the operation process of ViT from three perspectives: the overall model architecture, detailed explanation, and mathematical operations, enabling users to understand the underlying principles and the transition process between layers. The fourth interpretation view helps ViT users and experts gain a deeper understanding by calculating the cosine similarity between patches. Our two usage scenarios demonstrate the effectiveness and usability of EL-VIT in helping ViT users understand the working mechanism of ViT.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12666
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EL-VIT: Probing Vision Transformer with Interactive Visualization
Zhou, Hong
Zhang, Rui
Lai, Peifeng
Guo, Chaoran
Wang, Yong
Sun, Zhida
Li, Junjie
Artificial Intelligence
Nowadays, Vision Transformer (ViT) is widely utilized in various computer vision tasks, owing to its unique self-attention mechanism. However, the model architecture of ViT is complex and often challenging to comprehend, leading to a steep learning curve. ViT developers and users frequently encounter difficulties in interpreting its inner workings. Therefore, a visualization system is needed to assist ViT users in understanding its functionality. This paper introduces EL-VIT, an interactive visual analytics system designed to probe the Vision Transformer and facilitate a better understanding of its operations. The system consists of four layers of visualization views. The first three layers include model overview, knowledge background graph, and model detail view. These three layers elucidate the operation process of ViT from three perspectives: the overall model architecture, detailed explanation, and mathematical operations, enabling users to understand the underlying principles and the transition process between layers. The fourth interpretation view helps ViT users and experts gain a deeper understanding by calculating the cosine similarity between patches. Our two usage scenarios demonstrate the effectiveness and usability of EL-VIT in helping ViT users understand the working mechanism of ViT.
title EL-VIT: Probing Vision Transformer with Interactive Visualization
topic Artificial Intelligence
url https://arxiv.org/abs/2401.12666