ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline
Fuente:
arXiv
Salvato in:
| Autori principali: | Hernandez, Juan Manuel, Fernandez-Espinosa, Mariana, Parra, Denis, Gomez-Zara, Diego |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
di: Wróbel, Adam, et al.
Pubblicazione: (2026)
di: Wróbel, Adam, et al.
Pubblicazione: (2026)
Augmenting Teamwork through AI Agents as Spatial Collaborators
di: Fernandez-Espinosa, Mariana, et al.
Pubblicazione: (2025)
di: Fernandez-Espinosa, Mariana, et al.
Pubblicazione: (2025)
Interactivity x Explainability: Toward Understanding How Interactivity Can Improve Computer Vision Explanations
di: Panigrahi, Indu, et al.
Pubblicazione: (2025)
di: Panigrahi, Indu, et al.
Pubblicazione: (2025)
Behavioral Engagement in VR-Based Sign Language Learning: Visual Attention as a Predictor of Performance and Temporal Dynamics
di: Traini, Davide, et al.
Pubblicazione: (2026)
di: Traini, Davide, et al.
Pubblicazione: (2026)
QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
di: Wang, Yuxiao, et al.
Pubblicazione: (2025)
di: Wang, Yuxiao, et al.
Pubblicazione: (2025)
ChildCI Framework: Analysis of Motor and Cognitive Development in Children-Computer Interaction for Age Detection
di: Ruiz-Garcia, Juan Carlos, et al.
Pubblicazione: (2022)
di: Ruiz-Garcia, Juan Carlos, et al.
Pubblicazione: (2022)
ClickAIXR: On-Device Multimodal Vision-Language Interaction with Real-World Objects in Extended Reality
di: Khan, Dawar, et al.
Pubblicazione: (2026)
di: Khan, Dawar, et al.
Pubblicazione: (2026)
Interact with me: Joint Egocentric Forecasting of Intent to Interact, Attitude and Social Actions
di: Bian, Tongfei, et al.
Pubblicazione: (2024)
di: Bian, Tongfei, et al.
Pubblicazione: (2024)
Vision Language Models as Values Detectors
di: Abbo, Giulio Antonio, et al.
Pubblicazione: (2025)
di: Abbo, Giulio Antonio, et al.
Pubblicazione: (2025)
VisionCAD: An Integration-Free Radiology Copilot Framework
di: Li, Jiaming, et al.
Pubblicazione: (2025)
di: Li, Jiaming, et al.
Pubblicazione: (2025)
VFA: Vision Frequency Analysis of Foundation Models and Human
di: Darvishi-Bayazi, Mohammad-Javad, et al.
Pubblicazione: (2024)
di: Darvishi-Bayazi, Mohammad-Javad, et al.
Pubblicazione: (2024)
Computer Vision for Objects used in Group Work: Challenges and Opportunities
di: Jung, Changsoo, et al.
Pubblicazione: (2025)
di: Jung, Changsoo, et al.
Pubblicazione: (2025)
Real-Time Cellist Postural Evaluation With On-Device Computer Vision
di: Wang, Paolo, et al.
Pubblicazione: (2026)
di: Wang, Paolo, et al.
Pubblicazione: (2026)
Machine Vision-Based Surgical Lighting System:Design and Implementation
di: Gharghabi, Amir, et al.
Pubblicazione: (2025)
di: Gharghabi, Amir, et al.
Pubblicazione: (2025)
Weak-Annotation of HAR Datasets using Vision Foundation Models
di: Bock, Marius, et al.
Pubblicazione: (2024)
di: Bock, Marius, et al.
Pubblicazione: (2024)
Personalized Interpretability -- Interactive Alignment of Prototypical Parts Networks
di: Michalski, Tomasz, et al.
Pubblicazione: (2025)
di: Michalski, Tomasz, et al.
Pubblicazione: (2025)
UIPro: Unleashing Superior Interaction Capability For GUI Agents
di: Li, Hongxin, et al.
Pubblicazione: (2025)
di: Li, Hongxin, et al.
Pubblicazione: (2025)
Panda or not Panda? Understanding Adversarial Attacks with Interactive Visualization
di: You, Yuzhe, et al.
Pubblicazione: (2023)
di: You, Yuzhe, et al.
Pubblicazione: (2023)
ViEEG: Hierarchical Visual Neural Representation for EEG Brain Decoding
di: Liu, Minxu, et al.
Pubblicazione: (2025)
di: Liu, Minxu, et al.
Pubblicazione: (2025)
ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning
di: Li, Po-han, et al.
Pubblicazione: (2026)
di: Li, Po-han, et al.
Pubblicazione: (2026)
SymbolSight: Minimizing Inter-Symbol Interference for Reading with Prosthetic Vision
di: Lesner, Jasmine, et al.
Pubblicazione: (2026)
di: Lesner, Jasmine, et al.
Pubblicazione: (2026)
iTrace: Click-Based Gaze Visualization on the Apple Vision Pro
di: Mehmedova, Esra, et al.
Pubblicazione: (2025)
di: Mehmedova, Esra, et al.
Pubblicazione: (2025)
Collection Space Navigator: An Interactive Visualization Interface for Multidimensional Datasets
di: Ohm, Tillmann, et al.
Pubblicazione: (2023)
di: Ohm, Tillmann, et al.
Pubblicazione: (2023)
A Dataset for Crucial Object Recognition in Blind and Low-Vision Individuals' Navigation
di: Islam, Md Touhidul, et al.
Pubblicazione: (2024)
di: Islam, Md Touhidul, et al.
Pubblicazione: (2024)
An Egocentric Vision-Language Model based Portable Real-time Smart Assistant
di: Huang, Yifei, et al.
Pubblicazione: (2025)
di: Huang, Yifei, et al.
Pubblicazione: (2025)
Detecting Clues for Skill Levels and Machine Operation Difficulty from Egocentric Vision
di: Long-fei, Chen, et al.
Pubblicazione: (2019)
di: Long-fei, Chen, et al.
Pubblicazione: (2019)
Visual Affect Analysis: Predicting Emotions of Image Viewers with Vision-Language Models
di: Nowicki, Filip, et al.
Pubblicazione: (2026)
di: Nowicki, Filip, et al.
Pubblicazione: (2026)
Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision
di: Li, Chentao, et al.
Pubblicazione: (2026)
di: Li, Chentao, et al.
Pubblicazione: (2026)
EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision
di: Zhao, Yiming, et al.
Pubblicazione: (2024)
di: Zhao, Yiming, et al.
Pubblicazione: (2024)
Vid2Coach: Transforming How-To Videos into Task Assistants
di: Huh, Mina, et al.
Pubblicazione: (2025)
di: Huh, Mina, et al.
Pubblicazione: (2025)
Revisiting Human-in-the-Loop Object Retrieval with Pre-Trained Vision Transformers
di: Zaher, Kawtar, et al.
Pubblicazione: (2026)
di: Zaher, Kawtar, et al.
Pubblicazione: (2026)
Viewpoint Recommendation for Point Cloud Labeling through Interaction Cost Modeling
di: Zhang, Yu, et al.
Pubblicazione: (2026)
di: Zhang, Yu, et al.
Pubblicazione: (2026)
ICo3D: An Interactive Conversational 3D Virtual Human
di: Shaw, Richard, et al.
Pubblicazione: (2026)
di: Shaw, Richard, et al.
Pubblicazione: (2026)
Visually Grounded Narratives: Reducing Cognitive Burden in Researcher-Participant Interaction
di: Wu, Runtong, et al.
Pubblicazione: (2025)
di: Wu, Runtong, et al.
Pubblicazione: (2025)
HpEIS: Learning Hand Pose Embeddings for Multimedia Interactive Systems
di: Xu, Songpei, et al.
Pubblicazione: (2024)
di: Xu, Songpei, et al.
Pubblicazione: (2024)
When Technologies Are Not Enough: Understanding How Domestic Workers Employ (and Avoid) Online Technologies in Their Work Practices
di: Fernandez-Espinosa, Mariana, et al.
Pubblicazione: (2025)
di: Fernandez-Espinosa, Mariana, et al.
Pubblicazione: (2025)
Towards Context-aware Support for Color Vision Deficiency: An Approach Integrating LLM and AR
di: Morita, Shogo, et al.
Pubblicazione: (2024)
di: Morita, Shogo, et al.
Pubblicazione: (2024)
HarassGuard: Detecting Harassment Behaviors in Social Virtual Reality with Vision-Language Models
di: Lee, Junhee, et al.
Pubblicazione: (2026)
di: Lee, Junhee, et al.
Pubblicazione: (2026)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
di: Fei, Hao, et al.
Pubblicazione: (2024)
di: Fei, Hao, et al.
Pubblicazione: (2024)
Toward Scalable Co-located Practical Learning: Assisting with Computer Vision and Multimodal Analytics
di: Li, Xinyu, et al.
Pubblicazione: (2026)
di: Li, Xinyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
di: Wróbel, Adam, et al.
Pubblicazione: (2026) -
Augmenting Teamwork through AI Agents as Spatial Collaborators
di: Fernandez-Espinosa, Mariana, et al.
Pubblicazione: (2025) -
Interactivity x Explainability: Toward Understanding How Interactivity Can Improve Computer Vision Explanations
di: Panigrahi, Indu, et al.
Pubblicazione: (2025) -
Behavioral Engagement in VR-Based Sign Language Learning: Visual Attention as a Predictor of Performance and Temporal Dynamics
di: Traini, Davide, et al.
Pubblicazione: (2026) -
QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
di: Wang, Yuxiao, et al.
Pubblicazione: (2025)