Channel Vision Transformers: An Image Is Worth 1 x 16 x 16 Words
Fuente:
arXiv
Saved in:
| Main Authors: | Bao, Yujia, Sivanandan, Srinivasan, Karaletsos, Theofanis |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Low-Resolution Image is Worth 1x1 Words: Enabling Fine Image Super-Resolution with Transformers and TaylorShift
by: Nagaraju, Sanath Budakegowdanadoddi, et al.
Published: (2024)
by: Nagaraju, Sanath Budakegowdanadoddi, et al.
Published: (2024)
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
by: Nguyen, Duy-Kien, et al.
Published: (2024)
by: Nguyen, Duy-Kien, et al.
Published: (2024)
Variational Control for Guidance in Diffusion Models
by: Pandey, Kushagra, et al.
Published: (2025)
by: Pandey, Kushagra, et al.
Published: (2025)
Is an Image Also Worth 16x16=256 Superpixels? A Framework for Attentional Image Classification
by: Avelar, Pedro Henrique da Costa, et al.
Published: (2026)
by: Avelar, Pedro Henrique da Costa, et al.
Published: (2026)
QMViT: A Mushroom is worth 16x16 Words
by: Dutta, Siddhant, et al.
Published: (2024)
by: Dutta, Siddhant, et al.
Published: (2024)
Fast Vision Mamba: Pooling Spatial Dimensions for Accelerated Processing
by: Kapse, Saarthak, et al.
Published: (2025)
by: Kapse, Saarthak, et al.
Published: (2025)
Vision-LSTM: xLSTM as Generic Vision Backbone
by: Alkin, Benedikt, et al.
Published: (2024)
by: Alkin, Benedikt, et al.
Published: (2024)
MorphGen: Controllable and Morphologically Plausible Generative Cell-Imaging
by: Demirel, Berker, et al.
Published: (2025)
by: Demirel, Berker, et al.
Published: (2025)
An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning
by: Jin, Chen, et al.
Published: (2023)
by: Jin, Chen, et al.
Published: (2023)
IMTS is Worth Time $\times$ Channel Patches: Visual Masked Autoencoders for Irregular Multivariate Time Series Prediction
by: Hu, Zhangyi, et al.
Published: (2025)
by: Hu, Zhangyi, et al.
Published: (2025)
Seg-LSTM: Performance of xLSTM for Semantic Segmentation of Remotely Sensed Images
by: Zhu, Qinfeng, et al.
Published: (2024)
by: Zhu, Qinfeng, et al.
Published: (2024)
A Noise is Worth Diffusion Guidance
by: Ahn, Donghoon, et al.
Published: (2024)
by: Ahn, Donghoon, et al.
Published: (2024)
DiffiT: Diffusion Vision Transformers for Image Generation
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
VariViT: A Vision Transformer for Variable Image Sizes
by: Varma, Aswathi, et al.
Published: (2026)
by: Varma, Aswathi, et al.
Published: (2026)
One Image is Worth a Thousand Words: A Usability Preservable Text-Image Collaborative Erasing Framework
by: Li, Feiran, et al.
Published: (2025)
by: Li, Feiran, et al.
Published: (2025)
A More Word-like Image Tokenization for MLLMs
by: Lee, Hyun, et al.
Published: (2026)
by: Lee, Hyun, et al.
Published: (2026)
Adaptive Knowledge Distillation for Classification of Hand Images using Explainable Vision Transformers
by: Nguyen, Thanh Thi, et al.
Published: (2024)
by: Nguyen, Thanh Thi, et al.
Published: (2024)
Detection of subclinical atherosclerosis by image-based deep learning on chest x-ray
by: Gallone, Guglielmo, et al.
Published: (2024)
by: Gallone, Guglielmo, et al.
Published: (2024)
Enhancing Vision-Language Model Pre-training with Image-text Pair Pruning Based on Word Frequency
by: Liang, Mingliang, et al.
Published: (2024)
by: Liang, Mingliang, et al.
Published: (2024)
Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
by: Huybrechts, Goeric, et al.
Published: (2025)
by: Huybrechts, Goeric, et al.
Published: (2025)
WordVIS: A Color Worth A Thousand Words
by: Khan, Umar, et al.
Published: (2024)
by: Khan, Umar, et al.
Published: (2024)
One Prompt Word is Enough to Boost Adversarial Robustness for Pre-trained Vision-Language Models
by: Li, Lin, et al.
Published: (2024)
by: Li, Lin, et al.
Published: (2024)
Fusion of Pervasive RF Data with Spatial Images via Vision Transformers for Enhanced Mapping in Smart Cities
by: Mkrtchyan, Rafayel, et al.
Published: (2025)
by: Mkrtchyan, Rafayel, et al.
Published: (2025)
GLoG-CSUnet: Enhancing Vision Transformers with Adaptable Radiomic Features for Medical Image Segmentation
by: Eghbali, Niloufar, et al.
Published: (2025)
by: Eghbali, Niloufar, et al.
Published: (2025)
Knee-xRAI: An Explainable AI Framework for Automatic Kellgren-Lawrence Grading of Knee Osteoarthritis
by: Irfan, Azmul A., et al.
Published: (2026)
by: Irfan, Azmul A., et al.
Published: (2026)
Improving Interpretation Faithfulness for Vision Transformers
by: Hu, Lijie, et al.
Published: (2023)
by: Hu, Lijie, et al.
Published: (2023)
Block-Recurrent Dynamics in Vision Transformers
by: Jacobs, Mozes, et al.
Published: (2025)
by: Jacobs, Mozes, et al.
Published: (2025)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
by: Kim, Wonkyun, et al.
Published: (2024)
by: Kim, Wonkyun, et al.
Published: (2024)
Deterministic Continuous Replacement: Fast and Stable Module Replacement in Pretrained Transformers
by: Bradbury, Rowan, et al.
Published: (2025)
by: Bradbury, Rowan, et al.
Published: (2025)
Demographic Bias of Expert-Level Vision-Language Foundation Models in Medical Imaging
by: Yang, Yuzhe, et al.
Published: (2024)
by: Yang, Yuzhe, et al.
Published: (2024)
Continual Adaptation of Vision Transformers for Federated Learning
by: Halbe, Shaunak, et al.
Published: (2023)
by: Halbe, Shaunak, et al.
Published: (2023)
Class-Discriminative Attention Maps for Vision Transformers
by: Brocki, Lennart, et al.
Published: (2023)
by: Brocki, Lennart, et al.
Published: (2023)
Mechanisms of Non-Monotonic Scaling in Vision Transformers
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
Discovering Influential Neuron Path in Vision Transformers
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
ADAPT to Robustify Prompt Tuning Vision Transformers
by: Eskandar, Masih, et al.
Published: (2024)
by: Eskandar, Masih, et al.
Published: (2024)
Accelerating Vision Transformers with Adaptive Patch Sizes
by: Choudhury, Rohan, et al.
Published: (2025)
by: Choudhury, Rohan, et al.
Published: (2025)
Accessing Vision Foundation Models via ImageNet-1K
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
by: Wang, Jiayu, et al.
Published: (2024)
by: Wang, Jiayu, et al.
Published: (2024)
FasterViT: Fast Vision Transformers with Hierarchical Attention
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
Similar Items
-
A Low-Resolution Image is Worth 1x1 Words: Enabling Fine Image Super-Resolution with Transformers and TaylorShift
by: Nagaraju, Sanath Budakegowdanadoddi, et al.
Published: (2024) -
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
by: Nguyen, Duy-Kien, et al.
Published: (2024) -
Variational Control for Guidance in Diffusion Models
by: Pandey, Kushagra, et al.
Published: (2025) -
Is an Image Also Worth 16x16=256 Superpixels? A Framework for Attentional Image Classification
by: Avelar, Pedro Henrique da Costa, et al.
Published: (2026) -
QMViT: A Mushroom is worth 16x16 Words
by: Dutta, Siddhant, et al.
Published: (2024)