big.LITTLE Vision Transformer for Efficient Visual Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, He, Wang, Yulong, Ye, Zixuan, Dai, Jifeng, Xiong, Yuwen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
by: Hou, Zhi, et al.
Published: (2025)
by: Hou, Zhi, et al.
Published: (2025)
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
by: Xiong, Yuwen, et al.
Published: (2024)
by: Xiong, Yuwen, et al.
Published: (2024)
Revolutionizing Traffic Sign Recognition: Unveiling the Potential of Vision Transformers
by: Mingwin, Susano, et al.
Published: (2024)
by: Mingwin, Susano, et al.
Published: (2024)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
by: Duan, Yuchen, et al.
Published: (2024)
by: Duan, Yuchen, et al.
Published: (2024)
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
by: Wu, Zehuan, et al.
Published: (2024)
by: Wu, Zehuan, et al.
Published: (2024)
Diffusion Transformer Policy
by: Hou, Zhi, et al.
Published: (2024)
by: Hou, Zhi, et al.
Published: (2024)
Adaptively Bypassing Vision Transformer Blocks for Efficient Visual Tracking
by: Yang, Xiangyang, et al.
Published: (2024)
by: Yang, Xiangyang, et al.
Published: (2024)
EDTformer: An Efficient Decoder Transformer for Visual Place Recognition
by: Jin, Tong, et al.
Published: (2024)
by: Jin, Tong, et al.
Published: (2024)
FViT: A Focal Vision Transformer with Gabor Filter
by: Shi, Yulong, et al.
Published: (2024)
by: Shi, Yulong, et al.
Published: (2024)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
by: He, Xuanhua, et al.
Published: (2025)
by: He, Xuanhua, et al.
Published: (2025)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
by: Ge, Junqi, et al.
Published: (2024)
by: Ge, Junqi, et al.
Published: (2024)
EViT: An Eagle Vision Transformer with Bi-Fovea Self-Attention
by: Shi, Yulong, et al.
Published: (2023)
by: Shi, Yulong, et al.
Published: (2023)
Flexible and Efficient Spatio-Temporal Transformer for Sequential Visual Place Recognition
by: Kiu, Yu, et al.
Published: (2025)
by: Kiu, Yu, et al.
Published: (2025)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
by: Yang, Chenyu, et al.
Published: (2024)
by: Yang, Chenyu, et al.
Published: (2024)
Image Recognition with Online Lightweight Vision Transformer: A Survey
by: Zhang, Zherui, et al.
Published: (2025)
by: Zhang, Zherui, et al.
Published: (2025)
Shuffle Vision Transformer: Lightweight, Fast and Efficient Recognition of Driver Facial Expression
by: Saadi, Ibtissam, et al.
Published: (2024)
by: Saadi, Ibtissam, et al.
Published: (2024)
Efficient Vision-Language Pre-training by Cluster Masking
by: Wei, Zihao, et al.
Published: (2024)
by: Wei, Zihao, et al.
Published: (2024)
TAT-VPR: Ternary Adaptive Transformer for Dynamic and Efficient Visual Place Recognition
by: Grainge, Oliver, et al.
Published: (2025)
by: Grainge, Oliver, et al.
Published: (2025)
TransNeXt: Robust Foveal Visual Perception for Vision Transformers
by: Shi, Dai
Published: (2023)
by: Shi, Dai
Published: (2023)
TCFormer: Visual Recognition via Token Clustering Transformer
by: Zeng, Wang, et al.
Published: (2024)
by: Zeng, Wang, et al.
Published: (2024)
In-Context Matting
by: Guo, He, et al.
Published: (2024)
by: Guo, He, et al.
Published: (2024)
GlyphPattern: An Abstract Pattern Recognition Benchmark for Vision-Language Models
by: Wu, Zixuan, et al.
Published: (2024)
by: Wu, Zixuan, et al.
Published: (2024)
Region Matters: Efficient and Reliable Region-Aware Visual Place Recognition
by: Chen, Shunpeng, et al.
Published: (2026)
by: Chen, Shunpeng, et al.
Published: (2026)
Emotion Separation and Recognition from a Facial Expression by Generating the Poker Face with Vision Transformers
by: Li, Jia, et al.
Published: (2022)
by: Li, Jia, et al.
Published: (2022)
Improved EATFormer: A Vision Transformer for Medical Image Classification
by: Shisu, Yulong, et al.
Published: (2024)
by: Shisu, Yulong, et al.
Published: (2024)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
by: Chen, Zhe, et al.
Published: (2023)
by: Chen, Zhe, et al.
Published: (2023)
AIQViT: Architecture-Informed Post-Training Quantization for Vision Transformers
by: Jiang, Runqing, et al.
Published: (2025)
by: Jiang, Runqing, et al.
Published: (2025)
Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCR
by: Li, Zhenyang, et al.
Published: (2024)
by: Li, Zhenyang, et al.
Published: (2024)
Multi-Tailed Vision Transformer for Efficient Inference
by: Wang, Yunke, et al.
Published: (2022)
by: Wang, Yunke, et al.
Published: (2022)
Feature Complementation Architecture for Visual Place Recognition
by: Wang, Weiwei, et al.
Published: (2025)
by: Wang, Weiwei, et al.
Published: (2025)
TiC: Exploring Vision Transformer in Convolution
by: Zhang, Song, et al.
Published: (2023)
by: Zhang, Song, et al.
Published: (2023)
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
by: Ning, Shan, et al.
Published: (2026)
by: Ning, Shan, et al.
Published: (2026)
Synergistic Prompting for Robust Visual Recognition with Missing Modalities
by: Zhang, Zhihui, et al.
Published: (2025)
by: Zhang, Zhihui, et al.
Published: (2025)
BSViT: A Burst Spiking Vision Transformer for Expressive and Efficient Visual Representation Learning
by: Peng, Hongxiang, et al.
Published: (2026)
by: Peng, Hongxiang, et al.
Published: (2026)
LM-MCVT: A Lightweight Multi-modal Multi-view Convolutional-Vision Transformer Approach for 3D Object Recognition
by: Xiong, Songsong, et al.
Published: (2025)
by: Xiong, Songsong, et al.
Published: (2025)
HTR-VT: Handwritten Text Recognition with Vision Transformer
by: Li, Yuting, et al.
Published: (2024)
by: Li, Yuting, et al.
Published: (2024)
ColorSense: A Study on Color Vision in Machine Visual Recognition
by: Chiu, Ming-Chang, et al.
Published: (2022)
by: Chiu, Ming-Chang, et al.
Published: (2022)
SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition
by: Chen, Shunpeng, et al.
Published: (2025)
by: Chen, Shunpeng, et al.
Published: (2025)
VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference
by: Zhu, Hao, et al.
Published: (2026)
by: Zhu, Hao, et al.
Published: (2026)
Similar Items
-
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
by: Hou, Zhi, et al.
Published: (2025) -
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
by: Xiong, Yuwen, et al.
Published: (2024) -
Revolutionizing Traffic Sign Recognition: Unveiling the Potential of Vision Transformers
by: Mingwin, Susano, et al.
Published: (2024) -
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
by: Duan, Yuchen, et al.
Published: (2024) -
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
by: Wu, Zehuan, et al.
Published: (2024)