Convolutional Initialization for Data-Efficient Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Jianqiao, Li, Xueqian, Lucey, Simon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structured Initialization for Attention in Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2024)
by: Zheng, Jianqiao, et al.
Published: (2024)
Structured Initialization for Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2025)
by: Zheng, Jianqiao, et al.
Published: (2025)
Trading Positional Complexity vs. Deepness in Coordinate Networks
by: Zheng, Jianqiao, et al.
Published: (2022)
by: Zheng, Jianqiao, et al.
Published: (2022)
Fast Kernel Scene Flow
by: Li, Xueqian, et al.
Published: (2024)
by: Li, Xueqian, et al.
Published: (2024)
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Multi-Body Neural Scene Flow
by: Vidanapathirana, Kavisha, et al.
Published: (2023)
by: Vidanapathirana, Kavisha, et al.
Published: (2023)
Enhancing Transformers Through Conditioned Embedded Tokens
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
From Tables to Signals: Revealing Spectral Adaptivity in TabPFN
by: Zheng, Jianqiao, et al.
Published: (2025)
by: Zheng, Jianqiao, et al.
Published: (2025)
Convolutional Networks as Extremely Small Foundation Models: Visual Prompting and Theoretical Perspective
by: Wangni, Jianqiao
Published: (2024)
by: Wangni, Jianqiao
Published: (2024)
From Activation to Initialization: Scaling Insights for Optimizing Neural Fields
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
SineProject: Machine Unlearning for Stable Vision Language Alignment
by: Garg, Arpit, et al.
Published: (2025)
by: Garg, Arpit, et al.
Published: (2025)
Dense Cross-Connected Ensemble Convolutional Neural Networks for Enhanced Model Robustness
by: Wang, Longwei, et al.
Published: (2024)
by: Wang, Longwei, et al.
Published: (2024)
Gradient Descent as a Shrinkage Operator for Spectral Bias
by: Lucey, Simon
Published: (2025)
by: Lucey, Simon
Published: (2025)
Leaner Transformers: More Heads, Less Depth
by: Saratchandran, Hemanth, et al.
Published: (2025)
by: Saratchandran, Hemanth, et al.
Published: (2025)
Depth-Wise Convolutions in Vision Transformers for Efficient Training on Small Datasets
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
by: Zhang, Tianfang, et al.
Published: (2024)
by: Zhang, Tianfang, et al.
Published: (2024)
On Convolutional Vision Transformers for Yield Prediction
by: Inderka, Alvin, et al.
Published: (2024)
by: Inderka, Alvin, et al.
Published: (2024)
TiC: Exploring Vision Transformer in Convolution
by: Zhang, Song, et al.
Published: (2023)
by: Zhang, Song, et al.
Published: (2023)
MPM: Mutual Pair Merging for Efficient Vision Transformers
by: Ravé, Simon, et al.
Published: (2026)
by: Ravé, Simon, et al.
Published: (2026)
Weight Conditioning for Smooth Optimization of Neural Networks
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
ACC-ViT : Atrous Convolution's Comeback in Vision Transformers
by: Ibtehaz, Nabil, et al.
Published: (2024)
by: Ibtehaz, Nabil, et al.
Published: (2024)
DeNAS-ViT: Data Efficient NAS-Optimized Vision Transformer for Ultrasound Image Segmentation
by: Chen, Renqi, et al.
Published: (2024)
by: Chen, Renqi, et al.
Published: (2024)
Multiple-Exit Tuning: Towards Inference-Efficient Adaptation for Vision Transformer
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
SMORE: Simultaneous Map and Object REconstruction
by: Chodosh, Nathaniel, et al.
Published: (2024)
by: Chodosh, Nathaniel, et al.
Published: (2024)
3D Gaussian Point Encoders
by: James, Jim, et al.
Published: (2025)
by: James, Jim, et al.
Published: (2025)
ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages
by: Qian, Zhoujie
Published: (2025)
by: Qian, Zhoujie
Published: (2025)
Vision Transformers and Convolutional Neural Networks for Land Use Scene Classification
by: Kulkarni, Arun D.
Published: (2026)
by: Kulkarni, Arun D.
Published: (2026)
Enhancing Learnable Descriptive Convolutional Vision Transformer for Face Anti-Spoofing
by: Huanga, Pei-Kai, et al.
Published: (2025)
by: Huanga, Pei-Kai, et al.
Published: (2025)
Rethinking the Role of Spatial Mixing
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
Invertible Neural Warp for NeRF
by: Chng, Shin-Fang, et al.
Published: (2024)
by: Chng, Shin-Fang, et al.
Published: (2024)
LowFormer: Hardware Efficient Design for Convolutional Transformer Backbones
by: Nottebaum, Moritz, et al.
Published: (2024)
by: Nottebaum, Moritz, et al.
Published: (2024)
FullLoRA: Efficiently Boosting the Robustness of Pretrained Vision Transformers
by: Yuan, Zheng, et al.
Published: (2024)
by: Yuan, Zheng, et al.
Published: (2024)
DCFormer: Efficient 3D Vision-Language Modeling with Decomposed Convolutions
by: Ates, Gorkem Can, et al.
Published: (2025)
by: Ates, Gorkem Can, et al.
Published: (2025)
Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting
by: Xu, Runze, et al.
Published: (2026)
by: Xu, Runze, et al.
Published: (2026)
On Quantizing Implicit Neural Representations
by: Gordon, Cameron, et al.
Published: (2022)
by: Gordon, Cameron, et al.
Published: (2022)
GenConViT: Deepfake Video Detection Using Generative Convolutional Vision Transformer
by: Deressa, Deressa Wodajo, et al.
Published: (2023)
by: Deressa, Deressa Wodajo, et al.
Published: (2023)
PDC-ViT : Source Camera Identification using Pixel Difference Convolution and Vision Transformer
by: Elharrouss, Omar, et al.
Published: (2025)
by: Elharrouss, Omar, et al.
Published: (2025)
Preconditioners for the Stochastic Training of Neural Fields
by: Chng, Shin-Fang, et al.
Published: (2024)
by: Chng, Shin-Fang, et al.
Published: (2024)
Centroid-centered Modeling for Efficient Vision Transformer Pre-training
by: Yan, Xin, et al.
Published: (2023)
by: Yan, Xin, et al.
Published: (2023)
Efficient Learning With Sine-Activated Low-rank Matrices
by: Ji, Yiping, et al.
Published: (2024)
by: Ji, Yiping, et al.
Published: (2024)
Similar Items
-
Structured Initialization for Attention in Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2024) -
Structured Initialization for Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2025) -
Trading Positional Complexity vs. Deepness in Coordinate Networks
by: Zheng, Jianqiao, et al.
Published: (2022) -
Fast Kernel Scene Flow
by: Li, Xueqian, et al.
Published: (2024) -
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024)