Block Selective Reprogramming for On-device Training of Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Sarkar, Sreetama, Kundu, Souvik, Zheng, Kai, Beerel, Peter A. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppression
by: Sarkar, Sreetama, et al.
Published: (2025)
by: Sarkar, Sreetama, et al.
Published: (2025)
MaskVD: Region Masking for Efficient Video Object Detection
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning
by: Xu, Jingqi, et al.
Published: (2025)
by: Xu, Jingqi, et al.
Published: (2025)
Linearizing Models for Efficient yet Robust Private Inference
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
HIVTP: A Training-Free Method to Improve VLMs Efficiency via Hierarchical Visual Token Pruning Using Middle-Layer-Based Importance Score
by: Xu, Jingqi, et al.
Published: (2025)
by: Xu, Jingqi, et al.
Published: (2025)
Energy-Efficient & Real-Time Computer Vision with Intelligent Skipping via Reconfigurable CMOS Image Sensors
by: Kaiser, Md Abdullah-Al, et al.
Published: (2024)
by: Kaiser, Md Abdullah-Al, et al.
Published: (2024)
MABViT -- Modified Attention Block Enhances Vision Transformers
by: Ramesh, Mahesh, et al.
Published: (2023)
by: Ramesh, Mahesh, et al.
Published: (2023)
Attribute-based Visual Reprogramming for Vision-Language Models
by: Cai, Chengyi, et al.
Published: (2025)
by: Cai, Chengyi, et al.
Published: (2025)
Block-Recurrent Dynamics in Vision Transformers
by: Jacobs, Mozes, et al.
Published: (2025)
by: Jacobs, Mozes, et al.
Published: (2025)
Permutation-equivariant quantum convolutional neural networks
by: Das, Sreetama, et al.
Published: (2024)
by: Das, Sreetama, et al.
Published: (2024)
Synthetic Designed Experiments for Diagnosing Vision Model Failure
by: Sarkar, Krisanu
Published: (2026)
by: Sarkar, Krisanu
Published: (2026)
Contrastive Forward-Forward: A Training Algorithm of Vision Transformer
by: Aghagolzadeh, Hossein, et al.
Published: (2025)
by: Aghagolzadeh, Hossein, et al.
Published: (2025)
Inducing Spatial Locality in Vision Transformers through the Training Protocol
by: Toledo, Eduardo Santiago, et al.
Published: (2026)
by: Toledo, Eduardo Santiago, et al.
Published: (2026)
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
by: Jiang, Jiarui, et al.
Published: (2024)
by: Jiang, Jiarui, et al.
Published: (2024)
Ternarization of Vision Language Models for use on edge devices
by: Crulis, Ben, et al.
Published: (2025)
by: Crulis, Ben, et al.
Published: (2025)
From Colors to Classes: Emergence of Concepts in Vision Transformers
by: Dorszewski, Teresa, et al.
Published: (2025)
by: Dorszewski, Teresa, et al.
Published: (2025)
ViTNT-FIQA: Training-Free Face Image Quality Assessment with Vision Transformers
by: Ozgur, Guray, et al.
Published: (2026)
by: Ozgur, Guray, et al.
Published: (2026)
The role of data embedding in equivariant quantum convolutional neural networks
by: Das, Sreetama, et al.
Published: (2023)
by: Das, Sreetama, et al.
Published: (2023)
Bayesian-guided Label Mapping for Visual Reprogramming
by: Cai, Chengyi, et al.
Published: (2024)
by: Cai, Chengyi, et al.
Published: (2024)
LetheViT: Selective Machine Unlearning for Vision Transformers via Attention-Guided Contrastive Learning
by: Tong, Yujia, et al.
Published: (2025)
by: Tong, Yujia, et al.
Published: (2025)
RLS3: RL-Based Synthetic Sample Selection to Enhance Spatial Reasoning in Vision-Language Models for Indoor Autonomous Perception
by: Waite, Joshua R., et al.
Published: (2025)
by: Waite, Joshua R., et al.
Published: (2025)
CortiNet: A Physics-Perception Hybrid Cortical-Inspired Dual-Stream Network for Gallbladder Disease Diagnosis from Ultrasound
by: Kumar, Vagish, et al.
Published: (2026)
by: Kumar, Vagish, et al.
Published: (2026)
Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models
by: Zheng, Shunjie-Fabian, et al.
Published: (2025)
by: Zheng, Shunjie-Fabian, et al.
Published: (2025)
Sample-specific Masks for Visual Reprogramming-based Prompting
by: Cai, Chengyi, et al.
Published: (2024)
by: Cai, Chengyi, et al.
Published: (2024)
Self-Supervised Vision Transformers for Writer Retrieval
by: Raven, Tim, et al.
Published: (2024)
by: Raven, Tim, et al.
Published: (2024)
Oscillation-Reduced MXFP4 Training for Vision Transformers
by: Chen, Yuxiang, et al.
Published: (2025)
by: Chen, Yuxiang, et al.
Published: (2025)
ELSA: Exploiting Layer-wise N:M Sparsity for Vision Transformer Acceleration
by: Huang, Ning-Chi, et al.
Published: (2024)
by: Huang, Ning-Chi, et al.
Published: (2024)
Vision Transformers Exhibit Human-Like Biases: Evidence of Orientation and Color Selectivity, Categorical Perception, and Phase Transitions
by: Bahador, Nooshin
Published: (2025)
by: Bahador, Nooshin
Published: (2025)
Native Segmentation Vision Transformers
by: Brasó, Guillem, et al.
Published: (2025)
by: Brasó, Guillem, et al.
Published: (2025)
Understanding Model Reprogramming for CLIP via Decoupling Visual Prompts
by: Cai, Chengyi, et al.
Published: (2025)
by: Cai, Chengyi, et al.
Published: (2025)
Region Masking to Accelerate Video Processing on Neuromorphic Hardware
by: Sarkar, Sreetama, et al.
Published: (2025)
by: Sarkar, Sreetama, et al.
Published: (2025)
On the Use of Anchoring for Training Vision Models
by: Narayanaswamy, Vivek, et al.
Published: (2024)
by: Narayanaswamy, Vivek, et al.
Published: (2024)
Training Feature Attribution for Vision Models
by: Bacha, Aziz, et al.
Published: (2025)
by: Bacha, Aziz, et al.
Published: (2025)
Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
by: Huo, Simin, et al.
Published: (2025)
by: Huo, Simin, et al.
Published: (2025)
Matryoshka Query Transformer for Large Vision-Language Models
by: Hu, Wenbo, et al.
Published: (2024)
by: Hu, Wenbo, et al.
Published: (2024)
S-GRPO: Unified Post-Training for Large Vision-Language Models
by: Yan, Yuming, et al.
Published: (2026)
by: Yan, Yuming, et al.
Published: (2026)
SpecFormer: Guarding Vision Transformer Robustness via Maximum Singular Value Penalization
by: Hu, Xixu, et al.
Published: (2024)
by: Hu, Xixu, et al.
Published: (2024)
Slicing Vision Transformer for Flexible Inference
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
RAViT: Resolution-Adaptive Vision Transformer
by: Guidez, Martial, et al.
Published: (2026)
by: Guidez, Martial, et al.
Published: (2026)
Similar Items
-
Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppression
by: Sarkar, Sreetama, et al.
Published: (2025) -
MaskVD: Region Masking for Efficient Video Object Detection
by: Sarkar, Sreetama, et al.
Published: (2024) -
RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning
by: Xu, Jingqi, et al.
Published: (2025) -
Linearizing Models for Efficient yet Robust Private Inference
by: Sarkar, Sreetama, et al.
Published: (2024) -
HIVTP: A Training-Free Method to Improve VLMs Efficiency via Hierarchical Visual Token Pruning Using Middle-Layer-Based Importance Score
by: Xu, Jingqi, et al.
Published: (2025)