Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Ben, Chen, Xin, Zhao, Jie, Bo, Chunjuan, Wang, Dong, Lu, Huchuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UETrack: A Unified and Efficient Framework for Single Object Tracking
by: Kang, Ben, et al.
Published: (2026)
by: Kang, Ben, et al.
Published: (2026)
Efficient Motion Prompt Learning for Robust Visual Tracking
by: Zhao, Jie, et al.
Published: (2025)
by: Zhao, Jie, et al.
Published: (2025)
SUTrack: Towards Simple and Unified Single Object Tracking
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
Unified Sequence-to-Sequence Learning for Single- and Multi-Modal Visual Object Tracking
by: Chen, Xin, et al.
Published: (2023)
by: Chen, Xin, et al.
Published: (2023)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
by: Chattopadhyay, Nandish, et al.
Published: (2026)
by: Chattopadhyay, Nandish, et al.
Published: (2026)
ERVD: An Efficient and Robust ViT-Based Distillation Framework for Remote Sensing Image Retrieval
by: Dong, Le, et al.
Published: (2024)
by: Dong, Le, et al.
Published: (2024)
ViTCAE: ViT-based Class-conditioned Autoencoder
by: Jebraeeli, Vahid, et al.
Published: (2025)
by: Jebraeeli, Vahid, et al.
Published: (2025)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
by: Chen, Lu, et al.
Published: (2025)
by: Chen, Lu, et al.
Published: (2025)
ProtoS-ViT: Visual foundation models for sparse self-explainable classifications
by: Turbé, Hugues, et al.
Published: (2024)
by: Turbé, Hugues, et al.
Published: (2024)
Two-stream Beats One-stream: Asymmetric Siamese Network for Efficient Visual Tracking
by: Zhu, Jiawen, et al.
Published: (2025)
by: Zhu, Jiawen, et al.
Published: (2025)
How to train your ViT for OOD Detection
by: Mueller, Maximilian, et al.
Published: (2024)
by: Mueller, Maximilian, et al.
Published: (2024)
SRRT: Exploring Search Region Regulation for Visual Object Tracking
by: Zhu, Jiawen, et al.
Published: (2022)
by: Zhu, Jiawen, et al.
Published: (2022)
HydraViT: Stacking Heads for a Scalable ViT
by: Haberer, Janek, et al.
Published: (2024)
by: Haberer, Janek, et al.
Published: (2024)
Exploring Dynamic Transformer for Efficient Object Tracking
by: Zhu, Jiawen, et al.
Published: (2024)
by: Zhu, Jiawen, et al.
Published: (2024)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
by: Zhu, Chen, et al.
Published: (2025)
by: Zhu, Chen, et al.
Published: (2025)
GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
by: Lee, Hyunju, et al.
Published: (2025)
by: Lee, Hyunju, et al.
Published: (2025)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
by: Zhong, Yunshan, et al.
Published: (2023)
by: Zhong, Yunshan, et al.
Published: (2023)
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
by: Alvetreti, Federico, et al.
Published: (2025)
by: Alvetreti, Federico, et al.
Published: (2025)
ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers
by: Sinha, Sanchit, et al.
Published: (2025)
by: Sinha, Sanchit, et al.
Published: (2025)
Token Cropr: Faster ViTs for Quite a Few Tasks
by: Bergner, Benjamin, et al.
Published: (2024)
by: Bergner, Benjamin, et al.
Published: (2024)
CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
by: Gong, Wenyi, et al.
Published: (2025)
by: Gong, Wenyi, et al.
Published: (2025)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
by: Li, Zhengang, et al.
Published: (2024)
by: Li, Zhengang, et al.
Published: (2024)
Scene-Graph ViT: End-to-End Open-Vocabulary Visual Relationship Detection
by: Salzmann, Tim, et al.
Published: (2024)
by: Salzmann, Tim, et al.
Published: (2024)
Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
by: Zhao, Wangbo, et al.
Published: (2024)
by: Zhao, Wangbo, et al.
Published: (2024)
Parameter-Efficient Subspace Decoupling ViT for Mitigating Multi-Task Negative Transfer in Histological Scoring
by: Huang, Youhan, et al.
Published: (2026)
by: Huang, Youhan, et al.
Published: (2026)
Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentation
by: Tang, Fenghe, et al.
Published: (2025)
by: Tang, Fenghe, et al.
Published: (2025)
RepViT: Revisiting Mobile CNN From ViT Perspective
by: Wang, Ao, et al.
Published: (2023)
by: Wang, Ao, et al.
Published: (2023)
LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
by: Nag, Shashank, et al.
Published: (2025)
by: Nag, Shashank, et al.
Published: (2025)
Layer by layer, module by module: Choose both for optimal OOD probing of ViT
by: Odonnat, Ambroise, et al.
Published: (2026)
by: Odonnat, Ambroise, et al.
Published: (2026)
LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition
by: Hu, Youbing, et al.
Published: (2024)
by: Hu, Youbing, et al.
Published: (2024)
LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors
by: Suri, Saksham, et al.
Published: (2024)
by: Suri, Saksham, et al.
Published: (2024)
RELO: Reinforcement Learning to Localize for Visual Object Tracking
by: Chen, Xin, et al.
Published: (2026)
by: Chen, Xin, et al.
Published: (2026)
Deeper Inside Deep ViT
by: Hong, Sungrae
Published: (2025)
by: Hong, Sungrae
Published: (2025)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
by: Hwang, Dongyoon, et al.
Published: (2024)
by: Hwang, Dongyoon, et al.
Published: (2024)
Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification
by: Gustineli, Murilo, et al.
Published: (2025)
by: Gustineli, Murilo, et al.
Published: (2025)
MambaVT: Spatio-Temporal Contextual Modeling for robust RGB-T Tracking
by: Lai, Simiao, et al.
Published: (2024)
by: Lai, Simiao, et al.
Published: (2024)
Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
by: Yunusa, Haruna, et al.
Published: (2024)
by: Yunusa, Haruna, et al.
Published: (2024)
Causality $\neq$ Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs
by: Huang, Lianghuan, et al.
Published: (2025)
by: Huang, Lianghuan, et al.
Published: (2025)
Training-Free Acceleration of ViTs with Delayed Spatial Merging
by: Heo, Jung Hwan, et al.
Published: (2023)
by: Heo, Jung Hwan, et al.
Published: (2023)
Octic Vision Transformers: Quicker ViTs Through Equivariance
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
Similar Items
-
UETrack: A Unified and Efficient Framework for Single Object Tracking
by: Kang, Ben, et al.
Published: (2026) -
Efficient Motion Prompt Learning for Robust Visual Tracking
by: Zhao, Jie, et al.
Published: (2025) -
SUTrack: Towards Simple and Unified Single Object Tracking
by: Chen, Xin, et al.
Published: (2024) -
Unified Sequence-to-Sequence Learning for Single- and Multi-Modal Visual Object Tracking
by: Chen, Xin, et al.
Published: (2023) -
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
by: Chattopadhyay, Nandish, et al.
Published: (2026)