LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors
Fuente:
arXiv
Salvato in:
| Autori principali: | Suri, Saksham, Walmer, Matthew, Gupta, Kamal, Shrivastava, Abhinav |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
di: Walmer, Matthew, et al.
Pubblicazione: (2026)
di: Walmer, Matthew, et al.
Pubblicazione: (2026)
(LiFT) Lightweight Fitness Transformer: A language-vision model for Remote Monitoring of Physical Training
di: Postlmayr, A., et al.
Pubblicazione: (2025)
di: Postlmayr, A., et al.
Pubblicazione: (2025)
UVIS: Unsupervised Video Instance Segmentation
di: Huang, Shuaiyi, et al.
Pubblicazione: (2024)
di: Huang, Shuaiyi, et al.
Pubblicazione: (2024)
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
di: Agarwal, Vatsal, et al.
Pubblicazione: (2026)
di: Agarwal, Vatsal, et al.
Pubblicazione: (2026)
EAGLES: Efficient Accelerated 3D Gaussians with Lightweight EncodingS
di: Girish, Sharath, et al.
Pubblicazione: (2023)
di: Girish, Sharath, et al.
Pubblicazione: (2023)
LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment
di: Wang, Yibin, et al.
Pubblicazione: (2024)
di: Wang, Yibin, et al.
Pubblicazione: (2024)
LiFT: Lightweight, FPGA-tailored 3D object detection based on LiDAR data
di: Lis, Konrad, et al.
Pubblicazione: (2025)
di: Lis, Konrad, et al.
Pubblicazione: (2025)
Multi-entity Video Transformers for Fine-Grained Video Representation Learning
di: Walmer, Matthew, et al.
Pubblicazione: (2023)
di: Walmer, Matthew, et al.
Pubblicazione: (2023)
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
di: Wang, Hanyu, et al.
Pubblicazione: (2024)
di: Wang, Hanyu, et al.
Pubblicazione: (2024)
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
di: Xia, Chunlong, et al.
Pubblicazione: (2024)
di: Xia, Chunlong, et al.
Pubblicazione: (2024)
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
di: Kumar, Pulkit, et al.
Pubblicazione: (2025)
di: Kumar, Pulkit, et al.
Pubblicazione: (2025)
OSR-ViT: A Simple and Modular Framework for Open-Set Object Detection and Discovery
di: Inkawhich, Matthew, et al.
Pubblicazione: (2024)
di: Inkawhich, Matthew, et al.
Pubblicazione: (2024)
LiFT: Lifted Inter-slice Feature Trajectories for 3D Image Generation from 2D Generators
di: Zhang, Xinhe, et al.
Pubblicazione: (2026)
di: Zhang, Xinhe, et al.
Pubblicazione: (2026)
ViT-5: Vision Transformers for The Mid-2020s
di: Wang, Feng, et al.
Pubblicazione: (2026)
di: Wang, Feng, et al.
Pubblicazione: (2026)
ViT-VS: On the Applicability of Pretrained Vision Transformer Features for Generalizable Visual Servoing
di: Scherl, Alessandro, et al.
Pubblicazione: (2025)
di: Scherl, Alessandro, et al.
Pubblicazione: (2025)
ACC-ViT : Atrous Convolution's Comeback in Vision Transformers
di: Ibtehaz, Nabil, et al.
Pubblicazione: (2024)
di: Ibtehaz, Nabil, et al.
Pubblicazione: (2024)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
di: Zhu, Chen, et al.
Pubblicazione: (2025)
di: Zhu, Chen, et al.
Pubblicazione: (2025)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
di: Chattopadhyay, Nandish, et al.
Pubblicazione: (2026)
di: Chattopadhyay, Nandish, et al.
Pubblicazione: (2026)
TFS-ViT: Token-Level Feature Stylization for Domain Generalization
di: Noori, Mehrdad, et al.
Pubblicazione: (2023)
di: Noori, Mehrdad, et al.
Pubblicazione: (2023)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
di: Li, Yifan, et al.
Pubblicazione: (2026)
di: Li, Yifan, et al.
Pubblicazione: (2026)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
di: Ma, Xiaochen, et al.
Pubblicazione: (2023)
di: Ma, Xiaochen, et al.
Pubblicazione: (2023)
Deeper Inside Deep ViT
di: Hong, Sungrae
Pubblicazione: (2025)
di: Hong, Sungrae
Pubblicazione: (2025)
Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR Classification
di: Zhang, Jiaqing, et al.
Pubblicazione: (2024)
di: Zhang, Jiaqing, et al.
Pubblicazione: (2024)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
di: Zhong, Yunshan, et al.
Pubblicazione: (2023)
di: Zhong, Yunshan, et al.
Pubblicazione: (2023)
InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models
di: Saini, Nirat, et al.
Pubblicazione: (2024)
di: Saini, Nirat, et al.
Pubblicazione: (2024)
Efficient Continuous Video Flow Model for Video Prediction
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
RD-ViT: Recurrent-Depth Vision Transformer for Semantic Segmentation with Reduced Data Dependence Extending the Recurrent-Depth Transformer Architecture to Dense Prediction
di: He, Renjie
Pubblicazione: (2026)
di: He, Renjie
Pubblicazione: (2026)
Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking
di: Kang, Ben, et al.
Pubblicazione: (2025)
di: Kang, Ben, et al.
Pubblicazione: (2025)
EdgeCrafter: Compact ViTs for Edge Dense Prediction via Task-Specialized Distillation
di: Liu, Longfei, et al.
Pubblicazione: (2026)
di: Liu, Longfei, et al.
Pubblicazione: (2026)
LEIA: Latent View-invariant Embeddings for Implicit 3D Articulation
di: Swaminathan, Archana, et al.
Pubblicazione: (2024)
di: Swaminathan, Archana, et al.
Pubblicazione: (2024)
ViT-FIQA: Assessing Face Image Quality using Vision Transformers
di: Atzori, Andrea, et al.
Pubblicazione: (2025)
di: Atzori, Andrea, et al.
Pubblicazione: (2025)
MPTQ-ViT: Mixed-Precision Post-Training Quantization for Vision Transformer
di: Tai, Yu-Shan, et al.
Pubblicazione: (2024)
di: Tai, Yu-Shan, et al.
Pubblicazione: (2024)
ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline
di: Hernandez, Juan Manuel, et al.
Pubblicazione: (2026)
di: Hernandez, Juan Manuel, et al.
Pubblicazione: (2026)
LQ-Adapter: ViT-Adapter with Learnable Queries for Gallbladder Cancer Detection from Ultrasound Image
di: Madan, Chetan, et al.
Pubblicazione: (2024)
di: Madan, Chetan, et al.
Pubblicazione: (2024)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
di: Yao, Ting, et al.
Pubblicazione: (2024)
di: Yao, Ting, et al.
Pubblicazione: (2024)
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
di: Li, Wenhao, et al.
Pubblicazione: (2025)
di: Li, Wenhao, et al.
Pubblicazione: (2025)
RepViT: Revisiting Mobile CNN From ViT Perspective
di: Wang, Ao, et al.
Pubblicazione: (2023)
di: Wang, Ao, et al.
Pubblicazione: (2023)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
di: Kuzucu, Selim, et al.
Pubblicazione: (2025)
di: Kuzucu, Selim, et al.
Pubblicazione: (2025)
ADFQ-ViT: Activation-Distribution-Friendly Post-Training Quantization for Vision Transformers
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer
di: Dey, Sainath, et al.
Pubblicazione: (2025)
di: Dey, Sainath, et al.
Pubblicazione: (2025)
Documenti analoghi
-
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
di: Walmer, Matthew, et al.
Pubblicazione: (2026) -
(LiFT) Lightweight Fitness Transformer: A language-vision model for Remote Monitoring of Physical Training
di: Postlmayr, A., et al.
Pubblicazione: (2025) -
UVIS: Unsupervised Video Instance Segmentation
di: Huang, Shuaiyi, et al.
Pubblicazione: (2024) -
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
di: Agarwal, Vatsal, et al.
Pubblicazione: (2026) -
EAGLES: Efficient Accelerated 3D Gaussians with Lightweight EncodingS
di: Girish, Sharath, et al.
Pubblicazione: (2023)