Frozen Vision Transformers for Dense Prediction on Small Datasets: A Case Study in Arrow Localization
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Shepherd, Maxwell |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdapterTune: Zero-Initialized Low-Rank Adapters for Frozen Vision Transformers
von: Khazem, Salim
Veröffentlicht: (2026)
von: Khazem, Salim
Veröffentlicht: (2026)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
von: Li, Yi, et al.
Veröffentlicht: (2026)
von: Li, Yi, et al.
Veröffentlicht: (2026)
Frozen Forecasting: A Unified Evaluation
von: Walker, Jacob C, et al.
Veröffentlicht: (2025)
von: Walker, Jacob C, et al.
Veröffentlicht: (2025)
Inference-Path Optimization via Circuit Duplication in Frozen Visual Transformers for Marine Species Classification
von: Rost, Thomas Manuel
Veröffentlicht: (2026)
von: Rost, Thomas Manuel
Veröffentlicht: (2026)
Convolutional Neural Nets vs Vision Transformers: A SpaceNet Case Study with Balanced vs Imbalanced Regimes
von: Gothi, Akshar
Veröffentlicht: (2025)
von: Gothi, Akshar
Veröffentlicht: (2025)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
von: Pang, Ziqi, et al.
Veröffentlicht: (2023)
von: Pang, Ziqi, et al.
Veröffentlicht: (2023)
DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction
von: Liao, Guiqiu, et al.
Veröffentlicht: (2026)
von: Liao, Guiqiu, et al.
Veröffentlicht: (2026)
DeiT-LT Distillation Strikes Back for Vision Transformer Training on Long-Tailed Datasets
von: Rangwani, Harsh, et al.
Veröffentlicht: (2024)
von: Rangwani, Harsh, et al.
Veröffentlicht: (2024)
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
Can Visual Encoder Learn to See Arrows?
von: Terashita, Naoyuki, et al.
Veröffentlicht: (2025)
von: Terashita, Naoyuki, et al.
Veröffentlicht: (2025)
Gaze-Informed Vision Transformers: Predicting Driving Decisions Under Uncertainty
von: Koorathota, Sharath, et al.
Veröffentlicht: (2023)
von: Koorathota, Sharath, et al.
Veröffentlicht: (2023)
Vertical LoRA: Dense Expectation-Maximization Interpretation of Transformers
von: Fu, Zhuolin
Veröffentlicht: (2024)
von: Fu, Zhuolin
Veröffentlicht: (2024)
Semantic Compositions Enhance Vision-Language Contrastive Learning
von: Aladago, Maxwell, et al.
Veröffentlicht: (2024)
von: Aladago, Maxwell, et al.
Veröffentlicht: (2024)
RGB-Th-Bench: A Dense benchmark for Visual-Thermal Understanding of Vision Language Models
von: Moshtaghi, Mehdi, et al.
Veröffentlicht: (2025)
von: Moshtaghi, Mehdi, et al.
Veröffentlicht: (2025)
Parameter Reduction Improves Vision Transformers: A Comparative Study of Sharing and Width Reduction
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
ALERT Open Dataset and Input-Size-Agnostic Vision Transformer for Driver Activity Recognition using IR-UWB
von: Park, Jeongjun, et al.
Veröffentlicht: (2025)
von: Park, Jeongjun, et al.
Veröffentlicht: (2025)
Case Study: Transformer-Based Solution for the Automatic Digitization of Gas Plants
von: Bailo, I., et al.
Veröffentlicht: (2025)
von: Bailo, I., et al.
Veröffentlicht: (2025)
A3-FPN: Asymptotic Content-Aware Pyramid Attention Network for Dense Visual Prediction
von: Qin, Meng'en, et al.
Veröffentlicht: (2026)
von: Qin, Meng'en, et al.
Veröffentlicht: (2026)
Using Synthetic Images to Augment Small Medical Image Datasets
von: Vu, Minh H., et al.
Veröffentlicht: (2025)
von: Vu, Minh H., et al.
Veröffentlicht: (2025)
COVID19 Prediction Based On CT Scans Of Lungs Using DenseNet Architecture
von: Sanyal, Deborup
Veröffentlicht: (2025)
von: Sanyal, Deborup
Veröffentlicht: (2025)
EGA: Adapting Frozen Encoders for Vector Search with Bounded Out-of-Distribution Degradation
von: Zhao, Dongfang
Veröffentlicht: (2026)
von: Zhao, Dongfang
Veröffentlicht: (2026)
Improving Interpretation Faithfulness for Vision Transformers
von: Hu, Lijie, et al.
Veröffentlicht: (2023)
von: Hu, Lijie, et al.
Veröffentlicht: (2023)
Block-Recurrent Dynamics in Vision Transformers
von: Jacobs, Mozes, et al.
Veröffentlicht: (2025)
von: Jacobs, Mozes, et al.
Veröffentlicht: (2025)
VariViT: A Vision Transformer for Variable Image Sizes
von: Varma, Aswathi, et al.
Veröffentlicht: (2026)
von: Varma, Aswathi, et al.
Veröffentlicht: (2026)
A Survey of the Self Supervised Learning Mechanisms for Vision Transformers
von: Khan, Asifullah, et al.
Veröffentlicht: (2024)
von: Khan, Asifullah, et al.
Veröffentlicht: (2024)
Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2026)
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2026)
Exploring Curriculum Learning for Vision-Language Tasks: A Study on Small-Scale Multimodal Training
von: Saha, Rohan, et al.
Veröffentlicht: (2024)
von: Saha, Rohan, et al.
Veröffentlicht: (2024)
Modelling and Simulation of Neuromorphic Datasets for Anomaly Detection in Computer Vision
von: Middleton, Mike, et al.
Veröffentlicht: (2026)
von: Middleton, Mike, et al.
Veröffentlicht: (2026)
Continual Adaptation of Vision Transformers for Federated Learning
von: Halbe, Shaunak, et al.
Veröffentlicht: (2023)
von: Halbe, Shaunak, et al.
Veröffentlicht: (2023)
Mechanisms of Non-Monotonic Scaling in Vision Transformers
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
von: Kumar, Anantha Padmanaban Krishna
Veröffentlicht: (2025)
Discovering Influential Neuron Path in Vision Transformers
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
DiffiT: Diffusion Vision Transformers for Image Generation
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
ADAPT to Robustify Prompt Tuning Vision Transformers
von: Eskandar, Masih, et al.
Veröffentlicht: (2024)
von: Eskandar, Masih, et al.
Veröffentlicht: (2024)
Accelerating Vision Transformers with Adaptive Patch Sizes
von: Choudhury, Rohan, et al.
Veröffentlicht: (2025)
von: Choudhury, Rohan, et al.
Veröffentlicht: (2025)
Class-Discriminative Attention Maps for Vision Transformers
von: Brocki, Lennart, et al.
Veröffentlicht: (2023)
von: Brocki, Lennart, et al.
Veröffentlicht: (2023)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
von: Cazenavette, George, et al.
Veröffentlicht: (2025)
von: Cazenavette, George, et al.
Veröffentlicht: (2025)
An Interpretable Implicit-Based Approach for Modeling Local Spatial Effects: A Case Study of Global Gross Primary Productivity
von: Du, Siqi, et al.
Veröffentlicht: (2025)
von: Du, Siqi, et al.
Veröffentlicht: (2025)
Generative Dataset Distillation: Balancing Global Structure and Local Details
von: Li, Longzhen, et al.
Veröffentlicht: (2024)
von: Li, Longzhen, et al.
Veröffentlicht: (2024)
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
von: Salman, Shaeke, et al.
Veröffentlicht: (2024)
Oscillation-Reduced MXFP4 Training for Vision Transformers
von: Chen, Yuxiang, et al.
Veröffentlicht: (2025)
von: Chen, Yuxiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AdapterTune: Zero-Initialized Low-Rank Adapters for Frozen Vision Transformers
von: Khazem, Salim
Veröffentlicht: (2026) -
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
von: Li, Yi, et al.
Veröffentlicht: (2026) -
Frozen Forecasting: A Unified Evaluation
von: Walker, Jacob C, et al.
Veröffentlicht: (2025) -
Inference-Path Optimization via Circuit Duplication in Frozen Visual Transformers for Marine Species Classification
von: Rost, Thomas Manuel
Veröffentlicht: (2026) -
Convolutional Neural Nets vs Vision Transformers: A SpaceNet Case Study with Balanced vs Imbalanced Regimes
von: Gothi, Akshar
Veröffentlicht: (2025)