Maximizing the Position Embedding for Vision Transformers with Global Average Pooling
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Wonjun, Ham, Bumsub, Kim, Suhyun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Instance-Aware Group Quantization for Vision Transformers
by: Moon, Jaehyeon, et al.
Published: (2024)
by: Moon, Jaehyeon, et al.
Published: (2024)
AZ-NAS: Assembling Zero-Cost Proxies for Network Architecture Search
by: Lee, Junghyup, et al.
Published: (2024)
by: Lee, Junghyup, et al.
Published: (2024)
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
by: Lee, Wonjun, et al.
Published: (2025)
by: Lee, Wonjun, et al.
Published: (2025)
Relational Feature Caching for Accelerating Diffusion Transformers
by: Son, Byunggwan, et al.
Published: (2026)
by: Son, Byunggwan, et al.
Published: (2026)
GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
by: Lee, Hyunju, et al.
Published: (2025)
by: Lee, Hyunju, et al.
Published: (2025)
ELITE: Enhanced Language-Image Toxicity Evaluation for Safety
by: Lee, Wonjun, et al.
Published: (2025)
by: Lee, Wonjun, et al.
Published: (2025)
D-TPT: Dimensional Entropy Maximization for Calibrating Test-Time Prompt Tuning in Vision-Language Models
by: Han, Jisu, et al.
Published: (2025)
by: Han, Jisu, et al.
Published: (2025)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
Stochastic Subsampling With Average Pooling
by: Kim, Bum Jun, et al.
Published: (2024)
by: Kim, Bum Jun, et al.
Published: (2024)
TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts
by: Jeon, Jeimin, et al.
Published: (2026)
by: Jeon, Jeimin, et al.
Published: (2026)
Probabilistic Precision and Recall Towards Reliable Evaluation of Generative Models
by: Park, Dogyun, et al.
Published: (2023)
by: Park, Dogyun, et al.
Published: (2023)
Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision Transformers
by: Kim, Bum Jun, et al.
Published: (2024)
by: Kim, Bum Jun, et al.
Published: (2024)
Efficient Few-Shot Neural Architecture Search by Counting the Number of Nonlinear Functions
by: Oh, Youngmin, et al.
Published: (2024)
by: Oh, Youngmin, et al.
Published: (2024)
Scheduling Weight Transitions for Quantization-Aware Training
by: Lee, Junghyup, et al.
Published: (2024)
by: Lee, Junghyup, et al.
Published: (2024)
From Pixels to Patches: Pooling Strategies for Earth Embeddings
by: Corley, Isaac, et al.
Published: (2026)
by: Corley, Isaac, et al.
Published: (2026)
Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection
by: Lee, Sanghoon, et al.
Published: (2026)
by: Lee, Sanghoon, et al.
Published: (2026)
Deep Learning for Melt Pool Depth Contour Prediction From Surface Thermal Images via Vision Transformers
by: Ogoke, Francis, et al.
Published: (2024)
by: Ogoke, Francis, et al.
Published: (2024)
HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
3DPillars: Pillar-based two-stage 3D object detection
by: Noh, Jongyoun, et al.
Published: (2025)
by: Noh, Jongyoun, et al.
Published: (2025)
Toward INT4 Fixed-Point Training via Exploring Quantization Error for Gradients
by: Kim, Dohyung, et al.
Published: (2024)
by: Kim, Dohyung, et al.
Published: (2024)
Towards a Better Evaluation of Out-of-Domain Generalization
by: Hwang, Duhun, et al.
Published: (2024)
by: Hwang, Duhun, et al.
Published: (2024)
Visualizing the loss landscape of Self-supervised Vision Transformer
by: Lee, Youngwan, et al.
Published: (2024)
by: Lee, Youngwan, et al.
Published: (2024)
Collaborative Learning with Multiple Foundation Models for Source-Free Domain Adaptation
by: Lee, Huisoo, et al.
Published: (2025)
by: Lee, Huisoo, et al.
Published: (2025)
Improving Hyperbolic Representations via Gromov-Wasserstein Regularization
by: Yang, Yifei, et al.
Published: (2024)
by: Yang, Yifei, et al.
Published: (2024)
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
by: Cho, Hyeonwoo, et al.
Published: (2026)
by: Cho, Hyeonwoo, et al.
Published: (2026)
Disentangled Representations for Short-Term and Long-Term Person Re-Identification
by: Eom, Chanho, et al.
Published: (2024)
by: Eom, Chanho, et al.
Published: (2024)
The Effects of Grouped Structural Global Pruning of Vision Transformers on Domain Generalisation
by: Riaz, Hamza, et al.
Published: (2025)
by: Riaz, Hamza, et al.
Published: (2025)
Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective on Robustness
by: Mannes, Mahmoud
Published: (2026)
by: Mannes, Mahmoud
Published: (2026)
Offline Meteorology-Pollution Coupling Global Air Pollution Forecasting Model with Bilinear Pooling
by: Fan, Xu, et al.
Published: (2025)
by: Fan, Xu, et al.
Published: (2025)
The Role of Masking for Efficient Supervised Knowledge Distillation of Vision Transformers
by: Son, Seungwoo, et al.
Published: (2023)
by: Son, Seungwoo, et al.
Published: (2023)
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
by: Kang, Suhyun, et al.
Published: (2024)
by: Kang, Suhyun, et al.
Published: (2024)
ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding
by: Setyawan, Novendra, et al.
Published: (2024)
by: Setyawan, Novendra, et al.
Published: (2024)
Ranked Entropy Minimization for Continual Test-Time Adaptation
by: Han, Jisu, et al.
Published: (2025)
by: Han, Jisu, et al.
Published: (2025)
Semantic Prompting with Image-Token for Continual Learning
by: Han, Jisu, et al.
Published: (2024)
by: Han, Jisu, et al.
Published: (2024)
Efficient Matrix Implementation for Rotary Position Embedding
by: Minqi, Chen, et al.
Published: (2026)
by: Minqi, Chen, et al.
Published: (2026)
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
FYI: Flip Your Images for Dataset Distillation
by: Son, Byunggwan, et al.
Published: (2024)
by: Son, Byunggwan, et al.
Published: (2024)
Rethinking the Use of Vision Transformers for AI-Generated Image Detection
by: Park, NaHyeon, et al.
Published: (2025)
by: Park, NaHyeon, et al.
Published: (2025)
PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding Projection
by: Molahasani, Mahdiyar, et al.
Published: (2025)
by: Molahasani, Mahdiyar, et al.
Published: (2025)
Activation Quantization of Vision Encoders Needs Prefixing Registers
by: Kim, Seunghyeon, et al.
Published: (2025)
by: Kim, Seunghyeon, et al.
Published: (2025)
Similar Items
-
Instance-Aware Group Quantization for Vision Transformers
by: Moon, Jaehyeon, et al.
Published: (2024) -
AZ-NAS: Assembling Zero-Cost Proxies for Network Architecture Search
by: Lee, Junghyup, et al.
Published: (2024) -
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
by: Lee, Wonjun, et al.
Published: (2025) -
Relational Feature Caching for Accelerating Diffusion Transformers
by: Son, Byunggwan, et al.
Published: (2026) -
GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
by: Lee, Hyunju, et al.
Published: (2025)