Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Bum Jun, Kim, Sang Woo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Disappearance of Timestep Embedding in Modern Time-Dependent Neural Networks
by: Kim, Bum Jun, et al.
Published: (2024)
by: Kim, Bum Jun, et al.
Published: (2024)
Stochastic Subsampling With Average Pooling
by: Kim, Bum Jun, et al.
Published: (2024)
by: Kim, Bum Jun, et al.
Published: (2024)
Safe Semi-Supervised Contrastive Learning Using In-Distribution Data as Positive Examples
by: Kwak, Min Gu, et al.
Published: (2024)
by: Kwak, Min Gu, et al.
Published: (2024)
Oscillation-Reduced MXFP4 Training for Vision Transformers
by: Chen, Yuxiang, et al.
Published: (2025)
by: Chen, Yuxiang, et al.
Published: (2025)
Improving Diffusion-Based Image Editing Faithfulness via Guidance and Scheduling
by: Cho, Hansam, et al.
Published: (2025)
by: Cho, Hansam, et al.
Published: (2025)
Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
by: Kim, Gahyeon, et al.
Published: (2025)
by: Kim, Gahyeon, et al.
Published: (2025)
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task
by: Kim, Yunho, et al.
Published: (2024)
by: Kim, Yunho, et al.
Published: (2024)
Temporal Pair Consistency for Variance-Reduced Flow Matching
by: Maduabuchi, Chika, et al.
Published: (2026)
by: Maduabuchi, Chika, et al.
Published: (2026)
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
LipShiFT: A Certifiably Robust Shift-based Vision Transformer
by: Menon, Rohan, et al.
Published: (2025)
by: Menon, Rohan, et al.
Published: (2025)
Maximizing the Position Embedding for Vision Transformers with Global Average Pooling
by: Lee, Wonjun, et al.
Published: (2025)
by: Lee, Wonjun, et al.
Published: (2025)
Variational Partial Group Convolutions for Input-Aware Partial Equivariance of Rotations and Color-Shifts
by: Kim, Hyunsu, et al.
Published: (2024)
by: Kim, Hyunsu, et al.
Published: (2024)
Accelerating Vision Transformers with Adaptive Patch Sizes
by: Choudhury, Rohan, et al.
Published: (2025)
by: Choudhury, Rohan, et al.
Published: (2025)
Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings
by: Kim, Bumjun, et al.
Published: (2026)
by: Kim, Bumjun, et al.
Published: (2026)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
by: Choi, Kanghyun, et al.
Published: (2024)
by: Choi, Kanghyun, et al.
Published: (2024)
RESQUE: Quantifying Estimator to Task and Distribution Shift for Sustainable Model Reusability
by: Sangarya, Vishwesh, et al.
Published: (2024)
by: Sangarya, Vishwesh, et al.
Published: (2024)
AAPL: Adding Attributes to Prompt Learning for Vision-Language Models
by: Kim, Gahyeon, et al.
Published: (2024)
by: Kim, Gahyeon, et al.
Published: (2024)
Self-Predictive Dynamics for Generalization of Vision-based Reinforcement Learning
by: Kim, Kyungsoo, et al.
Published: (2025)
by: Kim, Kyungsoo, et al.
Published: (2025)
GradMix: Gradient-based Selective Mixup for Robust Data Augmentation in Class-Incremental Learning
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Scratching Visual Transformer's Back with Uniform Attention
by: Hyeon-Woo, Nam, et al.
Published: (2022)
by: Hyeon-Woo, Nam, et al.
Published: (2022)
Cooperative Meta-Learning with Gradient Augmentation
by: Shin, Jongyun, et al.
Published: (2024)
by: Shin, Jongyun, et al.
Published: (2024)
VisTabNet: Adapting Vision Transformers for Tabular Data
by: Wydmański, Witold, et al.
Published: (2024)
by: Wydmański, Witold, et al.
Published: (2024)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
by: Sui, Elaine, et al.
Published: (2024)
by: Sui, Elaine, et al.
Published: (2024)
End-to-End Breast Cancer Radiotherapy Planning via LMMs with Consistency Embedding
by: Kim, Kwanyoung, et al.
Published: (2023)
by: Kim, Kwanyoung, et al.
Published: (2023)
Sample Selection via Contrastive Fragmentation for Noisy Label Regression
by: Kim, Chris Dongjoo, et al.
Published: (2025)
by: Kim, Chris Dongjoo, et al.
Published: (2025)
Topology-Preserving Polygon Augmentation for Segmentation in Structured Visual Domains
by: Laudari, Sudip, et al.
Published: (2026)
by: Laudari, Sudip, et al.
Published: (2026)
MRI Plane Orientation Detection using a Context-Aware 2.5D Model
by: Kim, SangHyuk, et al.
Published: (2025)
by: Kim, SangHyuk, et al.
Published: (2025)
Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models
by: Wang, Hongjun, et al.
Published: (2026)
by: Wang, Hongjun, et al.
Published: (2026)
VG3T: Visual Geometry Grounded Gaussian Transformer
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
Steering to Say No: Configurable Refusal via Activation Steering in Vision Language Models
by: Yang, Jiaxi, et al.
Published: (2026)
by: Yang, Jiaxi, et al.
Published: (2026)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
by: Lee, Soeun, et al.
Published: (2024)
by: Lee, Soeun, et al.
Published: (2024)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
by: Kim, Taewhan, et al.
Published: (2024)
by: Kim, Taewhan, et al.
Published: (2024)
Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation
by: Kim, Taehoon, et al.
Published: (2025)
by: Kim, Taehoon, et al.
Published: (2025)
Learning Unified Distance Metric Across Diverse Data Distributions with Parameter-Efficient Transfer Learning
by: Kim, Sungyeon, et al.
Published: (2023)
by: Kim, Sungyeon, et al.
Published: (2023)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Sparse Model Inversion: Efficient Inversion of Vision Transformers for Data-Free Applications
by: Hu, Zixuan, et al.
Published: (2025)
by: Hu, Zixuan, et al.
Published: (2025)
Semi-Supervised Masked Autoencoders: Unlocking Vision Transformer Potential with Limited Data
by: Faysal, Atik, et al.
Published: (2026)
by: Faysal, Atik, et al.
Published: (2026)
I0T: Embedding Standardization Method Towards Zero Modality Gap
by: An, Na Min, et al.
Published: (2024)
by: An, Na Min, et al.
Published: (2024)
SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data
by: Kim, Dong-Hee, et al.
Published: (2025)
by: Kim, Dong-Hee, et al.
Published: (2025)
Rethinking Post-Unlearning Behavior of Large Vision-Language Models
by: Kim, Minsung, et al.
Published: (2025)
by: Kim, Minsung, et al.
Published: (2025)
Similar Items
-
The Disappearance of Timestep Embedding in Modern Time-Dependent Neural Networks
by: Kim, Bum Jun, et al.
Published: (2024) -
Stochastic Subsampling With Average Pooling
by: Kim, Bum Jun, et al.
Published: (2024) -
Safe Semi-Supervised Contrastive Learning Using In-Distribution Data as Positive Examples
by: Kwak, Min Gu, et al.
Published: (2024) -
Oscillation-Reduced MXFP4 Training for Vision Transformers
by: Chen, Yuxiang, et al.
Published: (2025) -
Improving Diffusion-Based Image Editing Faithfulness via Guidance and Scheduling
by: Cho, Hansam, et al.
Published: (2025)