Accelerating Vision Transformers with Adaptive Patch Sizes
Fuente:
arXiv
Saved in:
| Main Authors: | Choudhury, Rohan, Kim, JungEun, Park, Jinhyung, Yang, Eunho, Jeni, László A., Kitani, Kris M. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data-Efficient Unsupervised Interpolation Without Any Intermediate Frame for 4D Medical Images
by: Kim, JungEun, et al.
Published: (2024)
by: Kim, JungEun, et al.
Published: (2024)
SkipSR: Faster Super Resolution with Token Skipping
by: Choudhury, Rohan, et al.
Published: (2025)
by: Choudhury, Rohan, et al.
Published: (2025)
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
by: Choudhury, Rohan, et al.
Published: (2024)
by: Choudhury, Rohan, et al.
Published: (2024)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
by: Kim, Sohee, et al.
Published: (2025)
by: Kim, Sohee, et al.
Published: (2025)
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
by: Kim, Jisoo, et al.
Published: (2026)
by: Kim, Jisoo, et al.
Published: (2026)
LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding
by: Jang, Doohyuk, et al.
Published: (2024)
by: Jang, Doohyuk, et al.
Published: (2024)
Generalizable Neural Human Renderer
by: Masuda, Mana, et al.
Published: (2024)
by: Masuda, Mana, et al.
Published: (2024)
PruNeRF: Segment-Centric Dataset Pruning via 3D Spatial Consistency
by: Jung, Yeonsung, et al.
Published: (2024)
by: Jung, Yeonsung, et al.
Published: (2024)
EXAONEPath 1.0 Patch-level Foundation Model for Pathology
by: Yun, Juseung, et al.
Published: (2024)
by: Yun, Juseung, et al.
Published: (2024)
3D-LFM: Lifting Foundation Model
by: Dabhi, Mosam, et al.
Published: (2023)
by: Dabhi, Mosam, et al.
Published: (2023)
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
by: Park, Joonhyung, et al.
Published: (2025)
by: Park, Joonhyung, et al.
Published: (2025)
Digital-to-Physical Transfer of Adversarial Patches for Aerial Vehicle Detection
by: Woo, Jung Heum, et al.
Published: (2026)
by: Woo, Jung Heum, et al.
Published: (2026)
Multi-Person 3D Pose Estimation from Multi-View Uncalibrated Depth Cameras
by: Li, Yu-Jhe, et al.
Published: (2024)
by: Li, Yu-Jhe, et al.
Published: (2024)
SceneAdapt: Scene-aware Adaptation of Human Motion Diffusion
by: Cho, Jungbin, et al.
Published: (2025)
by: Cho, Jungbin, et al.
Published: (2025)
S2GO: Streaming Sparse Gaussian Occupancy Prediction
by: Park, Jinhyung, et al.
Published: (2025)
by: Park, Jinhyung, et al.
Published: (2025)
AT-SNN: Adaptive Tokens for Vision Transformer on Spiking Neural Network
by: Kang, Donghwa, et al.
Published: (2024)
by: Kang, Donghwa, et al.
Published: (2024)
Fourier-Guided Attention Upsampling for Image Super-Resolution
by: Choi, Daejune, et al.
Published: (2025)
by: Choi, Daejune, et al.
Published: (2025)
Improved Ear Verification with Vision Transformers and Overlapping Patches
by: Arun, Deeksha, et al.
Published: (2025)
by: Arun, Deeksha, et al.
Published: (2025)
AlignDiff: Learning Physically-Grounded Camera Alignment via Diffusion
by: Xie, Liuyue, et al.
Published: (2025)
by: Xie, Liuyue, et al.
Published: (2025)
Integrating Multimodal Large Language Model Knowledge into Amodal Completion
by: Yun, Heecheol, et al.
Published: (2026)
by: Yun, Heecheol, et al.
Published: (2026)
Mixed Non-linear Quantization for Vision Transformers
by: Kim, Gihwan, et al.
Published: (2024)
by: Kim, Gihwan, et al.
Published: (2024)
Med-PerSAM: One-Shot Visual Prompt Tuning for Personalized Segment Anything Model in Medical Domain
by: Yoon, Hangyul, et al.
Published: (2024)
by: Yoon, Hangyul, et al.
Published: (2024)
Patch Rebirth: Toward Fast and Transferable Model Inversion of Vision Transformers
by: Heo, Seongsoo, et al.
Published: (2025)
by: Heo, Seongsoo, et al.
Published: (2025)
LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers
by: Chowdhury, Md Abtahi Majeed, et al.
Published: (2025)
by: Chowdhury, Md Abtahi Majeed, et al.
Published: (2025)
Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping
by: Park, Sunghyun, et al.
Published: (2026)
by: Park, Sunghyun, et al.
Published: (2026)
DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
by: Kim, Dahye, et al.
Published: (2026)
by: Kim, Dahye, et al.
Published: (2026)
Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing
by: Kim, Yoonjeon, et al.
Published: (2024)
by: Kim, Yoonjeon, et al.
Published: (2024)
Accelerating Vision Transformers on Brain Processing Unit
by: Tang, Jinchi, et al.
Published: (2026)
by: Tang, Jinchi, et al.
Published: (2026)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
Diffusion Model Patching via Mixture-of-Prompts
by: Ham, Seokil, et al.
Published: (2024)
by: Ham, Seokil, et al.
Published: (2024)
Progressive Growing of Patch Size: Curriculum Learning for Accelerated and Improved Medical Image Segmentation
by: Fischer, Stefan M., et al.
Published: (2025)
by: Fischer, Stefan M., et al.
Published: (2025)
When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
by: Lu, Hui, et al.
Published: (2025)
by: Lu, Hui, et al.
Published: (2025)
Adaptive Patching for High-resolution Image Segmentation with Transformers
by: Zhang, Enzhi, et al.
Published: (2024)
by: Zhang, Enzhi, et al.
Published: (2024)
Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
by: Im, Eun Woo, et al.
Published: (2025)
by: Im, Eun Woo, et al.
Published: (2025)
ALERT Open Dataset and Input-Size-Agnostic Vision Transformer for Driver Activity Recognition using IR-UWB
by: Park, Jeongjun, et al.
Published: (2025)
by: Park, Jeongjun, et al.
Published: (2025)
Patch is Enough: Naturalistic Adversarial Patch against Vision-Language Pre-training Models
by: Kong, Dehong, et al.
Published: (2024)
by: Kong, Dehong, et al.
Published: (2024)
NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning
by: Fu, Jiahui, et al.
Published: (2026)
by: Fu, Jiahui, et al.
Published: (2026)
RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
by: Xu, Xuwei, et al.
Published: (2025)
by: Xu, Xuwei, et al.
Published: (2025)
Estimating Environmental Cost Throughout Model's Adaptive Life Cycle
by: Sangarya, Vishwesh, et al.
Published: (2024)
by: Sangarya, Vishwesh, et al.
Published: (2024)
Vision Transformer-Conditioned UNet for Domain-Adaptive Semantic Segmentation
by: Ortega, Joel Valdivia, et al.
Published: (2026)
by: Ortega, Joel Valdivia, et al.
Published: (2026)
Similar Items
-
Data-Efficient Unsupervised Interpolation Without Any Intermediate Frame for 4D Medical Images
by: Kim, JungEun, et al.
Published: (2024) -
SkipSR: Faster Super Resolution with Token Skipping
by: Choudhury, Rohan, et al.
Published: (2025) -
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
by: Choudhury, Rohan, et al.
Published: (2024) -
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
by: Kim, Sohee, et al.
Published: (2025) -
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
by: Kim, Jisoo, et al.
Published: (2026)