Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hartman, Max, Jayaraman, Vidhata, Choraria, Moulik, Bhimaraju, Akhil, Varshney, Lav R. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
by: Choraria, Moulik, et al.
Published: (2025)
by: Choraria, Moulik, et al.
Published: (2025)
Watermarking Discrete Diffusion Language Models
by: Bagchi, Avi, et al.
Published: (2025)
by: Bagchi, Avi, et al.
Published: (2025)
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
by: Hartman, Max, et al.
Published: (2026)
by: Hartman, Max, et al.
Published: (2026)
Semantically Grounded QFormer for Efficient Vision Language Understanding
by: Choraria, Moulik, et al.
Published: (2023)
by: Choraria, Moulik, et al.
Published: (2023)
SwitchCIT: Switching for Continual Instruction Tuning
by: Wu, Xinbo, et al.
Published: (2024)
by: Wu, Xinbo, et al.
Published: (2024)
Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping
by: Zeng, Weili, et al.
Published: (2025)
by: Zeng, Weili, et al.
Published: (2025)
GM-Skip: Metric-Guided Transformer Block Skipping for Efficient Vision-Language Models
by: Huang, Lianming, et al.
Published: (2025)
by: Huang, Lianming, et al.
Published: (2025)
Skip \n: A Simple Method to Reduce Hallucination in Large Vision-Language Models
by: Han, Zongbo, et al.
Published: (2024)
by: Han, Zongbo, et al.
Published: (2024)
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
by: Huang, Yushi, et al.
Published: (2025)
by: Huang, Yushi, et al.
Published: (2025)
Skip and Skip: Segmenting Medical Images with Prompts
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Context-Gated Associative Retrieval: From Theory to Transformers
by: Choraria, Moulik, et al.
Published: (2026)
by: Choraria, Moulik, et al.
Published: (2026)
SkipVAR: Accelerating Visual Autoregressive Modeling via Adaptive Frequency-Aware Skipping
by: Li, Jiajun, et al.
Published: (2025)
by: Li, Jiajun, et al.
Published: (2025)
SkipGS: Post-Densification Backward Skipping for Efficient 3DGS Training
by: Li, Jingxing, et al.
Published: (2026)
by: Li, Jingxing, et al.
Published: (2026)
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
by: Wu, Shihan, et al.
Published: (2024)
by: Wu, Shihan, et al.
Published: (2024)
On the Vulnerability of Skip Connections to Model Inversion Attacks
by: Koh, Jun Hao, et al.
Published: (2024)
by: Koh, Jun Hao, et al.
Published: (2024)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
by: Zhang, Dongxu, et al.
Published: (2026)
by: Zhang, Dongxu, et al.
Published: (2026)
SkipViT: Speeding Up Vision Transformers with a Token-Level Skip Connection
by: Ataiefard, Foozhan, et al.
Published: (2024)
by: Ataiefard, Foozhan, et al.
Published: (2024)
ColPali: Efficient Document Retrieval with Vision Language Models
by: Faysse, Manuel, et al.
Published: (2024)
by: Faysse, Manuel, et al.
Published: (2024)
TokenCom: Vision-Language Model for Multimodal and Multitask Token Communications
by: Jiang, Feibo, et al.
Published: (2026)
by: Jiang, Feibo, et al.
Published: (2026)
Always Skip Attention
by: Ji, Yiping, et al.
Published: (2025)
by: Ji, Yiping, et al.
Published: (2025)
SkipSR: Faster Super Resolution with Token Skipping
by: Choudhury, Rohan, et al.
Published: (2025)
by: Choudhury, Rohan, et al.
Published: (2025)
SSI-DM: Singularity Skipping Inversion of Diffusion Models
by: Min, Chen, et al.
Published: (2026)
by: Min, Chen, et al.
Published: (2026)
ProSMA-UNet: Decoder Conditioning for Proximal-Sparse Skip Feature Selection
by: Cheng, Chun-Wun, et al.
Published: (2026)
by: Cheng, Chun-Wun, et al.
Published: (2026)
Skip-WaveNet: A Wavelet based Multi-scale Architecture to Trace Snow Layers in Radar Echograms
by: Varshney, Debvrat, et al.
Published: (2023)
by: Varshney, Debvrat, et al.
Published: (2023)
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering
by: Zhu, Yingjian, et al.
Published: (2026)
by: Zhu, Yingjian, et al.
Published: (2026)
Optimization of Layer Skipping and Frequency Scaling for Convolutional Neural Networks under Latency Constraint
by: Chan, Minh David Thao, et al.
Published: (2025)
by: Chan, Minh David Thao, et al.
Published: (2025)
Theoretical Guarantees of Data Augmented Last Layer Retraining Methods
by: Welfert, Monica, et al.
Published: (2024)
by: Welfert, Monica, et al.
Published: (2024)
Skipping Computations in Multimodal LLMs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Rethinking Skip Connections: Additive U-Net for Robust and Interpretable Denoising
by: Lakkavalli, Vikram R
Published: (2026)
by: Lakkavalli, Vikram R
Published: (2026)
Raw Instinct: Trust Your Classifiers and Skip the Conversion
by: Kantas, Christos, et al.
Published: (2024)
by: Kantas, Christos, et al.
Published: (2024)
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
by: Hyung, Junha, et al.
Published: (2024)
by: Hyung, Junha, et al.
Published: (2024)
Energy-Efficient & Real-Time Computer Vision with Intelligent Skipping via Reconfigurable CMOS Image Sensors
by: Kaiser, Md Abdullah-Al, et al.
Published: (2024)
by: Kaiser, Md Abdullah-Al, et al.
Published: (2024)
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models
by: Pandya, Pranshu, et al.
Published: (2024)
by: Pandya, Pranshu, et al.
Published: (2024)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
by: Yang, Jheng-Hong, et al.
Published: (2024)
by: Yang, Jheng-Hong, et al.
Published: (2024)
Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models
by: Bachu, Saketh, et al.
Published: (2024)
by: Bachu, Saketh, et al.
Published: (2024)
Visual Language Model based Cross-modal Semantic Communication Systems
by: Jiang, Feibo, et al.
Published: (2024)
by: Jiang, Feibo, et al.
Published: (2024)
Pretrained Diffusion Models Are Inherently Skipped-Step Samplers
by: Xu, Wenju
Published: (2025)
by: Xu, Wenju
Published: (2025)
Beyond Skip Connection: Pooling and Unpooling Design for Elimination Singularities
by: Sun, Chengkun, et al.
Published: (2024)
by: Sun, Chengkun, et al.
Published: (2024)
SkipcrossNets: Adaptive Skip-cross Fusion for Road Detection
by: Gong, Yan, et al.
Published: (2023)
by: Gong, Yan, et al.
Published: (2023)
ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
by: Wan, Zifu, et al.
Published: (2025)
by: Wan, Zifu, et al.
Published: (2025)
Similar Items
-
DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
by: Choraria, Moulik, et al.
Published: (2025) -
Watermarking Discrete Diffusion Language Models
by: Bagchi, Avi, et al.
Published: (2025) -
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
by: Hartman, Max, et al.
Published: (2026) -
Semantically Grounded QFormer for Efficient Vision Language Understanding
by: Choraria, Moulik, et al.
Published: (2023) -
SwitchCIT: Switching for Continual Instruction Tuning
by: Wu, Xinbo, et al.
Published: (2024)