Skipping Computations in Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Shukor, Mustafa, Cord, Matthieu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Learning to Steer: Input-dependent Steering for Multimodal LLMs
by: Parekh, Jayneel, et al.
Published: (2025)
by: Parekh, Jayneel, et al.
Published: (2025)
A Concept-Based Explainability Framework for Large Multimodal Models
by: Parekh, Jayneel, et al.
Published: (2024)
by: Parekh, Jayneel, et al.
Published: (2024)
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
by: Vallaeys, Théophane, et al.
Published: (2024)
by: Vallaeys, Théophane, et al.
Published: (2024)
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
by: Khayatan, Pegah, et al.
Published: (2025)
by: Khayatan, Pegah, et al.
Published: (2025)
Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
by: Shukor, Mustafa, et al.
Published: (2023)
by: Shukor, Mustafa, et al.
Published: (2023)
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
by: Khayatan, Pegah, et al.
Published: (2026)
by: Khayatan, Pegah, et al.
Published: (2026)
What Makes Multimodal In-Context Learning Work?
by: Baldassini, Folco Bertini, et al.
Published: (2024)
by: Baldassini, Folco Bertini, et al.
Published: (2024)
DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut
by: Couairon, Paul, et al.
Published: (2024)
by: Couairon, Paul, et al.
Published: (2024)
Scaling Laws for Native Multimodal Models
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion Models
by: Corradini, Barbara Toniella, et al.
Published: (2024)
by: Corradini, Barbara Toniella, et al.
Published: (2024)
Always Skip Attention
by: Ji, Yiping, et al.
Published: (2025)
by: Ji, Yiping, et al.
Published: (2025)
Reliability in Semantic Segmentation: Can We Use Synthetic Data?
by: Loiseau, Thibaut, et al.
Published: (2023)
by: Loiseau, Thibaut, et al.
Published: (2023)
ReGentS: Real-World Safety-Critical Driving Scenario Generation Made Stable
by: Yin, Yuan, et al.
Published: (2024)
by: Yin, Yuan, et al.
Published: (2024)
Zero-Shot Refinement of Buildings' Segmentation Models using SAM
by: Mayladan, Ali, et al.
Published: (2023)
by: Mayladan, Ali, et al.
Published: (2023)
SkipSR: Faster Super Resolution with Token Skipping
by: Choudhury, Rohan, et al.
Published: (2025)
by: Choudhury, Rohan, et al.
Published: (2025)
Multimodal Autoregressive Pre-training of Large Vision Encoders
by: Fini, Enrico, et al.
Published: (2024)
by: Fini, Enrico, et al.
Published: (2024)
Pretrained Diffusion Models Are Inherently Skipped-Step Samplers
by: Xu, Wenju
Published: (2025)
by: Xu, Wenju
Published: (2025)
SkipViT: Speeding Up Vision Transformers with a Token-Level Skip Connection
by: Ataiefard, Foozhan, et al.
Published: (2024)
by: Ataiefard, Foozhan, et al.
Published: (2024)
Multimodal LLMs under Pairwise Modalities
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
CHEEM: Continual Learning by Reuse, New, Adapt and Skip -- A Hierarchical Exploration-Exploitation Approach
by: Savadikar, Chinmay, et al.
Published: (2023)
by: Savadikar, Chinmay, et al.
Published: (2023)
LLM-I: LLMs are Naturally Interleaved Multimodal Creators
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
ManiPose: Manifold-Constrained Multi-Hypothesis 3D Human Pose Estimation
by: Rommel, Cédric, et al.
Published: (2023)
by: Rommel, Cédric, et al.
Published: (2023)
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
by: Gidaris, Spyros, et al.
Published: (2023)
by: Gidaris, Spyros, et al.
Published: (2023)
Catalogue Grounded Multimodal Attribution for Museum Video under Resource and Regulatory Constraints
by: Nanang, Minsak, et al.
Published: (2026)
by: Nanang, Minsak, et al.
Published: (2026)
How Visual Representations Map to Language Feature Space in Multimodal LLMs
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
GRIT-LP: Graph Transformer with Long-Range Skip Connection and Partitioned Spatial Graphs for Accurate Ice Layer Thickness Prediction
by: Liu, Zesheng, et al.
Published: (2025)
by: Liu, Zesheng, et al.
Published: (2025)
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
by: Hartman, Max, et al.
Published: (2025)
by: Hartman, Max, et al.
Published: (2025)
Development of Skip Connection in Deep Neural Networks for Computer Vision and Medical Image Analysis: A Survey
by: Xu, Guoping, et al.
Published: (2024)
by: Xu, Guoping, et al.
Published: (2024)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
by: Ye, Wenqian, et al.
Published: (2024)
by: Ye, Wenqian, et al.
Published: (2024)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
by: Neuhaus, Yannic, et al.
Published: (2026)
by: Neuhaus, Yannic, et al.
Published: (2026)
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2025)
Multimodal Deep Learning for Diabetic Foot Ulcer Staging Using Integrated RGB and Thermal Imaging
by: Mermer, Gulengul, et al.
Published: (2026)
by: Mermer, Gulengul, et al.
Published: (2026)
S$^{2}$-DMs:Skip-Step Diffusion Models
by: Wang, Yixuan, et al.
Published: (2024)
by: Wang, Yixuan, et al.
Published: (2024)
Boost-and-Skip: A Simple Guidance-Free Diffusion for Minority Generation
by: Um, Soobin, et al.
Published: (2025)
by: Um, Soobin, et al.
Published: (2025)
Lumos : Empowering Multimodal LLMs with Scene Text Recognition
by: Shenoy, Ashish, et al.
Published: (2024)
by: Shenoy, Ashish, et al.
Published: (2024)
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
by: Cheng, Jintao, et al.
Published: (2025)
by: Cheng, Jintao, et al.
Published: (2025)
Farm-LightSeek: An Edge-centric Multimodal Agricultural IoT Data Analytics Framework with Lightweight LLMs
by: Jiang, Dawen, et al.
Published: (2025)
by: Jiang, Dawen, et al.
Published: (2025)
Rethinking Image Skip Connections in StyleGAN2
by: Park, Seung, et al.
Published: (2024)
by: Park, Seung, et al.
Published: (2024)
SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
by: Hosseini, Parsa, et al.
Published: (2025)
by: Hosseini, Parsa, et al.
Published: (2025)
Similar Items
-
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024) -
Learning to Steer: Input-dependent Steering for Multimodal LLMs
by: Parekh, Jayneel, et al.
Published: (2025) -
A Concept-Based Explainability Framework for Large Multimodal Models
by: Parekh, Jayneel, et al.
Published: (2024) -
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
by: Vallaeys, Théophane, et al.
Published: (2024) -
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
by: Khayatan, Pegah, et al.
Published: (2025)