DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Choraria, Moulik, Wu, Xinbo, Bhimaraju, Akhil, Sekhar, Nitesh, Wu, Yue, Zhang, Xu, Singhal, Prateek, Varshney, Lav R. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantically Grounded QFormer for Efficient Vision Language Understanding
by: Choraria, Moulik, et al.
Published: (2023)
by: Choraria, Moulik, et al.
Published: (2023)
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
by: Hartman, Max, et al.
Published: (2025)
by: Hartman, Max, et al.
Published: (2025)
Watermarking Discrete Diffusion Language Models
by: Bagchi, Avi, et al.
Published: (2025)
by: Bagchi, Avi, et al.
Published: (2025)
Dynamic Batching of Online Arrivals to Leverage Economies of Scale
by: Bhimaraju, Akhil, et al.
Published: (2023)
by: Bhimaraju, Akhil, et al.
Published: (2023)
Context-Gated Associative Retrieval: From Theory to Transformers
by: Choraria, Moulik, et al.
Published: (2026)
by: Choraria, Moulik, et al.
Published: (2026)
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
by: Hartman, Max, et al.
Published: (2026)
by: Hartman, Max, et al.
Published: (2026)
Transformer-based Causal Language Models Perform Clustering
by: Wu, Xinbo, et al.
Published: (2024)
by: Wu, Xinbo, et al.
Published: (2024)
A Meta-Learning Perspective on Transformers for Causal Language Modeling
by: Wu, Xinbo, et al.
Published: (2023)
by: Wu, Xinbo, et al.
Published: (2023)
Fractional Budget Allocation for Influence Maximization under General Marketing Strategies
by: Bhimaraju, Akhil, et al.
Published: (2024)
by: Bhimaraju, Akhil, et al.
Published: (2024)
Concealment of Intent: A Game-Theoretic Analysis
by: Wu, Xinbo, et al.
Published: (2025)
by: Wu, Xinbo, et al.
Published: (2025)
Verifier Threshold: An Efficient Test-Time Scaling Approach for Image Generation
by: Sundaresha, Vignesh, et al.
Published: (2025)
by: Sundaresha, Vignesh, et al.
Published: (2025)
A Theoretical Game of Attacks via Compositional Skills
by: Wu, Xinbo, et al.
Published: (2026)
by: Wu, Xinbo, et al.
Published: (2026)
SwitchCIT: Switching for Continual Instruction Tuning
by: Wu, Xinbo, et al.
Published: (2024)
by: Wu, Xinbo, et al.
Published: (2024)
Convex Distillation: Efficient Compression of Deep Networks via Convex Optimization
by: Varshney, Prateek, et al.
Published: (2024)
by: Varshney, Prateek, et al.
Published: (2024)
Hallucination Basins: A Dynamic Framework for Understanding and Controlling LLM Hallucinations
by: Cherukuri, Kalyan, et al.
Published: (2026)
by: Cherukuri, Kalyan, et al.
Published: (2026)
Fed-SB: A Silver Bullet for Extreme Communication Efficiency and Performance in (Private) Federated LoRA Fine-Tuning
by: Singhal, Raghav, et al.
Published: (2025)
by: Singhal, Raghav, et al.
Published: (2025)
Learning from One and Only One Shot
by: Yu, Haizi, et al.
Published: (2022)
by: Yu, Haizi, et al.
Published: (2022)
SparseJEPA: Sparse Representation Learning of Joint Embedding Predictive Architectures
by: Hartman, Max, et al.
Published: (2025)
by: Hartman, Max, et al.
Published: (2025)
Adaptively Bypassing Vision Transformer Blocks for Efficient Visual Tracking
by: Yang, Xiangyang, et al.
Published: (2024)
by: Yang, Xiangyang, et al.
Published: (2024)
Efficient Model-Agnostic Multi-Group Equivariant Networks
by: Baltaji, Razan, et al.
Published: (2023)
by: Baltaji, Razan, et al.
Published: (2023)
Do Multimodal Large Language Models Understand Welding?
by: Khvatskii, Grigorii, et al.
Published: (2025)
by: Khvatskii, Grigorii, et al.
Published: (2025)
Task-specific regularization loss towards model calibration for reliable lung cancer detection
by: Kalra, Mehar Prateek, et al.
Published: (2024)
by: Kalra, Mehar Prateek, et al.
Published: (2024)
NeRF-Insert: 3D Local Editing with Multimodal Control Signals
by: Sabat, Benet Oriol, et al.
Published: (2024)
by: Sabat, Benet Oriol, et al.
Published: (2024)
StableIdentity: Inserting Anybody into Anywhere at First Sight
by: Wang, Qinghe, et al.
Published: (2024)
by: Wang, Qinghe, et al.
Published: (2024)
MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
by: Cao, Yue, et al.
Published: (2024)
by: Cao, Yue, et al.
Published: (2024)
InverseMeetInsert: Robust Real Image Editing via Geometric Accumulation Inversion in Guided Diffusion Models
by: Zheng, Yan, et al.
Published: (2024)
by: Zheng, Yan, et al.
Published: (2024)
Indian Sign Language Detection for Real-Time Translation using Machine Learning
by: Singhal, Rajat, et al.
Published: (2025)
by: Singhal, Rajat, et al.
Published: (2025)
Semantic Compression with Information Lattice Learning
by: Yu, Haizi, et al.
Published: (2024)
by: Yu, Haizi, et al.
Published: (2024)
Containment Verification: AI Safety Guarantees Independent of Alignment
by: Moon, Royce, et al.
Published: (2026)
by: Moon, Royce, et al.
Published: (2026)
Online Reinforcement Learning with Passive Memory
by: Pattanaik, Anay, et al.
Published: (2024)
by: Pattanaik, Anay, et al.
Published: (2024)
SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass
by: Qian, Chen, et al.
Published: (2026)
by: Qian, Chen, et al.
Published: (2026)
PIXELS: Progressive Image Xemplar-based Editing with Latent Surgery
by: Biswas, Shristi Das, et al.
Published: (2025)
by: Biswas, Shristi Das, et al.
Published: (2025)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
by: Huang, Muye, et al.
Published: (2025)
by: Huang, Muye, et al.
Published: (2025)
Bypassing LLM Watermarks with Color-Aware Substitutions
by: Wu, Qilong, et al.
Published: (2024)
by: Wu, Qilong, et al.
Published: (2024)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
Quick Bypass Mechanism of Zero-Shot Diffusion-Based Image Restoration
by: Tai, Yu-Shan, et al.
Published: (2025)
by: Tai, Yu-Shan, et al.
Published: (2025)
An Information Theory of Compute-Optimal Size Scaling, Emergence, and Plateaus in Language Models
by: Nayak, Anuj K., et al.
Published: (2024)
by: Nayak, Anuj K., et al.
Published: (2024)
Structure-Guided Histopathology Synthesis via Dual-LoRA Diffusion
by: Xu, Xuan, et al.
Published: (2026)
by: Xu, Xuan, et al.
Published: (2026)
Similar Items
-
Semantically Grounded QFormer for Efficient Vision Language Understanding
by: Choraria, Moulik, et al.
Published: (2023) -
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
by: Hartman, Max, et al.
Published: (2025) -
Watermarking Discrete Diffusion Language Models
by: Bagchi, Avi, et al.
Published: (2025) -
Dynamic Batching of Online Arrivals to Leverage Economies of Scale
by: Bhimaraju, Akhil, et al.
Published: (2023) -
Context-Gated Associative Retrieval: From Theory to Transformers
by: Choraria, Moulik, et al.
Published: (2026)