$\mathcal{V}isi\mathcal{P}runer$: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Yingqi, Zhao, Anhao, Fu, Jinlan, Tong, Junlong, Su, Hui, Pan, Yijie, Zhang, Wei, Shen, Xiaoyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position Encoding
by: Tong, Junlong, et al.
Published: (2025)
by: Tong, Junlong, et al.
Published: (2025)
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
by: Fan, Yingqi, et al.
Published: (2026)
by: Fan, Yingqi, et al.
Published: (2026)
StreamingThinker: Large Language Models Can Think While Reading
by: Tong, Junlong, et al.
Published: (2025)
by: Tong, Junlong, et al.
Published: (2025)
SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling
by: Zhao, Anhao, et al.
Published: (2025)
by: Zhao, Anhao, et al.
Published: (2025)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
by: Zhao, Anhao, et al.
Published: (2026)
by: Zhao, Anhao, et al.
Published: (2026)
Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices
by: Lin, Junyan, et al.
Published: (2025)
by: Lin, Junyan, et al.
Published: (2025)
Unveiling In-Context Learning: A Coordinate System to Understand Its Working Mechanism
by: Zhao, Anhao, et al.
Published: (2024)
by: Zhao, Anhao, et al.
Published: (2024)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
by: Liu, Wenjie, et al.
Published: (2026)
by: Liu, Wenjie, et al.
Published: (2026)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
by: Fu, Jinlan, et al.
Published: (2025)
by: Fu, Jinlan, et al.
Published: (2025)
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
by: Wu, Hao, et al.
Published: (2026)
by: Wu, Hao, et al.
Published: (2026)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
by: Fu, Jinlan, et al.
Published: (2025)
by: Fu, Jinlan, et al.
Published: (2025)
Multimodal Language Models See Better When They Look Shallower
by: Chen, Haoran, et al.
Published: (2025)
by: Chen, Haoran, et al.
Published: (2025)
Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts
by: Zhu, Zhihao, et al.
Published: (2026)
by: Zhu, Zhihao, et al.
Published: (2026)
Rethinking the Role of LLMs in Time Series Forecasting
by: Qiu, Xin, et al.
Published: (2026)
by: Qiu, Xin, et al.
Published: (2026)
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
by: Lin, Junyan, et al.
Published: (2026)
by: Lin, Junyan, et al.
Published: (2026)
On The Axioms Of $\mathcal{M},\mathcal{N}$-Adhesive Categories
by: Castelnovo, Davide, et al.
Published: (2024)
by: Castelnovo, Davide, et al.
Published: (2024)
Context Guided Transformer Entropy Modeling for Video Compression
by: Tong, Junlong, et al.
Published: (2025)
by: Tong, Junlong, et al.
Published: (2025)
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
by: Tong, Bingkui, et al.
Published: (2025)
by: Tong, Bingkui, et al.
Published: (2025)
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
by: Ethayarajh, Kawin, et al.
Published: (2021)
by: Ethayarajh, Kawin, et al.
Published: (2021)
On-Policy Supervised Fine-Tuning for Efficient Reasoning
by: Zhao, Anhao, et al.
Published: (2026)
by: Zhao, Anhao, et al.
Published: (2026)
The Few Govern the Many:Unveiling Few-Layer Dominance for Time Series Models
by: Qiu, Xin, et al.
Published: (2025)
by: Qiu, Xin, et al.
Published: (2025)
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Solving unification in the description logic $\mathcal{FL}_\bot$
by: Morawska, Barbara, et al.
Published: (2024)
by: Morawska, Barbara, et al.
Published: (2024)
Text Detoxification in isiXhosa and Yorùbá: A Cross-Lingual Machine Learning Approach for Low-Resource African Languages
by: Agbeyangi, Abayomi O.
Published: (2026)
by: Agbeyangi, Abayomi O.
Published: (2026)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
by: Xiao, Yang, et al.
Published: (2023)
by: Xiao, Yang, et al.
Published: (2023)
Automatically assessing oral narratives of Afrikaans and isiXhosa children
by: Louw, Retief, et al.
Published: (2025)
by: Louw, Retief, et al.
Published: (2025)
Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality Rebalance
by: Qu, Mingcheng, et al.
Published: (2025)
by: Qu, Mingcheng, et al.
Published: (2025)
FILO -- automated unification in $\mathcal{FL}_0$
by: Morawska, Barbara, et al.
Published: (2025)
by: Morawska, Barbara, et al.
Published: (2025)
Hint-before-Solving Prompting: Guiding LLMs to Effectively Utilize Encoded Knowledge
by: Fu, Jinlan, et al.
Published: (2024)
by: Fu, Jinlan, et al.
Published: (2024)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
by: Verma, Gaurav, et al.
Published: (2024)
by: Verma, Gaurav, et al.
Published: (2024)
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
by: Lv, Zhengyao, et al.
Published: (2025)
by: Lv, Zhengyao, et al.
Published: (2025)
HiP-AD: Hierarchical and Multi-Granularity Planning with Deformable Attention for Autonomous Driving in a Single Decoder
by: Tang, Yingqi, et al.
Published: (2025)
by: Tang, Yingqi, et al.
Published: (2025)
UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking
by: Wu, Hao, et al.
Published: (2026)
by: Wu, Hao, et al.
Published: (2026)
Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency
by: Liang, Jiafeng, et al.
Published: (2025)
by: Liang, Jiafeng, et al.
Published: (2025)
GSIFN: A Graph-Structured and Interlaced-Masked Multimodal Transformer-based Fusion Network for Multimodal Sentiment Analysis
by: Jin, Yijie
Published: (2024)
by: Jin, Yijie
Published: (2024)
Complex Valued Deep Operator Network (DeepONet) $[\mathcal{G}]$ for Three Dimensional Maxwell's Equations: $\mathcal{G} \in \mathbb{C}^{m \times n}$
by: Jiang, Qile, et al.
Published: (2024)
by: Jiang, Qile, et al.
Published: (2024)
BabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context
by: Matzopoulos, Alexis, et al.
Published: (2025)
by: Matzopoulos, Alexis, et al.
Published: (2025)
Dynamic Multimodal Sentiment Analysis: Leveraging Cross-Modal Attention for Enabled Classification
by: Lee, Hui, et al.
Published: (2025)
by: Lee, Hui, et al.
Published: (2025)
Feature-based analysis of oral narratives from Afrikaans and isiXhosa children
by: Sharratt, Emma, et al.
Published: (2025)
by: Sharratt, Emma, et al.
Published: (2025)
Similar Items
-
LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position Encoding
by: Tong, Junlong, et al.
Published: (2025) -
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
by: Fan, Yingqi, et al.
Published: (2026) -
StreamingThinker: Large Language Models Can Think While Reading
by: Tong, Junlong, et al.
Published: (2025) -
SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling
by: Zhao, Anhao, et al.
Published: (2025) -
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
by: Zhao, Anhao, et al.
Published: (2026)