Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Weihao, Wu, Qiong, Lin, Wenhao, Zhou, Yiyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph
von: Jiang, Yutao, et al.
Veröffentlicht: (2025)
von: Jiang, Yutao, et al.
Veröffentlicht: (2025)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Grounded Chain-of-Thought for Multimodal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
von: An, Wenbin, et al.
Veröffentlicht: (2024)
von: An, Wenbin, et al.
Veröffentlicht: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
Improving Gloss-free Sign Language Translation by Reducing Representation Density
von: Ye, Jinhui, et al.
Veröffentlicht: (2024)
von: Ye, Jinhui, et al.
Veröffentlicht: (2024)
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models
von: Guo, Yansong, et al.
Veröffentlicht: (2026)
von: Guo, Yansong, et al.
Veröffentlicht: (2026)
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
von: Bin, Yi, et al.
Veröffentlicht: (2024)
von: Bin, Yi, et al.
Veröffentlicht: (2024)
WordArt Designer API: User-Driven Artistic Typography Synthesis with Large Language Models on ModelScope
von: He, Jun-Yan, et al.
Veröffentlicht: (2024)
von: He, Jun-Yan, et al.
Veröffentlicht: (2024)
ChronusOmni: Improving Time Awareness of Omni Large Language Models
von: Chen, Yijing, et al.
Veröffentlicht: (2025)
von: Chen, Yijing, et al.
Veröffentlicht: (2025)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
von: Ding, Peng, et al.
Veröffentlicht: (2024)
von: Ding, Peng, et al.
Veröffentlicht: (2024)
Towards Training-free Multimodal Hate Localisation with Large Language Models
von: Sun, Yueming, et al.
Veröffentlicht: (2026)
von: Sun, Yueming, et al.
Veröffentlicht: (2026)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
von: Fazli, Mehrdad, et al.
Veröffentlicht: (2025)
von: Fazli, Mehrdad, et al.
Veröffentlicht: (2025)
Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
von: Hu, Xiaowan, et al.
Veröffentlicht: (2024)
von: Hu, Xiaowan, et al.
Veröffentlicht: (2024)
Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
von: Caffagni, Davide, et al.
Veröffentlicht: (2025)
Visual Grounding with Multi-modal Conditional Adaptation
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling
von: Lan, Jing, et al.
Veröffentlicht: (2026)
von: Lan, Jing, et al.
Veröffentlicht: (2026)
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model
von: Cuong, Dinh Viet, et al.
Veröffentlicht: (2025)
von: Cuong, Dinh Viet, et al.
Veröffentlicht: (2025)
Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs
von: Mo, Wentao, et al.
Veröffentlicht: (2026)
von: Mo, Wentao, et al.
Veröffentlicht: (2026)
EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model
von: Li, Deng, et al.
Veröffentlicht: (2024)
von: Li, Deng, et al.
Veröffentlicht: (2024)
Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
von: Zhan, Wengyi, et al.
Veröffentlicht: (2025)
von: Zhan, Wengyi, et al.
Veröffentlicht: (2025)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
Ask Questions with Double Hints: Visual Question Generation with Answer-awareness and Region-reference
von: Shen, Kai, et al.
Veröffentlicht: (2024)
von: Shen, Kai, et al.
Veröffentlicht: (2024)
MLANet: Multi-Level Attention Network with Sub-instruction for Continuous Vision-and-Language Navigation
von: He, Zongtao, et al.
Veröffentlicht: (2023)
von: He, Zongtao, et al.
Veröffentlicht: (2023)
CalliffusionV2: Personalized Natural Calligraphy Generation with Flexible Multi-modal Control
von: Liao, Qisheng, et al.
Veröffentlicht: (2024)
von: Liao, Qisheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
von: Wu, Qiong, et al.
Veröffentlicht: (2024) -
Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2024) -
What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph
von: Jiang, Yutao, et al.
Veröffentlicht: (2025) -
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
von: Wang, Xiao, et al.
Veröffentlicht: (2025) -
Grounded Chain-of-Thought for Multimodal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2025)