Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Qiong, Ye, Weihao, Zhou, Yiyi, Sun, Xiaoshuai, Ji, Rongrong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Routing Experts: Learning to Route Dynamic Experts in Multi-modal Large Language Models
di: Wu, Qiong, et al.
Pubblicazione: (2024)
di: Wu, Qiong, et al.
Pubblicazione: (2024)
Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
di: Ye, Weihao, et al.
Pubblicazione: (2024)
di: Ye, Weihao, et al.
Pubblicazione: (2024)
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
di: Wu, Qiong, et al.
Pubblicazione: (2024)
di: Wu, Qiong, et al.
Pubblicazione: (2024)
Grounded Chain-of-Thought for Multimodal Large Language Models
di: Wu, Qiong, et al.
Pubblicazione: (2025)
di: Wu, Qiong, et al.
Pubblicazione: (2025)
MMAPS: End-to-End Multi-Grained Multi-Modal Attribute-Aware Product Summarization
di: Chen, Tao, et al.
Pubblicazione: (2023)
di: Chen, Tao, et al.
Pubblicazione: (2023)
AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning
di: Gao, Jun, et al.
Pubblicazione: (2024)
di: Gao, Jun, et al.
Pubblicazione: (2024)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
di: Wu, Zichen, et al.
Pubblicazione: (2024)
di: Wu, Zichen, et al.
Pubblicazione: (2024)
Collaborative Evolution: Multi-Round Learning Between Large and Small Language Models for Emergent Fake News Detection
di: Zhou, Ziyi, et al.
Pubblicazione: (2025)
di: Zhou, Ziyi, et al.
Pubblicazione: (2025)
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model
di: Hu, Anwen, et al.
Pubblicazione: (2023)
di: Hu, Anwen, et al.
Pubblicazione: (2023)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
di: Wang, Xiao, et al.
Pubblicazione: (2025)
di: Wang, Xiao, et al.
Pubblicazione: (2025)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
di: He, Zheqi, et al.
Pubblicazione: (2024)
di: He, Zheqi, et al.
Pubblicazione: (2024)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
di: Lu, Jinghui, et al.
Pubblicazione: (2025)
di: Lu, Jinghui, et al.
Pubblicazione: (2025)
IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity Alignment
di: Su, Taoyu, et al.
Pubblicazione: (2024)
di: Su, Taoyu, et al.
Pubblicazione: (2024)
EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
di: Gao, Lancheng, et al.
Pubblicazione: (2025)
di: Gao, Lancheng, et al.
Pubblicazione: (2025)
Multimodal Large Language Models for Medicine: A Comprehensive Survey
di: Ye, Jiarui, et al.
Pubblicazione: (2025)
di: Ye, Jiarui, et al.
Pubblicazione: (2025)
MMPKUBase: A Comprehensive and High-quality Chinese Multi-modal Knowledge Graph
di: Yi, Xuan, et al.
Pubblicazione: (2024)
di: Yi, Xuan, et al.
Pubblicazione: (2024)
LoginMEA: Local-to-Global Interaction Network for Multi-modal Entity Alignment
di: Su, Taoyu, et al.
Pubblicazione: (2024)
di: Su, Taoyu, et al.
Pubblicazione: (2024)
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
di: Chen, Lichang, et al.
Pubblicazione: (2024)
di: Chen, Lichang, et al.
Pubblicazione: (2024)
A Benchmark and Robustness Study of In-Context-Learning with Large Language Models in Music Entity Detection
di: Hachmeier, Simon, et al.
Pubblicazione: (2024)
di: Hachmeier, Simon, et al.
Pubblicazione: (2024)
SoMeLVLM: A Large Vision Language Model for Social Media Processing
di: Zhang, Xinnong, et al.
Pubblicazione: (2024)
di: Zhang, Xinnong, et al.
Pubblicazione: (2024)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
di: Jiang, Chaoya, et al.
Pubblicazione: (2024)
di: Jiang, Chaoya, et al.
Pubblicazione: (2024)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
di: Liang, Zhengyang, et al.
Pubblicazione: (2024)
di: Liang, Zhengyang, et al.
Pubblicazione: (2024)
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
di: Li, Zhu, et al.
Pubblicazione: (2025)
di: Li, Zhu, et al.
Pubblicazione: (2025)
ChartAdapter: Large Vision-Language Model for Chart Summarization
di: Xu, Peixin, et al.
Pubblicazione: (2024)
di: Xu, Peixin, et al.
Pubblicazione: (2024)
Large Language Models for Computer-Aided Design: A Survey
di: Zhang, Licheng, et al.
Pubblicazione: (2025)
di: Zhang, Licheng, et al.
Pubblicazione: (2025)
MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
di: Wu, Shih-Lun, et al.
Pubblicazione: (2025)
di: Wu, Shih-Lun, et al.
Pubblicazione: (2025)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
di: Yang, Danni, et al.
Pubblicazione: (2024)
di: Yang, Danni, et al.
Pubblicazione: (2024)
Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
di: Wang, Junyu, et al.
Pubblicazione: (2025)
di: Wang, Junyu, et al.
Pubblicazione: (2025)
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
di: Zhang, Bo, et al.
Pubblicazione: (2024)
di: Zhang, Bo, et al.
Pubblicazione: (2024)
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025)
di: Fazli, Mehrdad, et al.
Pubblicazione: (2025)
TCAN: Text-oriented Cross Attention Network for Multimodal Sentiment Analysis
di: Quan, Weize, et al.
Pubblicazione: (2024)
di: Quan, Weize, et al.
Pubblicazione: (2024)
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
di: Zhang, Hanlei, et al.
Pubblicazione: (2025)
di: Zhang, Hanlei, et al.
Pubblicazione: (2025)
Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing
di: Wu, Zichen, et al.
Pubblicazione: (2025)
di: Wu, Zichen, et al.
Pubblicazione: (2025)
SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
di: Jin, Zeyu, et al.
Pubblicazione: (2024)
di: Jin, Zeyu, et al.
Pubblicazione: (2024)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
di: Lu, Jinghui, et al.
Pubblicazione: (2024)
di: Lu, Jinghui, et al.
Pubblicazione: (2024)
Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
di: Wang, Bing, et al.
Pubblicazione: (2025)
di: Wang, Bing, et al.
Pubblicazione: (2025)
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
di: Wei, Jingxuan, et al.
Pubblicazione: (2025)
di: Wei, Jingxuan, et al.
Pubblicazione: (2025)
Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal Perspective
di: Su, Taoyu, et al.
Pubblicazione: (2025)
di: Su, Taoyu, et al.
Pubblicazione: (2025)
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
di: Liu, Shansong, et al.
Pubblicazione: (2023)
di: Liu, Shansong, et al.
Pubblicazione: (2023)
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
di: Yang, Danni, et al.
Pubblicazione: (2024)
di: Yang, Danni, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Routing Experts: Learning to Route Dynamic Experts in Multi-modal Large Language Models
di: Wu, Qiong, et al.
Pubblicazione: (2024) -
Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
di: Ye, Weihao, et al.
Pubblicazione: (2024) -
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
di: Wu, Qiong, et al.
Pubblicazione: (2024) -
Grounded Chain-of-Thought for Multimodal Large Language Models
di: Wu, Qiong, et al.
Pubblicazione: (2025) -
MMAPS: End-to-End Multi-Grained Multi-Modal Attribute-Aware Product Summarization
di: Chen, Tao, et al.
Pubblicazione: (2023)