Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Liang, Zhengyang, Liang, Meiyu, Huang, Wei, Li, Yawen, Xue, Zhe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning
di: Gao, Jun, et al.
Pubblicazione: (2024)
di: Gao, Jun, et al.
Pubblicazione: (2024)
Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models
di: Wu, Qiong, et al.
Pubblicazione: (2024)
di: Wu, Qiong, et al.
Pubblicazione: (2024)
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation
di: Zhu, Xiaofei, et al.
Pubblicazione: (2024)
di: Zhu, Xiaofei, et al.
Pubblicazione: (2024)
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
di: Zhang, Bo, et al.
Pubblicazione: (2024)
di: Zhang, Bo, et al.
Pubblicazione: (2024)
A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification
di: Calbucura, Nicolas, et al.
Pubblicazione: (2025)
di: Calbucura, Nicolas, et al.
Pubblicazione: (2025)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
di: Chen, Junyi, et al.
Pubblicazione: (2023)
di: Chen, Junyi, et al.
Pubblicazione: (2023)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
di: Wu, Zichen, et al.
Pubblicazione: (2024)
di: Wu, Zichen, et al.
Pubblicazione: (2024)
Mutual Information-based Representations Disentanglement for Unaligned Multimodal Language Sequences
di: Qian, Fan, et al.
Pubblicazione: (2024)
di: Qian, Fan, et al.
Pubblicazione: (2024)
Dependency Structure Augmented Contextual Scoping Framework for Multimodal Aspect-Based Sentiment Analysis
di: Liu, Hao, et al.
Pubblicazione: (2025)
di: Liu, Hao, et al.
Pubblicazione: (2025)
TCAN: Text-oriented Cross Attention Network for Multimodal Sentiment Analysis
di: Quan, Weize, et al.
Pubblicazione: (2024)
di: Quan, Weize, et al.
Pubblicazione: (2024)
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
di: Wei, Jingxuan, et al.
Pubblicazione: (2025)
di: Wei, Jingxuan, et al.
Pubblicazione: (2025)
Medical MLLM is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language Models
di: Huang, Xijie, et al.
Pubblicazione: (2024)
di: Huang, Xijie, et al.
Pubblicazione: (2024)
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model
di: Hu, Anwen, et al.
Pubblicazione: (2023)
di: Hu, Anwen, et al.
Pubblicazione: (2023)
Knowledge-Guided Dynamic Modality Attention Fusion Framework for Multimodal Sentiment Analysis
di: Feng, Xinyu, et al.
Pubblicazione: (2024)
di: Feng, Xinyu, et al.
Pubblicazione: (2024)
MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
di: Wang, Bingbing, et al.
Pubblicazione: (2025)
di: Wang, Bingbing, et al.
Pubblicazione: (2025)
Multimodal Sentiment Analysis Based on Causal Reasoning
di: Chen, Fuhai, et al.
Pubblicazione: (2024)
di: Chen, Fuhai, et al.
Pubblicazione: (2024)
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
di: Li, Zhu, et al.
Pubblicazione: (2025)
di: Li, Zhu, et al.
Pubblicazione: (2025)
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
di: Liu, Peipei, et al.
Pubblicazione: (2023)
di: Liu, Peipei, et al.
Pubblicazione: (2023)
IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity Alignment
di: Su, Taoyu, et al.
Pubblicazione: (2024)
di: Su, Taoyu, et al.
Pubblicazione: (2024)
Verifying Cross-modal Entity Consistency in News using Vision-language Models
di: Tahmasebi, Sahar, et al.
Pubblicazione: (2025)
di: Tahmasebi, Sahar, et al.
Pubblicazione: (2025)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
di: Gan, Chengguang, et al.
Pubblicazione: (2025)
di: Gan, Chengguang, et al.
Pubblicazione: (2025)
Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
di: Khac, Phúc H. Le, et al.
Pubblicazione: (2024)
di: Khac, Phúc H. Le, et al.
Pubblicazione: (2024)
Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
di: Wang, Bing, et al.
Pubblicazione: (2025)
di: Wang, Bing, et al.
Pubblicazione: (2025)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
di: Tang, Hao, et al.
Pubblicazione: (2025)
di: Tang, Hao, et al.
Pubblicazione: (2025)
MMPKUBase: A Comprehensive and High-quality Chinese Multi-modal Knowledge Graph
di: Yi, Xuan, et al.
Pubblicazione: (2024)
di: Yi, Xuan, et al.
Pubblicazione: (2024)
LoginMEA: Local-to-Global Interaction Network for Multi-modal Entity Alignment
di: Su, Taoyu, et al.
Pubblicazione: (2024)
di: Su, Taoyu, et al.
Pubblicazione: (2024)
Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
di: Meng, Chunlei, et al.
Pubblicazione: (2026)
di: Meng, Chunlei, et al.
Pubblicazione: (2026)
ChartEditor: A Reinforcement Learning Framework for Robust Chart Editing
di: Chen, Liangyu, et al.
Pubblicazione: (2025)
di: Chen, Liangyu, et al.
Pubblicazione: (2025)
Multi-MLLM Knowledge Distillation for Out-of-Context News Detection
di: Gu, Yimeng, et al.
Pubblicazione: (2025)
di: Gu, Yimeng, et al.
Pubblicazione: (2025)
Veagle: Advancements in Multimodal Representation Learning
di: Chawla, Rajat, et al.
Pubblicazione: (2024)
di: Chawla, Rajat, et al.
Pubblicazione: (2024)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
di: Zhang, Lei, et al.
Pubblicazione: (2025)
di: Zhang, Lei, et al.
Pubblicazione: (2025)
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction
di: Zhao, Yuan, et al.
Pubblicazione: (2024)
di: Zhao, Yuan, et al.
Pubblicazione: (2024)
MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
di: Zhang, Hanlei, et al.
Pubblicazione: (2024)
di: Zhang, Hanlei, et al.
Pubblicazione: (2024)
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
di: Cheng, Fenghua, et al.
Pubblicazione: (2025)
di: Cheng, Fenghua, et al.
Pubblicazione: (2025)
SoMeLVLM: A Large Vision Language Model for Social Media Processing
di: Zhang, Xinnong, et al.
Pubblicazione: (2024)
di: Zhang, Xinnong, et al.
Pubblicazione: (2024)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
di: Cui, Shiyao, et al.
Pubblicazione: (2025)
di: Cui, Shiyao, et al.
Pubblicazione: (2025)
Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
di: Niu, Fuqiang, et al.
Pubblicazione: (2024)
di: Niu, Fuqiang, et al.
Pubblicazione: (2024)
Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs
di: Mo, Wentao, et al.
Pubblicazione: (2026)
di: Mo, Wentao, et al.
Pubblicazione: (2026)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
di: Shah, Siddhant Bikram, et al.
Pubblicazione: (2024)
di: Shah, Siddhant Bikram, et al.
Pubblicazione: (2024)
Documenti analoghi
-
AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning
di: Gao, Jun, et al.
Pubblicazione: (2024) -
Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models
di: Wu, Qiong, et al.
Pubblicazione: (2024) -
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation
di: Zhu, Xiaofei, et al.
Pubblicazione: (2024) -
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
di: Zhang, Bo, et al.
Pubblicazione: (2024) -
A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification
di: Calbucura, Nicolas, et al.
Pubblicazione: (2025)