MoExtend: Tuning New Experts for Modality and Task Extension
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Shanshan, Gao, Shanghua, Huang, Zhongzhan, Wen, Wushao, Zitnik, Marinka, Zhou, Pan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
von: Zhong, Shanshan, et al.
Veröffentlicht: (2023)
von: Zhong, Shanshan, et al.
Veröffentlicht: (2023)
A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models
von: Huang, Zhongzhan, et al.
Veröffentlicht: (2025)
von: Huang, Zhongzhan, et al.
Veröffentlicht: (2025)
ASR: Attention-alike Structural Re-parameterization
von: Zhong, Shanshan, et al.
Veröffentlicht: (2023)
von: Zhong, Shanshan, et al.
Veröffentlicht: (2023)
AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
Qworld: Question-Specific Evaluation Criteria for LLMs
von: Gao, Shanghua, et al.
Veröffentlicht: (2026)
von: Gao, Shanghua, et al.
Veröffentlicht: (2026)
Expert Pyramid Tuning: Efficient Parameter Fine-Tuning for Expertise-Driven Task Allocation
von: Zhang, Jia-Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Jia-Chen, et al.
Veröffentlicht: (2026)
Is Extending Modality The Right Path Towards Omni-Modality?
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
MoDE: CLIP Data Experts via Clustering
von: Ma, Jiawei, et al.
Veröffentlicht: (2024)
von: Ma, Jiawei, et al.
Veröffentlicht: (2024)
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
MoEController: Instruction-based Arbitrary Image Manipulation with Mixture-of-Expert Controllers
von: Li, Sijia, et al.
Veröffentlicht: (2023)
von: Li, Sijia, et al.
Veröffentlicht: (2023)
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion
von: Han, Feng, et al.
Veröffentlicht: (2026)
von: Han, Feng, et al.
Veröffentlicht: (2026)
Less is More: High-value Data Selection for Visual Instruction Tuning
von: Liu, Zikang, et al.
Veröffentlicht: (2024)
von: Liu, Zikang, et al.
Veröffentlicht: (2024)
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
Otter: A Multi-Modal Model with In-Context Instruction Tuning
von: Li, Bo, et al.
Veröffentlicht: (2023)
von: Li, Bo, et al.
Veröffentlicht: (2023)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
von: Huang, Qidong, et al.
Veröffentlicht: (2024)
von: Huang, Qidong, et al.
Veröffentlicht: (2024)
Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
MMGR: Multi-Modal Generative Reasoning
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
von: Cai, Zefan, et al.
Veröffentlicht: (2025)
MDTv2: Masked Diffusion Transformer is a Strong Image Synthesizer
von: Gao, Shanghua, et al.
Veröffentlicht: (2023)
von: Gao, Shanghua, et al.
Veröffentlicht: (2023)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
von: Chen, Zeren, et al.
Veröffentlicht: (2023)
von: Chen, Zeren, et al.
Veröffentlicht: (2023)
CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks
von: Suharitdamrong, Wish, et al.
Veröffentlicht: (2026)
von: Suharitdamrong, Wish, et al.
Veröffentlicht: (2026)
Head-wise Modality Specialization within MLLMs for Robust Fake News Detection under Missing Modality
von: Qian, Kai, et al.
Veröffentlicht: (2026)
von: Qian, Kai, et al.
Veröffentlicht: (2026)
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
von: Yang, Honglong, et al.
Veröffentlicht: (2025)
von: Yang, Honglong, et al.
Veröffentlicht: (2025)
Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection
von: Hossain, Md. Mithun, et al.
Veröffentlicht: (2025)
von: Hossain, Md. Mithun, et al.
Veröffentlicht: (2025)
LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
Multimodal Medical Code Tokenizer
von: Su, Xiaorui, et al.
Veröffentlicht: (2025)
von: Su, Xiaorui, et al.
Veröffentlicht: (2025)
ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2024)
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2024)
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
von: Zhang, Wenqiao, et al.
Veröffentlicht: (2024)
von: Zhang, Wenqiao, et al.
Veröffentlicht: (2024)
Learning Generalized Medical Image Representations through Image-Graph Contrastive Pretraining
von: Khanna, Sameer, et al.
Veröffentlicht: (2024)
von: Khanna, Sameer, et al.
Veröffentlicht: (2024)
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
von: Bi, Jinhe, et al.
Veröffentlicht: (2024)
von: Bi, Jinhe, et al.
Veröffentlicht: (2024)
Cross-Modal Adapter for Vision-Language Retrieval
von: Jiang, Haojun, et al.
Veröffentlicht: (2022)
von: Jiang, Haojun, et al.
Veröffentlicht: (2022)
Sandboxed Coding Agents are Competitive Omni-modal Task Solvers
von: Chen, Dongping, et al.
Veröffentlicht: (2026)
von: Chen, Dongping, et al.
Veröffentlicht: (2026)
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Vision-Language Models Create Cross-Modal Task Representations
von: Luo, Grace, et al.
Veröffentlicht: (2024)
von: Luo, Grace, et al.
Veröffentlicht: (2024)
Shared and Private Information Learning in Multimodal Sentiment Analysis with Deep Modal Alignment and Self-supervised Multi-Task Learning
von: Lai, Songning, et al.
Veröffentlicht: (2023)
von: Lai, Songning, et al.
Veröffentlicht: (2023)
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs
von: Bo, Zi-Hao, et al.
Veröffentlicht: (2026)
von: Bo, Zi-Hao, et al.
Veröffentlicht: (2026)
MoKus: Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization
von: Zhu, Chenyang, et al.
Veröffentlicht: (2026)
von: Zhu, Chenyang, et al.
Veröffentlicht: (2026)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
von: Xing, Long, et al.
Veröffentlicht: (2025)
von: Xing, Long, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
von: Zhong, Shanshan, et al.
Veröffentlicht: (2023) -
A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models
von: Huang, Zhongzhan, et al.
Veröffentlicht: (2025) -
ASR: Attention-alike Structural Re-parameterization
von: Zhong, Shanshan, et al.
Veröffentlicht: (2023) -
AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
von: Liu, Yifan, et al.
Veröffentlicht: (2025) -
Qworld: Question-Specific Evaluation Criteria for LLMs
von: Gao, Shanghua, et al.
Veröffentlicht: (2026)