Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yunxin, Jiang, Shenyuan, Hu, Baotian, Wang, Longyue, Zhong, Wanqi, Luo, Wenhan, Ma, Lin, Zhang, Min |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
di: Li, Yunxin, et al.
Pubblicazione: (2025)
di: Li, Yunxin, et al.
Pubblicazione: (2025)
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
di: Shi, Haoyuan, et al.
Pubblicazione: (2026)
di: Shi, Haoyuan, et al.
Pubblicazione: (2026)
UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
di: Chen, Yanzhe, et al.
Pubblicazione: (2025)
di: Chen, Yanzhe, et al.
Pubblicazione: (2025)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
di: Jiang, Jingjing, et al.
Pubblicazione: (2025)
di: Jiang, Jingjing, et al.
Pubblicazione: (2025)
A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
Multi-level Mixture of Experts for Multimodal Entity Linking
di: Hu, Zhiwei, et al.
Pubblicazione: (2025)
di: Hu, Zhiwei, et al.
Pubblicazione: (2025)
Mixture-of-Prompt-Experts for Multi-modal Semantic Understanding
di: Wu, Zichen, et al.
Pubblicazione: (2024)
di: Wu, Zichen, et al.
Pubblicazione: (2024)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
di: Chen, Shuang, et al.
Pubblicazione: (2026)
di: Chen, Shuang, et al.
Pubblicazione: (2026)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
di: Chen, Junyi, et al.
Pubblicazione: (2023)
di: Chen, Junyi, et al.
Pubblicazione: (2023)
Mixture of LoRA Experts
di: Wu, Xun, et al.
Pubblicazione: (2024)
di: Wu, Xun, et al.
Pubblicazione: (2024)
Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child Interaction
di: Shi, Weiyan, et al.
Pubblicazione: (2025)
di: Shi, Weiyan, et al.
Pubblicazione: (2025)
VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization
di: Li, Yunxin, et al.
Pubblicazione: (2025)
di: Li, Yunxin, et al.
Pubblicazione: (2025)
UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning
di: Bai, Hayes, et al.
Pubblicazione: (2026)
di: Bai, Hayes, et al.
Pubblicazione: (2026)
Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
di: Ma, Xueqi, et al.
Pubblicazione: (2025)
di: Ma, Xueqi, et al.
Pubblicazione: (2025)
Lost in Overlap: Exploring Logit-based Watermark Collision in LLMs
di: Luo, Yiyang, et al.
Pubblicazione: (2024)
di: Luo, Yiyang, et al.
Pubblicazione: (2024)
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
di: Jiang, Fan, et al.
Pubblicazione: (2026)
di: Jiang, Fan, et al.
Pubblicazione: (2026)
UniScene: Multi-Camera Unified Pre-training via 3D Scene Reconstruction for Autonomous Driving
di: Min, Chen, et al.
Pubblicazione: (2023)
di: Min, Chen, et al.
Pubblicazione: (2023)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
di: Sun, Hao, et al.
Pubblicazione: (2024)
di: Sun, Hao, et al.
Pubblicazione: (2024)
MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter
di: Liu, Zhiyuan, et al.
Pubblicazione: (2023)
di: Liu, Zhiyuan, et al.
Pubblicazione: (2023)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
di: Wen, Haokun, et al.
Pubblicazione: (2026)
di: Wen, Haokun, et al.
Pubblicazione: (2026)
UniMuMo: Unified Text, Music and Motion Generation
di: Yang, Han, et al.
Pubblicazione: (2024)
di: Yang, Han, et al.
Pubblicazione: (2024)
UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
di: Cheng, Zhi-Qi, et al.
Pubblicazione: (2024)
di: Cheng, Zhi-Qi, et al.
Pubblicazione: (2024)
Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP Supervision
di: Yin, Kangsheng, et al.
Pubblicazione: (2025)
di: Yin, Kangsheng, et al.
Pubblicazione: (2025)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
di: Li, Hebeizi, et al.
Pubblicazione: (2026)
di: Li, Hebeizi, et al.
Pubblicazione: (2026)
A Comprehensive Evaluation of GPT-4V on Knowledge-Intensive Visual Question Answering
di: Li, Yunxin, et al.
Pubblicazione: (2023)
di: Li, Yunxin, et al.
Pubblicazione: (2023)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
di: Wang, Xiang, et al.
Pubblicazione: (2025)
di: Wang, Xiang, et al.
Pubblicazione: (2025)
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
di: Zhang, Bo, et al.
Pubblicazione: (2024)
di: Zhang, Bo, et al.
Pubblicazione: (2024)
MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening
di: Shao, Yongqi, et al.
Pubblicazione: (2025)
di: Shao, Yongqi, et al.
Pubblicazione: (2025)
Scaling up Multimodal Pre-training for Sign Language Understanding
di: Zhou, Wengang, et al.
Pubblicazione: (2024)
di: Zhou, Wengang, et al.
Pubblicazione: (2024)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
di: Zhang, Jiaxu, et al.
Pubblicazione: (2025)
di: Zhang, Jiaxu, et al.
Pubblicazione: (2025)
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
di: Yan, Qianqi, et al.
Pubblicazione: (2026)
di: Yan, Qianqi, et al.
Pubblicazione: (2026)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
di: Jiang, Ruixiang, et al.
Pubblicazione: (2025)
di: Jiang, Ruixiang, et al.
Pubblicazione: (2025)
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
di: Rho, Kyeongha, et al.
Pubblicazione: (2025)
di: Rho, Kyeongha, et al.
Pubblicazione: (2025)
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
di: He, Xin, et al.
Pubblicazione: (2024)
di: He, Xin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
di: Li, Yunxin, et al.
Pubblicazione: (2024) -
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
di: Li, Yunxin, et al.
Pubblicazione: (2025) -
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
di: Shi, Haoyuan, et al.
Pubblicazione: (2026) -
UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
di: Liu, Zhenyu, et al.
Pubblicazione: (2025) -
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
di: Chen, Yanzhe, et al.
Pubblicazione: (2025)