Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kang, Junmo, Karlinsky, Leonid, Luo, Hongyin, Wang, Zhen, Hansen, Jacob, Glass, James, Cox, David, Panda, Rameswar, Feris, Rogerio, Ritter, Alan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Specialization: Uncovering Latent Expertise within Large Language Models
von: Kang, Junmo, et al.
Veröffentlicht: (2023)
von: Kang, Junmo, et al.
Veröffentlicht: (2023)
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
von: Hansen, Jacob, et al.
Veröffentlicht: (2025)
von: Hansen, Jacob, et al.
Veröffentlicht: (2025)
State-Space Large Audio Language Models
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024)
Listen, Think, and Understand
von: Gong, Yuan, et al.
Veröffentlicht: (2023)
von: Gong, Yuan, et al.
Veröffentlicht: (2023)
CAMELoT: Towards Large Language Models with Training-Free Consolidated Associative Memory
von: He, Zexue, et al.
Veröffentlicht: (2024)
von: He, Zexue, et al.
Veröffentlicht: (2024)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
Navigating the Labyrinth: Evaluating LLMs' Ability to Reason About Search Problems
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2024)
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2024)
$\textit{Trans-LoRA}$: towards data-free Transferable Parameter Efficient Finetuning
von: Wang, Runqian, et al.
Veröffentlicht: (2024)
von: Wang, Runqian, et al.
Veröffentlicht: (2024)
$\texttt{BATCLIP}$: Bimodal Online Test-Time Adaptation for CLIP
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2024)
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2024)
Towards Audio Token Compression in Large Audio Language Models
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2025)
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2025)
Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2025)
von: Borazjanizadeh, Nasim, et al.
Veröffentlicht: (2025)
Latent Implicit Visual Reasoning
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
Advancing Expert Specialization for Better MoE
von: Guo, Hongcan, et al.
Veröffentlicht: (2025)
von: Guo, Hongcan, et al.
Veröffentlicht: (2025)
Large Scale Generative AI Text Applied to Sports and Music
von: Baughman, Aaron, et al.
Veröffentlicht: (2024)
von: Baughman, Aaron, et al.
Veröffentlicht: (2024)
Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based Finetuning
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2025)
von: Raghavendra, Mohit, et al.
Veröffentlicht: (2025)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
von: Araujo, Edson, et al.
Veröffentlicht: (2025)
von: Araujo, Edson, et al.
Veröffentlicht: (2025)
LangNav: Language as a Perceptual Representation for Navigation
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
von: Pan, Bowen, et al.
Veröffentlicht: (2023)
SD-MoE: Spectral Decomposition for Effective Expert Specialization
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
Comparison Visual Instruction Tuning
von: Lin, Wei, et al.
Veröffentlicht: (2024)
von: Lin, Wei, et al.
Veröffentlicht: (2024)
Adaptive Memory Replay for Continual Learning
von: Smith, James Seale, et al.
Veröffentlicht: (2024)
von: Smith, James Seale, et al.
Veröffentlicht: (2024)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
Post-Trained MoE Can Skip Half Experts via Self-Distillation
von: Lv, Xingtai, et al.
Veröffentlicht: (2026)
von: Lv, Xingtai, et al.
Veröffentlicht: (2026)
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
Scattered Mixture-of-Experts Implementation
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs
von: Mirvakhabova, Leyla, et al.
Veröffentlicht: (2025)
von: Mirvakhabova, Leyla, et al.
Veröffentlicht: (2025)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
von: Huang, Irene, et al.
Veröffentlicht: (2024)
von: Huang, Irene, et al.
Veröffentlicht: (2024)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert Specialization
von: Hu, Rizhen, et al.
Veröffentlicht: (2026)
von: Hu, Rizhen, et al.
Veröffentlicht: (2026)
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs
von: Bo, Zi-Hao, et al.
Veröffentlicht: (2026)
von: Bo, Zi-Hao, et al.
Veröffentlicht: (2026)
DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs
von: Wang, Jing, et al.
Veröffentlicht: (2026)
von: Wang, Jing, et al.
Veröffentlicht: (2026)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
von: Xue, Leyang, et al.
Veröffentlicht: (2024)
von: Xue, Leyang, et al.
Veröffentlicht: (2024)
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models
von: Aghdam, Maryam Akhavan, et al.
Veröffentlicht: (2024)
von: Aghdam, Maryam Akhavan, et al.
Veröffentlicht: (2024)
Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
Teaching VLMs to Localize Specific Objects from In-context Examples
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
von: Doveh, Sivan, et al.
Veröffentlicht: (2024)
MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale
von: Falke, Tobias, et al.
Veröffentlicht: (2026)
von: Falke, Tobias, et al.
Veröffentlicht: (2026)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
von: Farhat, Yehya, et al.
Veröffentlicht: (2023)
von: Farhat, Yehya, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Self-Specialization: Uncovering Latent Expertise within Large Language Models
von: Kang, Junmo, et al.
Veröffentlicht: (2023) -
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
von: Hansen, Jacob, et al.
Veröffentlicht: (2025) -
State-Space Large Audio Language Models
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024) -
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2024) -
Listen, Think, and Understand
von: Gong, Yuan, et al.
Veröffentlicht: (2023)