Geometry-Preserving Aggregation for Mixture-of-Experts Embedding Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Kachuee, Sajjad, Sharifkhani, Mohammad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Large Language Models with Zero-Shot Adjustable Acceleration
di: Kachuee, Sajjad, et al.
Pubblicazione: (2025)
di: Kachuee, Sajjad, et al.
Pubblicazione: (2025)
Latency Adjustable Transformer Encoder for Language Understanding
di: Kachuee, Sajjad, et al.
Pubblicazione: (2022)
di: Kachuee, Sajjad, et al.
Pubblicazione: (2022)
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
di: He, Haoze, et al.
Pubblicazione: (2026)
di: He, Haoze, et al.
Pubblicazione: (2026)
Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free
di: Li, Ziyue, et al.
Pubblicazione: (2024)
di: Li, Ziyue, et al.
Pubblicazione: (2024)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
di: Ahrac, Sagi, et al.
Pubblicazione: (2026)
di: Ahrac, Sagi, et al.
Pubblicazione: (2026)
Beyond instruction-conditioning, MoTE: Mixture of Task Experts for Multi-task Embedding Models
di: Romero, Miguel, et al.
Pubblicazione: (2025)
di: Romero, Miguel, et al.
Pubblicazione: (2025)
Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
di: Belenki, Lior, et al.
Pubblicazione: (2025)
di: Belenki, Lior, et al.
Pubblicazione: (2025)
Mixture of Lookup Experts
di: Jie, Shibo, et al.
Pubblicazione: (2025)
di: Jie, Shibo, et al.
Pubblicazione: (2025)
Upcycling Large Language Models into Mixture of Experts
di: He, Ethan, et al.
Pubblicazione: (2024)
di: He, Ethan, et al.
Pubblicazione: (2024)
Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings
di: Hallee, Logan, et al.
Pubblicazione: (2024)
di: Hallee, Logan, et al.
Pubblicazione: (2024)
Bayesian Mixture of Experts For Large Language Models
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models
di: Yoon, Youngsik, et al.
Pubblicazione: (2026)
di: Yoon, Youngsik, et al.
Pubblicazione: (2026)
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
di: Zheng, Chen, et al.
Pubblicazione: (2025)
di: Zheng, Chen, et al.
Pubblicazione: (2025)
Not Eliminate but Aggregate: Post-Hoc Control over Mixture-of-Experts to Address Shortcut Shifts in Natural Language Understanding
di: Honda, Ukyo, et al.
Pubblicazione: (2024)
di: Honda, Ukyo, et al.
Pubblicazione: (2024)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
di: Huang, Wei, et al.
Pubblicazione: (2024)
di: Huang, Wei, et al.
Pubblicazione: (2024)
A Survey on Mixture of Experts in Large Language Models
di: Cai, Weilin, et al.
Pubblicazione: (2024)
di: Cai, Weilin, et al.
Pubblicazione: (2024)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
di: Liew, Seng Pei, et al.
Pubblicazione: (2025)
di: Liew, Seng Pei, et al.
Pubblicazione: (2025)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
di: Wang, An, et al.
Pubblicazione: (2024)
di: Wang, An, et al.
Pubblicazione: (2024)
Binary-Integer-Programming Based Algorithm for Expert Load Balancing in Mixture-of-Experts Models
di: Sun, Yuan
Pubblicazione: (2025)
di: Sun, Yuan
Pubblicazione: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
di: Takashiro, Shota, et al.
Pubblicazione: (2026)
di: Takashiro, Shota, et al.
Pubblicazione: (2026)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
di: Lv, Ang, et al.
Pubblicazione: (2025)
di: Lv, Ang, et al.
Pubblicazione: (2025)
A Closer Look into Mixture-of-Experts in Large Language Models
di: Lo, Ka Man, et al.
Pubblicazione: (2024)
di: Lo, Ka Man, et al.
Pubblicazione: (2024)
Mixture of Experts Made Intrinsically Interpretable
di: Yang, Xingyi, et al.
Pubblicazione: (2025)
di: Yang, Xingyi, et al.
Pubblicazione: (2025)
Maximum Score Routing For Mixture-of-Experts
di: Dong, Bowen, et al.
Pubblicazione: (2025)
di: Dong, Bowen, et al.
Pubblicazione: (2025)
Mixture-of-Experts as Soft Clustering: A Dual Jacobian-PCA Spectral Geometry Perspective
di: Liu, Feilong
Pubblicazione: (2026)
di: Liu, Feilong
Pubblicazione: (2026)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs
di: Sel, Bilgehan, et al.
Pubblicazione: (2024)
di: Sel, Bilgehan, et al.
Pubblicazione: (2024)
Mixture of LoRA Experts
di: Wu, Xun, et al.
Pubblicazione: (2024)
di: Wu, Xun, et al.
Pubblicazione: (2024)
MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language Models
di: Tastan, Nurbek, et al.
Pubblicazione: (2026)
di: Tastan, Nurbek, et al.
Pubblicazione: (2026)
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
di: Wang, Haoxiang, et al.
Pubblicazione: (2024)
di: Wang, Haoxiang, et al.
Pubblicazione: (2024)
MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Models
di: Liu, Zehua, et al.
Pubblicazione: (2025)
di: Liu, Zehua, et al.
Pubblicazione: (2025)
Beyond Benchmarks: Understanding Mixture-of-Experts Models through Internal Mechanisms
di: Ying, Jiahao, et al.
Pubblicazione: (2025)
di: Ying, Jiahao, et al.
Pubblicazione: (2025)
Slicing and Dicing: Configuring Optimal Mixtures of Experts
di: Li, Margaret, et al.
Pubblicazione: (2026)
di: Li, Margaret, et al.
Pubblicazione: (2026)
MoFE: Mixture of Frozen Experts Architecture
di: Seo, Jean, et al.
Pubblicazione: (2025)
di: Seo, Jean, et al.
Pubblicazione: (2025)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
di: Muzio, Alexandre, et al.
Pubblicazione: (2024)
di: Muzio, Alexandre, et al.
Pubblicazione: (2024)
Shared Global and Local Geometry of Language Model Embeddings
di: Lee, Andrew, et al.
Pubblicazione: (2025)
di: Lee, Andrew, et al.
Pubblicazione: (2025)
Mixture of Heterogeneous Grouped Experts for Language Modeling
di: Ma, Zhicheng, et al.
Pubblicazione: (2026)
di: Ma, Zhicheng, et al.
Pubblicazione: (2026)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
di: Liu, Yang, et al.
Pubblicazione: (2026)
di: Liu, Yang, et al.
Pubblicazione: (2026)
Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
di: Dong, Zican, et al.
Pubblicazione: (2025)
di: Dong, Zican, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Efficient Large Language Models with Zero-Shot Adjustable Acceleration
di: Kachuee, Sajjad, et al.
Pubblicazione: (2025) -
Latency Adjustable Transformer Encoder for Language Understanding
di: Kachuee, Sajjad, et al.
Pubblicazione: (2022) -
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
di: He, Haoze, et al.
Pubblicazione: (2026) -
Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free
di: Li, Ziyue, et al.
Pubblicazione: (2024) -
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
di: Ahrac, Sagi, et al.
Pubblicazione: (2026)