Mixture of Lookup Key-Value Experts
Fuente:
arXiv
Saved in:
| Main Author: | Wang, Zongcheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixture of Lookup Experts
by: Jie, Shibo, et al.
Published: (2025)
by: Jie, Shibo, et al.
Published: (2025)
LOOKAT: Lookup-Optimized Key-Attention for Memory-Efficient Transformers
by: Karmore, Aryan
Published: (2026)
by: Karmore, Aryan
Published: (2026)
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
by: Park, Sejik
Published: (2024)
by: Park, Sejik
Published: (2024)
Generalizing GNNs with Tokenized Mixture of Experts
by: Guo, Xiaoguang, et al.
Published: (2026)
by: Guo, Xiaoguang, et al.
Published: (2026)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
by: Nguyen, Dung V., et al.
Published: (2025)
by: Nguyen, Dung V., et al.
Published: (2025)
$μ$-Parametrization for Mixture of Experts
by: Małaśnicki, Jan, et al.
Published: (2025)
by: Małaśnicki, Jan, et al.
Published: (2025)
Path-Constrained Mixture-of-Experts
by: Gu, Zijin, et al.
Published: (2026)
by: Gu, Zijin, et al.
Published: (2026)
Mixture of Diverse Size Experts
by: Sun, Manxi, et al.
Published: (2024)
by: Sun, Manxi, et al.
Published: (2024)
On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks
by: Wang, Mingze, et al.
Published: (2025)
by: Wang, Mingze, et al.
Published: (2025)
L$^3$: Large Lookup Layers
by: Tseng, Albert, et al.
Published: (2026)
by: Tseng, Albert, et al.
Published: (2026)
Self-Augmented Mixture-of-Experts for QoS Prediction
by: Cai, Kecheng, et al.
Published: (2026)
by: Cai, Kecheng, et al.
Published: (2026)
Mixture of Raytraced Experts
by: Perin, Andrea, et al.
Published: (2025)
by: Perin, Andrea, et al.
Published: (2025)
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
by: Chowdhury, Mohammed Nowaz Rabbani, et al.
Published: (2024)
by: Chowdhury, Mohammed Nowaz Rabbani, et al.
Published: (2024)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
by: Gao, Yuting, et al.
Published: (2025)
by: Gao, Yuting, et al.
Published: (2025)
Mixture of Experts in a Mixture of RL settings
by: Willi, Timon, et al.
Published: (2024)
by: Willi, Timon, et al.
Published: (2024)
Robustness of Mixtures of Experts to Feature Noise
by: Sun, Dong, et al.
Published: (2026)
by: Sun, Dong, et al.
Published: (2026)
Temporally Extended Mixture-of-Experts Models
by: Shen, Zeyu, et al.
Published: (2026)
by: Shen, Zeyu, et al.
Published: (2026)
Hyperparameter Transfer with Mixture-of-Expert Layers
by: Jiang, Tianze, et al.
Published: (2026)
by: Jiang, Tianze, et al.
Published: (2026)
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026)
by: Madan, Vivan, et al.
Published: (2026)
Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuning
by: Eo, Sugyeong, et al.
Published: (2025)
by: Eo, Sugyeong, et al.
Published: (2025)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Maximum Score Routing For Mixture-of-Experts
by: Dong, Bowen, et al.
Published: (2025)
by: Dong, Bowen, et al.
Published: (2025)
Dynamic Expert Sharing: Decoupling Memory from Parallelism in Mixture-of-Experts Diffusion LLMs
by: Chen, Hao Mark, et al.
Published: (2026)
by: Chen, Hao Mark, et al.
Published: (2026)
Mixture of Experts in Large Language Models
by: Zhang, Danyang, et al.
Published: (2025)
by: Zhang, Danyang, et al.
Published: (2025)
$ϕ$-Balancing for Mixture-of-Experts Training
by: Chen, Lizhang, et al.
Published: (2026)
by: Chen, Lizhang, et al.
Published: (2026)
Multi-Head Mixture-of-Experts
by: Wu, Xun, et al.
Published: (2024)
by: Wu, Xun, et al.
Published: (2024)
LookupFFN: Making Transformers Compute-lite for CPU inference
by: Zeng, Zhanpeng, et al.
Published: (2024)
by: Zeng, Zhanpeng, et al.
Published: (2024)
Generalization Error Analysis for Sparse Mixture-of-Experts: A Preliminary Study
by: Zhao, Jinze, et al.
Published: (2024)
by: Zhao, Jinze, et al.
Published: (2024)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Sparsity and Superposition in Mixture of Experts
by: Chaudhari, Marmik, et al.
Published: (2025)
by: Chaudhari, Marmik, et al.
Published: (2025)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
by: He, Yifei, et al.
Published: (2025)
by: He, Yifei, et al.
Published: (2025)
Mixture of Concept Bottleneck Experts
by: De Santis, Francesco, et al.
Published: (2026)
by: De Santis, Francesco, et al.
Published: (2026)
Buffer Overflow in Mixture of Experts
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Mixture of A Million Experts
by: He, Xu Owen
Published: (2024)
by: He, Xu Owen
Published: (2024)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
by: Li, Lujun, et al.
Published: (2025)
by: Li, Lujun, et al.
Published: (2025)
On Linear Mode Connectivity of Mixture-of-Experts Architectures
by: Tran, Viet-Hoang, et al.
Published: (2025)
by: Tran, Viet-Hoang, et al.
Published: (2025)
Convergence Rates for Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
Dynamic Mixture-of-Experts for Incremental Graph Learning
by: Kong, Lecheng, et al.
Published: (2025)
by: Kong, Lecheng, et al.
Published: (2025)
Similar Items
-
Mixture of Lookup Experts
by: Jie, Shibo, et al.
Published: (2025) -
LOOKAT: Lookup-Optimized Key-Attention for Memory-Efficient Transformers
by: Karmore, Aryan
Published: (2026) -
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
by: Park, Sejik
Published: (2024) -
Generalizing GNNs with Tokenized Mixture of Experts
by: Guo, Xiaoguang, et al.
Published: (2026) -
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
by: Nguyen, Dung V., et al.
Published: (2025)