Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Song, Guanghui, Liao, Dongping, Zhao, Yiren, Ye, Kejiang, Xu, Cheng-zhong, Gao, Xitong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BridgeNet: A Unified Multimodal Framework for Bridging 2D and 3D Industrial Anomaly Detection
di: Xiang, An, et al.
Pubblicazione: (2025)
di: Xiang, An, et al.
Pubblicazione: (2025)
Optimised Grouped-Query Attention Mechanism for Transformers
di: Chen, Yuang, et al.
Pubblicazione: (2024)
di: Chen, Yuang, et al.
Pubblicazione: (2024)
FLIP: Towards Comprehensive and Reliable Evaluation of Federated Prompt Learning
di: Liao, Dongping, et al.
Pubblicazione: (2025)
di: Liao, Dongping, et al.
Pubblicazione: (2025)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
di: Wang, An, et al.
Pubblicazione: (2024)
di: Wang, An, et al.
Pubblicazione: (2024)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
di: Li, Cheng, et al.
Pubblicazione: (2025)
di: Li, Cheng, et al.
Pubblicazione: (2025)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
Offline Map Matching Based on Localization Error Distribution Modeling
di: Xu, Ruilin, et al.
Pubblicazione: (2025)
di: Xu, Ruilin, et al.
Pubblicazione: (2025)
Mixture of Heterogeneous Grouped Experts for Language Modeling
di: Ma, Zhicheng, et al.
Pubblicazione: (2026)
di: Ma, Zhicheng, et al.
Pubblicazione: (2026)
PiKV: KV Cache Management System for Mixture of Experts
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
Refining Salience-Aware Sparse Fine-Tuning Strategies for Language Models
di: Liu, Xinxin, et al.
Pubblicazione: (2024)
di: Liu, Xinxin, et al.
Pubblicazione: (2024)
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
di: Xiong, Yi, et al.
Pubblicazione: (2024)
di: Xiong, Yi, et al.
Pubblicazione: (2024)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
di: Zhao, Guoliang, et al.
Pubblicazione: (2025)
di: Zhao, Guoliang, et al.
Pubblicazione: (2025)
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
di: Shen, Yiqun, et al.
Pubblicazione: (2025)
di: Shen, Yiqun, et al.
Pubblicazione: (2025)
TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks
di: Shen, Hanzhang, et al.
Pubblicazione: (2026)
di: Shen, Hanzhang, et al.
Pubblicazione: (2026)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
di: Gao, Yizhao, et al.
Pubblicazione: (2026)
di: Gao, Yizhao, et al.
Pubblicazione: (2026)
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
di: Liao, Fangshuo, et al.
Pubblicazione: (2025)
di: Liao, Fangshuo, et al.
Pubblicazione: (2025)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
di: Go, Seokjin, et al.
Pubblicazione: (2025)
di: Go, Seokjin, et al.
Pubblicazione: (2025)
TokenPure: Watermark Removal through Tokenized Appearance and Structural Guidance
di: Yang, Pei, et al.
Pubblicazione: (2025)
di: Yang, Pei, et al.
Pubblicazione: (2025)
A Time Series is Worth Five Experts: Heterogeneous Mixture of Experts for Traffic Flow Prediction
di: Wang, Guangyu, et al.
Pubblicazione: (2024)
di: Wang, Guangyu, et al.
Pubblicazione: (2024)
GroupedMixer: An Entropy Model with Group-wise Token-Mixers for Learned Image Compression
di: Li, Daxin, et al.
Pubblicazione: (2024)
di: Li, Daxin, et al.
Pubblicazione: (2024)
Topology Controls the Phase Separation Dynamics of Multicomponent Fluid Mixtures
di: Rennick, Michael, et al.
Pubblicazione: (2025)
di: Rennick, Michael, et al.
Pubblicazione: (2025)
Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers
di: Zhao, Xin, et al.
Pubblicazione: (2025)
di: Zhao, Xin, et al.
Pubblicazione: (2025)
Generalizing GNNs with Tokenized Mixture of Experts
di: Guo, Xiaoguang, et al.
Pubblicazione: (2026)
di: Guo, Xiaoguang, et al.
Pubblicazione: (2026)
Unlocking the Global Synergies in Low-Rank Adapters
di: Zhang, Zixi, et al.
Pubblicazione: (2024)
di: Zhang, Zixi, et al.
Pubblicazione: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: Yiyuan He, et al.
Pubblicazione: (2026)
di: Yiyuan He, et al.
Pubblicazione: (2026)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
di: Zhang, Zeliang, et al.
Pubblicazione: (2024)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
di: Li, Chong, et al.
Pubblicazione: (2025)
di: Li, Chong, et al.
Pubblicazione: (2025)
AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models
di: Zeng, Zihao, et al.
Pubblicazione: (2024)
di: Zeng, Zihao, et al.
Pubblicazione: (2024)
Optimizing Mixture of Block Attention
di: Xiao, Guangxuan, et al.
Pubblicazione: (2025)
di: Xiao, Guangxuan, et al.
Pubblicazione: (2025)
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
di: Zhao, Yiren, et al.
Pubblicazione: (2026)
di: Zhao, Yiren, et al.
Pubblicazione: (2026)
Scaling Laws For Mixed Quantization
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts
di: Zhang, Xue, et al.
Pubblicazione: (2025)
di: Zhang, Xue, et al.
Pubblicazione: (2025)
OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
di: Gao, Yuting, et al.
Pubblicazione: (2025)
di: Gao, Yuting, et al.
Pubblicazione: (2025)
KV Shifting Attention Enhances Language Modeling
di: Xu, Mingyu, et al.
Pubblicazione: (2024)
di: Xu, Mingyu, et al.
Pubblicazione: (2024)
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
di: Jia, Haojie, et al.
Pubblicazione: (2025)
di: Jia, Haojie, et al.
Pubblicazione: (2025)
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
di: Zhou, Enshuai, et al.
Pubblicazione: (2026)
di: Zhou, Enshuai, et al.
Pubblicazione: (2026)
GRA: Detecting Oriented Objects through Group-wise Rotating and Attention
di: Wang, Jiangshan, et al.
Pubblicazione: (2024)
di: Wang, Jiangshan, et al.
Pubblicazione: (2024)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
di: Li, Junzhuo, et al.
Pubblicazione: (2026)
di: Li, Junzhuo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
BridgeNet: A Unified Multimodal Framework for Bridging 2D and 3D Industrial Anomaly Detection
di: Xiang, An, et al.
Pubblicazione: (2025) -
Optimised Grouped-Query Attention Mechanism for Transformers
di: Chen, Yuang, et al.
Pubblicazione: (2024) -
FLIP: Towards Comprehensive and Reliable Evaluation of Federated Prompt Learning
di: Liao, Dongping, et al.
Pubblicazione: (2025) -
HMoE: Heterogeneous Mixture of Experts for Language Modeling
di: Wang, An, et al.
Pubblicazione: (2024) -
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
di: Li, Cheng, et al.
Pubblicazione: (2025)