Enregistré dans:
| Auteurs principaux: | Song, Guanghui, Liao, Dongping, Zhao, Yiren, Ye, Kejiang, Xu, Cheng-zhong, Gao, Xitong |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2506.13541 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
BridgeNet: A Unified Multimodal Framework for Bridging 2D and 3D Industrial Anomaly Detection
par: Xiang, An, et autres
Publié: (2025)
par: Xiang, An, et autres
Publié: (2025)
Optimised Grouped-Query Attention Mechanism for Transformers
par: Chen, Yuang, et autres
Publié: (2024)
par: Chen, Yuang, et autres
Publié: (2024)
FLIP: Towards Comprehensive and Reliable Evaluation of Federated Prompt Learning
par: Liao, Dongping, et autres
Publié: (2025)
par: Liao, Dongping, et autres
Publié: (2025)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
par: Wu, Jingfeng, et autres
Publié: (2025)
par: Wu, Jingfeng, et autres
Publié: (2025)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
par: Wang, An, et autres
Publié: (2024)
par: Wang, An, et autres
Publié: (2024)
Offline Map Matching Based on Localization Error Distribution Modeling
par: Xu, Ruilin, et autres
Publié: (2025)
par: Xu, Ruilin, et autres
Publié: (2025)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
par: Li, Cheng, et autres
Publié: (2025)
par: Li, Cheng, et autres
Publié: (2025)
Refining Salience-Aware Sparse Fine-Tuning Strategies for Language Models
par: Liu, Xinxin, et autres
Publié: (2024)
par: Liu, Xinxin, et autres
Publié: (2024)
Mixture of Heterogeneous Grouped Experts for Language Modeling
par: Ma, Zhicheng, et autres
Publié: (2026)
par: Ma, Zhicheng, et autres
Publié: (2026)
PiKV: KV Cache Management System for Mixture of Experts
par: Liu, Dong, et autres
Publié: (2025)
par: Liu, Dong, et autres
Publié: (2025)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
par: Zhao, Guoliang, et autres
Publié: (2025)
par: Zhao, Guoliang, et autres
Publié: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
par: Zheng, Wanyi, et autres
Publié: (2025)
par: Zheng, Wanyi, et autres
Publié: (2025)
TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks
par: Shen, Hanzhang, et autres
Publié: (2026)
par: Shen, Hanzhang, et autres
Publié: (2026)
Unlocking the Global Synergies in Low-Rank Adapters
par: Zhang, Zixi, et autres
Publié: (2024)
par: Zhang, Zixi, et autres
Publié: (2024)
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
par: Xiong, Yi, et autres
Publié: (2024)
par: Xiong, Yi, et autres
Publié: (2024)
TokenPure: Watermark Removal through Tokenized Appearance and Structural Guidance
par: Yang, Pei, et autres
Publié: (2025)
par: Yang, Pei, et autres
Publié: (2025)
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
par: Shen, Yiqun, et autres
Publié: (2025)
par: Shen, Yiqun, et autres
Publié: (2025)
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
par: Gao, Yizhao, et autres
Publié: (2026)
par: Gao, Yizhao, et autres
Publié: (2026)
Topology Controls the Phase Separation Dynamics of Multicomponent Fluid Mixtures
par: Rennick, Michael, et autres
Publié: (2025)
par: Rennick, Michael, et autres
Publié: (2025)
Scaling Laws For Mixed Quantization
par: Cao, Zeyu, et autres
Publié: (2024)
par: Cao, Zeyu, et autres
Publié: (2024)
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
par: Liao, Fangshuo, et autres
Publié: (2025)
par: Liao, Fangshuo, et autres
Publié: (2025)
A Time Series is Worth Five Experts: Heterogeneous Mixture of Experts for Traffic Flow Prediction
par: Wang, Guangyu, et autres
Publié: (2024)
par: Wang, Guangyu, et autres
Publié: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: He, Yiyuan, et autres
Publié: (2025)
par: He, Yiyuan, et autres
Publié: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: Yiyuan He, et autres
Publié: (2026)
par: Yiyuan He, et autres
Publié: (2026)
GroupedMixer: An Entropy Model with Group-wise Token-Mixers for Learned Image Compression
par: Li, Daxin, et autres
Publié: (2024)
par: Li, Daxin, et autres
Publié: (2024)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
par: Go, Seokjin, et autres
Publié: (2025)
par: Go, Seokjin, et autres
Publié: (2025)
Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers
par: Zhao, Xin, et autres
Publié: (2025)
par: Zhao, Xin, et autres
Publié: (2025)
Generalizing GNNs with Tokenized Mixture of Experts
par: Guo, Xiaoguang, et autres
Publié: (2026)
par: Guo, Xiaoguang, et autres
Publié: (2026)
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
par: Zhao, Yiren, et autres
Publié: (2026)
par: Zhao, Yiren, et autres
Publié: (2026)
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
par: Jia, Haojie, et autres
Publié: (2025)
par: Jia, Haojie, et autres
Publié: (2025)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
par: Zhang, Zeliang, et autres
Publié: (2024)
par: Zhang, Zeliang, et autres
Publié: (2024)
AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models
par: Zeng, Zihao, et autres
Publié: (2024)
par: Zeng, Zihao, et autres
Publié: (2024)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
par: Li, Chong, et autres
Publié: (2025)
par: Li, Chong, et autres
Publié: (2025)
Optimizing Mixture of Block Attention
par: Xiao, Guangxuan, et autres
Publié: (2025)
par: Xiao, Guangxuan, et autres
Publié: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
par: Liao, Junhan, et autres
Publié: (2025)
par: Liao, Junhan, et autres
Publié: (2025)
OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
par: Gao, Yuting, et autres
Publié: (2025)
par: Gao, Yuting, et autres
Publié: (2025)
Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts
par: Zhang, Xue, et autres
Publié: (2025)
par: Zhang, Xue, et autres
Publié: (2025)
KV Shifting Attention Enhances Language Modeling
par: Xu, Mingyu, et autres
Publié: (2024)
par: Xu, Mingyu, et autres
Publié: (2024)
GRA: Detecting Oriented Objects through Group-wise Rotating and Attention
par: Wang, Jiangshan, et autres
Publié: (2024)
par: Wang, Jiangshan, et autres
Publié: (2024)
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
par: Pu, Yifan, et autres
Publié: (2024)
par: Pu, Yifan, et autres
Publié: (2024)
Documents similaires
-
BridgeNet: A Unified Multimodal Framework for Bridging 2D and 3D Industrial Anomaly Detection
par: Xiang, An, et autres
Publié: (2025) -
Optimised Grouped-Query Attention Mechanism for Transformers
par: Chen, Yuang, et autres
Publié: (2024) -
FLIP: Towards Comprehensive and Reliable Evaluation of Federated Prompt Learning
par: Liao, Dongping, et autres
Publié: (2025) -
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
par: Wu, Jingfeng, et autres
Publié: (2025) -
HMoE: Heterogeneous Mixture of Experts for Language Modeling
par: Wang, An, et autres
Publié: (2024)