TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ren, Feng, Qin, Ruoyu, Ma, Teng, Cai, Shangming, Liu, Zheng, Lei, Chao, Zhu, Dejiang, Yang, Ke, Li, Zheming, Cui, Jialei, Huang, Weixiao, Zhao, Yikai, Zhang, Yineng, Wu, Hao, Gao, Xiang, Fu, Yuhao, Jiang, Jinlei, Wu, Yongwei, Zhang, Mingxing |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
par: Qin, Ruoyu, et autres
Publié: (2024)
par: Qin, Ruoyu, et autres
Publié: (2024)
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
par: Chen, Shaoyuan, et autres
Publié: (2024)
par: Chen, Shaoyuan, et autres
Publié: (2024)
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
par: Qin, Ruoyu, et autres
Publié: (2025)
par: Qin, Ruoyu, et autres
Publié: (2025)
Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
par: Li, Zongze, et autres
Publié: (2026)
par: Li, Zongze, et autres
Publié: (2026)
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
par: Qin, Ruoyu, et autres
Publié: (2026)
par: Qin, Ruoyu, et autres
Publié: (2026)
Efficient Graph-Based Approximate Nearest Neighbor Search Achieving: Low Latency Without Throughput Loss
par: Luo, Jingjia, et autres
Publié: (2025)
par: Luo, Jingjia, et autres
Publié: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
par: Dong, Xianzhe, et autres
Publié: (2025)
par: Dong, Xianzhe, et autres
Publié: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
par: Ye, Zihao, et autres
Publié: (2025)
par: Ye, Zihao, et autres
Publié: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
par: Ruan, Chaoyi, et autres
Publié: (2025)
par: Ruan, Chaoyi, et autres
Publié: (2025)
Physical parameter regression from black hole images via a multiscale adaptive neural network
par: Wei, Jialei, et autres
Publié: (2025)
par: Wei, Jialei, et autres
Publié: (2025)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
par: Jin, Yibo, et autres
Publié: (2024)
par: Jin, Yibo, et autres
Publié: (2024)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
par: Hu, Xiannan, et autres
Publié: (2025)
par: Hu, Xiannan, et autres
Publié: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
par: Zhong, Yinmin, et autres
Publié: (2024)
par: Zhong, Yinmin, et autres
Publié: (2024)
Trinity: Disaggregating Vector Search from Prefill-Decode Disaggregation in LLM Serving
par: Liu, Yi, et autres
Publié: (2025)
par: Liu, Yi, et autres
Publié: (2025)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
par: Qiu, Haoran, et autres
Publié: (2025)
par: Qiu, Haoran, et autres
Publié: (2025)
SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
par: Tian, Yuhao, et autres
Publié: (2025)
par: Tian, Yuhao, et autres
Publié: (2025)
TrEnv-X: Transparently Share Serverless Execution Environments Across Different Functions and Nodes
par: Huang, Jialiang, et autres
Publié: (2025)
par: Huang, Jialiang, et autres
Publié: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
par: Zhu, Ruidong, et autres
Publié: (2025)
par: Zhu, Ruidong, et autres
Publié: (2025)
Efficiently Serving Large Multimodal Models Using EPD Disaggregation
par: Singh, Gursimran, et autres
Publié: (2024)
par: Singh, Gursimran, et autres
Publié: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: He, Yiyuan, et autres
Publié: (2025)
par: He, Yiyuan, et autres
Publié: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: Yiyuan He, et autres
Publié: (2026)
par: Yiyuan He, et autres
Publié: (2026)
End Khovanov homology and exotic Lagrangian planes
par: Teng, Yikai
Publié: (2025)
par: Teng, Yikai
Publié: (2025)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
par: Shi, Tianyao, et autres
Publié: (2024)
par: Shi, Tianyao, et autres
Publié: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
par: Kumar, Satyam, et autres
Publié: (2026)
par: Kumar, Satyam, et autres
Publié: (2026)
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
par: Hong, Ke, et autres
Publié: (2025)
par: Hong, Ke, et autres
Publié: (2025)
Uncertainty Quantification and Flow Dynamics in Rotating Detonation Engines
par: Kumar, Vinay, et autres
Publié: (2025)
par: Kumar, Vinay, et autres
Publié: (2025)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
par: Wu, Hanjiang, et autres
Publié: (2026)
par: Wu, Hanjiang, et autres
Publié: (2026)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
par: Bai, Fan, et autres
Publié: (2026)
par: Bai, Fan, et autres
Publié: (2026)
TENT5/FAM46: An Enigmatic Family of Secretory Tuners
par: Daniel Lacidogna, et autres
Publié: (2025)
par: Daniel Lacidogna, et autres
Publié: (2025)
MCN-CL: Multimodal Cross-Attention Network and Contrastive Learning for Multimodal Emotion Recognition
par: Li, Feng, et autres
Publié: (2025)
par: Li, Feng, et autres
Publié: (2025)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
par: Lai, Ruiqi, et autres
Publié: (2025)
par: Lai, Ruiqi, et autres
Publié: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
par: Wu, Yu, et autres
Publié: (2025)
par: Wu, Yu, et autres
Publié: (2025)
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
par: Ding, Jianru, et autres
Publié: (2026)
par: Ding, Jianru, et autres
Publié: (2026)
Efficient Multi-round LLM Inference over Disaggregated Serving
par: He, Wenhao, et autres
Publié: (2026)
par: He, Wenhao, et autres
Publié: (2026)
OBSERVATIONS ON TENT-USING IN THE CAROLLINE BAT RHINOPYLLA PUMILIO IN SOUTHEASTERN BRAZIL
par: Zortéa, Marlon
Publié: (1995)
par: Zortéa, Marlon
Publié: (1995)
The Effects of Circadian Rhythms and Exercise Preconditioning on Cardiac Troponin T Levels Following Graded Exercise
par: Jinlei Nie, et autres
Publié: (2025)
par: Jinlei Nie, et autres
Publié: (2025)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
par: Li, Jiaxi, et autres
Publié: (2025)
par: Li, Jiaxi, et autres
Publié: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
par: Shi, Xiaoxiang, et autres
Publié: (2025)
par: Shi, Xiaoxiang, et autres
Publié: (2025)
Documents similaires
-
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
par: Qin, Ruoyu, et autres
Publié: (2024) -
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
par: Chen, Shaoyuan, et autres
Publié: (2024) -
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
par: Qin, Ruoyu, et autres
Publié: (2025) -
Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
par: Li, Zongze, et autres
Publié: (2026) -
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
par: Qin, Ruoyu, et autres
Publié: (2026)