SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Woo, Sunghyeon, Seo, Ahreum, Lee, Jaegwang, Kil, Jaeeun, Seo, Hanbae, Kim, Joonghoon, Park, Baeseong, Kwon, Se Jung, Lee, Dongsoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
ICaRus: Identical Cache Reuse for Efficient Multi Model Inference
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward Propagation
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2024)
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2024)
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
von: Lee, Jeongtae, et al.
Veröffentlicht: (2026)
von: Lee, Jeongtae, et al.
Veröffentlicht: (2026)
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
von: Bae, Jeongin, et al.
Veröffentlicht: (2026)
von: Bae, Jeongin, et al.
Veröffentlicht: (2026)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
von: Park, Gunho, et al.
Veröffentlicht: (2022)
von: Park, Gunho, et al.
Veröffentlicht: (2022)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
von: Park, Gunho, et al.
Veröffentlicht: (2025)
von: Park, Gunho, et al.
Veröffentlicht: (2025)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
von: Park, Gunho, et al.
Veröffentlicht: (2025)
von: Park, Gunho, et al.
Veröffentlicht: (2025)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
von: Lee, Joonhyung, et al.
Veröffentlicht: (2024)
An Inquiry into Datacenter TCO for LLM Inference with FP8
von: Kim, Jiwoo, et al.
Veröffentlicht: (2025)
von: Kim, Jiwoo, et al.
Veröffentlicht: (2025)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2023)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2023)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
von: Park, Gunho, et al.
Veröffentlicht: (2025)
von: Park, Gunho, et al.
Veröffentlicht: (2025)
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
von: Heo, Jung Hwan, et al.
Veröffentlicht: (2023)
von: Heo, Jung Hwan, et al.
Veröffentlicht: (2023)
High‐performance graphitic nanoplatelets & high‐density polyethylene nanocomposites
von: Seo Jeong Yoon, et al.
Veröffentlicht: (2024)
von: Seo Jeong Yoon, et al.
Veröffentlicht: (2024)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
SCANNER: Knowledge-Enhanced Approach for Robust Multi-modal Named Entity Recognition of Unseen Entities
von: Ok, Hyunjong, et al.
Veröffentlicht: (2024)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2024)
DIAL: Distribution-Informed Adaptive Learning of Multi-Task Constraints for Safety-Critical Systems
von: Yoo, Se-Wook, et al.
Veröffentlicht: (2025)
von: Yoo, Se-Wook, et al.
Veröffentlicht: (2025)
Autonomous Algorithm for Training Autonomous Vehicles with Minimal Human Intervention
von: Lee, Sang-Hyun, et al.
Veröffentlicht: (2024)
von: Lee, Sang-Hyun, et al.
Veröffentlicht: (2024)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
Physics in Next-token Prediction
von: An, Hongjun, et al.
Veröffentlicht: (2024)
von: An, Hongjun, et al.
Veröffentlicht: (2024)
DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
von: Kwon, Sangwoo, et al.
Veröffentlicht: (2025)
von: Kwon, Sangwoo, et al.
Veröffentlicht: (2025)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2024)
von: Lee, Jung Hyun, et al.
Veröffentlicht: (2024)
Imagination-Augmented Hierarchical Reinforcement Learning for Safe and Interactive Autonomous Driving in Urban Environments
von: Lee, Sang-Hyun, et al.
Veröffentlicht: (2023)
von: Lee, Sang-Hyun, et al.
Veröffentlicht: (2023)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
Trinity: Disaggregating Vector Search from Prefill-Decode Disaggregation in LLM Serving
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
Who Moves and Where Do They Go? Generational Differences in Urban‐to‐Rural Migration in South Korea
von: Hyeongju Seo, et al.
Veröffentlicht: (2026)
von: Hyeongju Seo, et al.
Veröffentlicht: (2026)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
von: Wang, Shaoyu, et al.
Veröffentlicht: (2025)
von: Wang, Shaoyu, et al.
Veröffentlicht: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
MT-RAIG: Novel Benchmark and Evaluation Framework for Retrieval-Augmented Insight Generation over Multiple Tables
von: Seo, Kwangwook, et al.
Veröffentlicht: (2025)
von: Seo, Kwangwook, et al.
Veröffentlicht: (2025)
Current-driven dynamics of antiferromagnetic domain-wall skyrmions
von: Kim, Wooyon, et al.
Veröffentlicht: (2025)
von: Kim, Wooyon, et al.
Veröffentlicht: (2025)
Traversability-aware Adaptive Optimization for Path Planning and Control in Mountainous Terrain
von: Yoo, Se-Wook, et al.
Veröffentlicht: (2024)
von: Yoo, Se-Wook, et al.
Veröffentlicht: (2024)
Self-Supervised Curriculum Generation for Autonomous Reinforcement Learning without Task-Specific Knowledge
von: Lee, Sang-Hyun, et al.
Veröffentlicht: (2023)
von: Lee, Sang-Hyun, et al.
Veröffentlicht: (2023)
Silicon‐Based Anodes for Sulfide Solid‐State Batteries: Failure Mechanisms and Multiscale Design Strategies
von: Murugesan Karuppaiah, et al.
Veröffentlicht: (2026)
von: Murugesan Karuppaiah, et al.
Veröffentlicht: (2026)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
von: Cho, Jaehong, et al.
Veröffentlicht: (2026)
von: Cho, Jaehong, et al.
Veröffentlicht: (2026)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
On the Convergence of Density-Based Predictive Control for Multi-Agent Non-Uniform Area Coverage
von: Seo, Sungjun, et al.
Veröffentlicht: (2025)
von: Seo, Sungjun, et al.
Veröffentlicht: (2025)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
von: Qin, Ruoyu, et al.
Veröffentlicht: (2024)
von: Qin, Ruoyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026) -
ICaRus: Identical Cache Reuse for Efficient Multi Model Inference
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026) -
DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward Propagation
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2024) -
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
von: Lee, Jeongtae, et al.
Veröffentlicht: (2026) -
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
von: Bae, Jeongin, et al.
Veröffentlicht: (2026)