Guardado en:
| Autores principales: | Gim, In, Ma, Zhiyao, Lee, Seung-seob, Zhong, Lin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2510.24051 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Asynchronous LLM Function Calling
por: Gim, In, et al.
Publicado: (2024)
por: Gim, In, et al.
Publicado: (2024)
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
por: Gim, In, et al.
Publicado: (2023)
por: Gim, In, et al.
Publicado: (2023)
Cacheback: Speculative Decoding With Nothing But Cache
por: Ma, Zhiyao, et al.
Publicado: (2025)
por: Ma, Zhiyao, et al.
Publicado: (2025)
Serve Programs, Not Prompts
por: Gim, In, et al.
Publicado: (2025)
por: Gim, In, et al.
Publicado: (2025)
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
por: Li, Caihua, et al.
Publicado: (2024)
por: Li, Caihua, et al.
Publicado: (2024)
MappedTrace: Tracing Pointer Remotely with Compiler-generated Maps
por: Ma, Zhiyao, et al.
Publicado: (2025)
por: Ma, Zhiyao, et al.
Publicado: (2025)
Blindfold: Confidential Memory Management by Untrusted Operating System
por: Li, Caihua, et al.
Publicado: (2024)
por: Li, Caihua, et al.
Publicado: (2024)
WebPie: A Tiny Slice of Dependent Typing
por: Scholliers, Christophe
Publicado: (2024)
por: Scholliers, Christophe
Publicado: (2024)
NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
por: Tang, Yupeng, et al.
Publicado: (2023)
por: Tang, Yupeng, et al.
Publicado: (2023)
AccelGen: Heterogeneous SLO-Guaranteed High-Throughput LLM Inference Serving for Diverse Applications
por: Shen, Haiying, et al.
Publicado: (2025)
por: Shen, Haiying, et al.
Publicado: (2025)
TimelyLLM: Segmented LLM Serving System for Time-sensitive Robotic Applications
por: Ling, Neiwen, et al.
Publicado: (2024)
por: Ling, Neiwen, et al.
Publicado: (2024)
InstCache: A Predictive Cache for LLM Serving
por: Zou, Longwei, et al.
Publicado: (2024)
por: Zou, Longwei, et al.
Publicado: (2024)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
por: Li, Zikun, et al.
Publicado: (2025)
por: Li, Zikun, et al.
Publicado: (2025)
LAPIS: Language Model-Augmented Police Investigation System
por: Kim, Heedou, et al.
Publicado: (2024)
por: Kim, Heedou, et al.
Publicado: (2024)
Regulating Branch Parallelism in LLM Serving
por: Gandhi, Swapnil, et al.
Publicado: (2026)
por: Gandhi, Swapnil, et al.
Publicado: (2026)
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
por: Hong, Ke, et al.
Publicado: (2025)
por: Hong, Ke, et al.
Publicado: (2025)
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
por: Wang, Ying, et al.
Publicado: (2025)
por: Wang, Ying, et al.
Publicado: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
por: Lin, Yujun, et al.
Publicado: (2024)
por: Lin, Yujun, et al.
Publicado: (2024)
Learned Best-Effort LLM Serving
por: Jha, Siddharth, et al.
Publicado: (2024)
por: Jha, Siddharth, et al.
Publicado: (2024)
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
por: Li, Haoyang, et al.
Publicado: (2025)
por: Li, Haoyang, et al.
Publicado: (2025)
Continuous Semantic Caching for Low-Cost LLM Serving
por: Atalar, Baran, et al.
Publicado: (2026)
por: Atalar, Baran, et al.
Publicado: (2026)
LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoring
por: Jang, Jinhee, et al.
Publicado: (2025)
por: Jang, Jinhee, et al.
Publicado: (2025)
MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
por: Li, Yufei, et al.
Publicado: (2025)
por: Li, Yufei, et al.
Publicado: (2025)
Dual-Pool Token-Budget Routing for Cost-Efficient and Reliable LLM Serving
por: Liu, Xunzhuo, et al.
Publicado: (2026)
por: Liu, Xunzhuo, et al.
Publicado: (2026)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
por: Jin, Yibo, et al.
Publicado: (2024)
por: Jin, Yibo, et al.
Publicado: (2024)
Identifying Emerging Concepts in Large Corpora
por: Ma, Sibo, et al.
Publicado: (2025)
por: Ma, Sibo, et al.
Publicado: (2025)
Taming the Titans: A Survey of Efficient LLM Inference Serving
por: Zhen, Ranran, et al.
Publicado: (2025)
por: Zhen, Ranran, et al.
Publicado: (2025)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
por: Yang, Shang, et al.
Publicado: (2025)
por: Yang, Shang, et al.
Publicado: (2025)
Dual-Scale World Models for LLM Agents Towards Hard-Exploration Problems
por: Kim, Minsoo, et al.
Publicado: (2025)
por: Kim, Minsoo, et al.
Publicado: (2025)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
por: Oliaro, Gabriele, et al.
Publicado: (2024)
por: Oliaro, Gabriele, et al.
Publicado: (2024)
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
por: Agullo, Ferran, et al.
Publicado: (2025)
por: Agullo, Ferran, et al.
Publicado: (2025)
Culinary Class Wars: Evaluating LLMs using ASH in Cuisine Transfer Task
por: Lee, Hoonick, et al.
Publicado: (2024)
por: Lee, Hoonick, et al.
Publicado: (2024)
Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation
por: Ren, Zhiyao, et al.
Publicado: (2026)
por: Ren, Zhiyao, et al.
Publicado: (2026)
Smoothie-Qwen: Post-Hoc Smoothing to Reduce Language Bias in Multilingual LLMs
por: Ji, SeungWon, et al.
Publicado: (2025)
por: Ji, SeungWon, et al.
Publicado: (2025)
Dataverse: Open-Source ETL (Extract, Transform, Load) Pipeline for Large Language Models
por: Park, Hyunbyung, et al.
Publicado: (2024)
por: Park, Hyunbyung, et al.
Publicado: (2024)
Cognitive Policy-Driven LLM for Diagnosis and Intervention of Cognitive Distortions in Emotional Support Conversation
por: Zhong, Lin, et al.
Publicado: (2026)
por: Zhong, Lin, et al.
Publicado: (2026)
HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving
por: Kumar, Avinash, et al.
Publicado: (2025)
por: Kumar, Avinash, et al.
Publicado: (2025)
SOMA: Efficient Multi-turn LLM Serving via Small Language Model
por: Cheng, Xueqi, et al.
Publicado: (2026)
por: Cheng, Xueqi, et al.
Publicado: (2026)
Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
por: Xu, Yechen, et al.
Publicado: (2024)
por: Xu, Yechen, et al.
Publicado: (2024)
Ejemplares similares
-
Asynchronous LLM Function Calling
por: Gim, In, et al.
Publicado: (2024) -
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
por: Gim, In, et al.
Publicado: (2023) -
Cacheback: Speculative Decoding With Nothing But Cache
por: Ma, Zhiyao, et al.
Publicado: (2025) -
Serve Programs, Not Prompts
por: Gim, In, et al.
Publicado: (2025) -
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
por: Li, Caihua, et al.
Publicado: (2024)