An Empirical Study of LLM Serving in Confidential GPUs
Fuente:
Zenodo
Saved in:
| Main Authors: | Park, Eunseong, Xiong, Wenjie, Thans, Timo, Kumar Kalidasan, Vishnu, Hu, Qinghao |
|---|---|
| Format: | Recurso digital |
| Published: |
Zenodo
2026
|
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Performance of Confidential Computing GPUs
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
by: Ibarra, Antonio Martínez, et al.
Published: (2025)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
by: Zhu, Jianwei, et al.
Published: (2024)
by: Zhu, Jianwei, et al.
Published: (2024)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
by: Yao, Xiaozhe, et al.
Published: (2023)
by: Yao, Xiaozhe, et al.
Published: (2023)
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
by: Mei, Yixuan, et al.
Published: (2026)
by: Mei, Yixuan, et al.
Published: (2026)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
by: He, Xuan, et al.
Published: (2025)
by: He, Xuan, et al.
Published: (2025)
Serving Compound Inference Systems on Datacenter GPUs
by: Devata, Sriram, et al.
Published: (2026)
by: Devata, Sriram, et al.
Published: (2026)
CBEAOR: An Energy Aware Optimal Clustering and Routing Protocol for Sustainable IoT Healthcare Networks
by: Arthi Kalidasan, et al.
Published: (2025)
by: Arthi Kalidasan, et al.
Published: (2025)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
by: Yang, Shang, et al.
Published: (2025)
by: Yang, Shang, et al.
Published: (2025)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
by: Shi, Tianyao, et al.
Published: (2024)
by: Shi, Tianyao, et al.
Published: (2024)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
Asymptotic expansions for spectral convergence of compact self-adjoint operators on general spectral subsets, with application to kernel Gram matrices
by: Bae, Eunseong, et al.
Published: (2026)
by: Bae, Eunseong, et al.
Published: (2026)
Kernel smoothing on manifolds
by: Bae, Eunseong, et al.
Published: (2026)
by: Bae, Eunseong, et al.
Published: (2026)
Global Existence and Incompressible Limit for Compressible Navier-Stokes Equations in Bounded Domains with Large Bulk Viscosity Coefficient and Large Initial Data
by: Lei, Qinghao, et al.
Published: (2025)
by: Lei, Qinghao, et al.
Published: (2025)
Global Existence and Incompressible Limit of the Cauchy Problem for 2D Compressible Navier-Stokes Equations with Large Bulk Viscosity and Large Initial Data
by: Lei, Qinghao, et al.
Published: (2025)
by: Lei, Qinghao, et al.
Published: (2025)
Global Existence and Incompressible Limit for Compressible Navier-Stokes Equations with Large Bulk Viscosity Coefficient and Large Initial Data
by: Lei, Qinghao, et al.
Published: (2025)
by: Lei, Qinghao, et al.
Published: (2025)
Realism, Idealism and Progressive Outlook of U. R. Anathamurthy’s Bharthipura: A Critical Study
by: Vishnu Kumar
Published: (2019)
by: Vishnu Kumar
Published: (2019)
Conflict-Aware Soft Prompting for Retrieval-Augmented Generation
by: Choi, Eunseong, et al.
Published: (2025)
by: Choi, Eunseong, et al.
Published: (2025)
Polars inside Intel SGX2 Enclaves: An Empirical Study of Confidential Analytical Query Processing
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
Housing market connectedness and transmission of monetary policy
by: Woo Suk Lee, et al.
Published: (2025)
by: Woo Suk Lee, et al.
Published: (2025)
Confidential FRIT via Homomorphic Encryption
by: Hoshino, Haruki, et al.
Published: (2025)
by: Hoshino, Haruki, et al.
Published: (2025)
GPUs, CPUs, and... NICs: Rethinking the Network's Role in Serving Complex AI Pipelines
by: Wong, Mike, et al.
Published: (2025)
by: Wong, Mike, et al.
Published: (2025)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
by: Kim, Heehoon, et al.
Published: (2026)
by: Kim, Heehoon, et al.
Published: (2026)
Reasoning Language Model Inference Serving Unveiled: An Empirical Study
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
A Systematic Characterization of LLM Inference on GPUs
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
LLM-based Vulnerability Detection at Project Scale: An Empirical Study
by: Li, Fengjie, et al.
Published: (2026)
by: Li, Fengjie, et al.
Published: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
by: Kumar, Satyam, et al.
Published: (2026)
by: Kumar, Satyam, et al.
Published: (2026)
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
by: Li, Caihua, et al.
Published: (2024)
by: Li, Caihua, et al.
Published: (2024)
WebWeaver: Breaking Topology Confidentiality in LLM Multi-Agent Systems with Stealthy Context-Based Inference
by: Xiong, Zixun, et al.
Published: (2026)
by: Xiong, Zixun, et al.
Published: (2026)
MaskAdapt: Learning Flexible Motion Adaptation via Mask-Invariant Prior for Physics-Based Characters
by: Park, Soomin, et al.
Published: (2026)
by: Park, Soomin, et al.
Published: (2026)
From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression
by: Choi, Eunseong, et al.
Published: (2024)
by: Choi, Eunseong, et al.
Published: (2024)
Measurement-induced bistability in the excited states of a transmon
by: Choi, Jeakyung, et al.
Published: (2024)
by: Choi, Jeakyung, et al.
Published: (2024)
Multi-Granularity Guided Fusion-in-Decoder
by: Choi, Eunseong, et al.
Published: (2024)
by: Choi, Eunseong, et al.
Published: (2024)
Cosmological inference from combining Planck and ACT cluster counts
by: Lee, Eunseong, et al.
Published: (2024)
by: Lee, Eunseong, et al.
Published: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
Reliable and Resilient Collective Communication Library for LLM Training and Serving
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
by: Xu, Jiale, et al.
Published: (2025)
by: Xu, Jiale, et al.
Published: (2025)
Catalyst‐Free Synthesis of Quinazolinone‐Fused Quinoxaline Derivatives under Mild Conditions: Characterization, In Silico Prediction and Biological Evaluation
by: Venkadachalam Rahimiya, et al.
Published: (2025)
by: Venkadachalam Rahimiya, et al.
Published: (2025)
Quantization and Security Parameter Design for Overflow-Free Confidential FRIT
by: Park, Jungjin, et al.
Published: (2025)
by: Park, Jungjin, et al.
Published: (2025)
Similar Items
-
Performance of Confidential Computing GPUs
by: Ibarra, Antonio Martínez, et al.
Published: (2025) -
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
by: Zhu, Jianwei, et al.
Published: (2024) -
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
by: Jiang, Youhe, et al.
Published: (2025) -
DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
by: Yao, Xiaozhe, et al.
Published: (2023) -
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
by: Mei, Yixuan, et al.
Published: (2026)