Accuracy-Delay Trade-Off in LLM Offloading via Token-Level Uncertainty
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Yumin, Lyu, Hyeonsu, Lee, Minjae, Yang, Hyun Jong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On-demand Cold Start Frequency Reduction with Off-Policy Reinforcement Learning in Serverless Computing
por: Agarwal, Siddharth, et al.
Publicado: (2023)
por: Agarwal, Siddharth, et al.
Publicado: (2023)
A Case Study of API Design for Interoperability and Security of the Internet of Things
por: Kim, Dongha, et al.
Publicado: (2024)
por: Kim, Dongha, et al.
Publicado: (2024)
AdapTBF: Decentralized Bandwidth Control via Adaptive Token Borrowing for HPC Storage
por: Rashid, Md Hasanur, et al.
Publicado: (2026)
por: Rashid, Md Hasanur, et al.
Publicado: (2026)
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
por: Lee, Jin, et al.
Publicado: (2026)
por: Lee, Jin, et al.
Publicado: (2026)
Efficient Coordination for Distributed Discrete-Event Systems
por: Jun, Byeonggil, et al.
Publicado: (2024)
por: Jun, Byeonggil, et al.
Publicado: (2024)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
por: Yu, Hanfei, et al.
Publicado: (2025)
por: Yu, Hanfei, et al.
Publicado: (2025)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
por: Jain, Kunal, et al.
Publicado: (2024)
por: Jain, Kunal, et al.
Publicado: (2024)
Pythia: Exploiting Workflow Predictability for Efficient Agent-Native LLM Serving
por: Yu, Shan, et al.
Publicado: (2026)
por: Yu, Shan, et al.
Publicado: (2026)
From Barrier to Bridge: The Case for AI Data Center/Power Grid Co-Design
por: Bashir, Noman, et al.
Publicado: (2026)
por: Bashir, Noman, et al.
Publicado: (2026)
Multi-Agentic AI for Fairness-Aware and Accelerated Multi-modal Large Model Inference in Real-world Mobile Edge Networks
por: Li, Haiyuan, et al.
Publicado: (2026)
por: Li, Haiyuan, et al.
Publicado: (2026)
Systemic approach for modeling a generic smart grid
por: Amor, Sofiane Ben, et al.
Publicado: (2025)
por: Amor, Sofiane Ben, et al.
Publicado: (2025)
A Deep Recurrent-Reinforcement Learning Method for Intelligent AutoScaling of Serverless Functions
por: Agarwal, Siddharth, et al.
Publicado: (2023)
por: Agarwal, Siddharth, et al.
Publicado: (2023)
Input-Based Ensemble-Learning Method for Dynamic Memory Configuration of Serverless Computing Functions
por: Agarwal, Siddharth, et al.
Publicado: (2024)
por: Agarwal, Siddharth, et al.
Publicado: (2024)
Tackling the Crowdsourced Shared-Trip Delivery Problem at Scale with a Novel Decomposition Heuristic
por: Yang, Dingtong, et al.
Publicado: (2022)
por: Yang, Dingtong, et al.
Publicado: (2022)
PMU-based Distributed Non-iterative Algorithm for Real-time Voltage Stability Monitoring
por: Guddanti, Kishan Prudhvi, et al.
Publicado: (2019)
por: Guddanti, Kishan Prudhvi, et al.
Publicado: (2019)
Hybrid Cooperative Co-Evolution Algorithm for Deadlock-prone Distributed Assembly Flowshop Scheduling with Limited buffers Using Petri nets
por: Wang, Siyi, et al.
Publicado: (2024)
por: Wang, Siyi, et al.
Publicado: (2024)
Plug & Offload: Transparently Offloading TCP Stack onto Off-path SmartNIC with PnO-TCP
por: Nan, Hailong, et al.
Publicado: (2025)
por: Nan, Hailong, et al.
Publicado: (2025)
Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge Platforms
por: Taufique, Zain, et al.
Publicado: (2023)
por: Taufique, Zain, et al.
Publicado: (2023)
Rethinking Inter-Process Communication with Memory Operation Offloading
por: Park, Misun, et al.
Publicado: (2026)
por: Park, Misun, et al.
Publicado: (2026)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
por: Jeong, Bodon, et al.
Publicado: (2026)
por: Jeong, Bodon, et al.
Publicado: (2026)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
por: Zhang, Wei, et al.
Publicado: (2025)
por: Zhang, Wei, et al.
Publicado: (2025)
AMV-L: Lifecycle-Managed Agent Memory for Tail-Latency Control in Long-Running LLM Systems
por: Bamidele, Emmanuel
Publicado: (2026)
por: Bamidele, Emmanuel
Publicado: (2026)
Distributed Difference of Convex Optimization
por: Khatana, Vivek, et al.
Publicado: (2024)
por: Khatana, Vivek, et al.
Publicado: (2024)
Robust Set Partitioning Strategy for Malicious Information Detection in Large-Scale Internet of Things
por: Suo, Yuhan, et al.
Publicado: (2025)
por: Suo, Yuhan, et al.
Publicado: (2025)
Designing Dense Satellite Clusters for Distributed Space-based Datacenters
por: Pénot, Jules, et al.
Publicado: (2026)
por: Pénot, Jules, et al.
Publicado: (2026)
Load Balancing Using Sparse Communication
por: Mendelson, Gal, et al.
Publicado: (2022)
por: Mendelson, Gal, et al.
Publicado: (2022)
Vessim: A Testbed for Carbon-Aware Applications and Systems
por: Wiesner, Philipp, et al.
Publicado: (2023)
por: Wiesner, Philipp, et al.
Publicado: (2023)
Iterative Thresholding and Projection Algorithms and Model-Based Deep Neural Networks for Sparse LQR Control Design
por: Cho, Myung
Publicado: (2022)
por: Cho, Myung
Publicado: (2022)
Deployment of Containerized Simulations in an API-Driven Distributed Infrastructure
por: Kraus, Tim, et al.
Publicado: (2025)
por: Kraus, Tim, et al.
Publicado: (2025)
Distribution and Management of Datacenter Load Decoupling
por: Lin, Liuzixuan, et al.
Publicado: (2025)
por: Lin, Liuzixuan, et al.
Publicado: (2025)
Stannic: Systolic STochAstic ONliNe SchedulIng AcCelerator
por: Ross, Adam H., et al.
Publicado: (2025)
por: Ross, Adam H., et al.
Publicado: (2025)
Thinking fast and slow -- a cognitive inspired framework for decision intelligence for power systems
por: Mathur, Apoorv
Publicado: (2026)
por: Mathur, Apoorv
Publicado: (2026)
H-MBR: Hypervisor-level Memory Bandwidth Reservation for Mixed Criticality Systems
por: Oliveira, Afonso, et al.
Publicado: (2025)
por: Oliveira, Afonso, et al.
Publicado: (2025)
MQFQ-Sticky: Fair Queueing For Serverless GPU Functions
por: Fuerst, Alexander, et al.
Publicado: (2025)
por: Fuerst, Alexander, et al.
Publicado: (2025)
GPUArmor: A Hardware-Software Co-design for Efficient and Scalable Memory Safety on GPUs
por: Ziad, Mohamed Tarek Ibn, et al.
Publicado: (2025)
por: Ziad, Mohamed Tarek Ibn, et al.
Publicado: (2025)
TALICS$^3$: Tape Library Cloud Storage System Simulator
por: Arslan, Suayb S., et al.
Publicado: (2024)
por: Arslan, Suayb S., et al.
Publicado: (2024)
Finite State Machines-Based Path-Following Collaborative Computing Strategy for Emergency UAV Swarms
por: Hu, Jialin, et al.
Publicado: (2024)
por: Hu, Jialin, et al.
Publicado: (2024)
Carbon-Aware Quality Adaptation for Energy-Intensive Services
por: Wiesner, Philipp, et al.
Publicado: (2024)
por: Wiesner, Philipp, et al.
Publicado: (2024)
Computationally Efficient Laplacian CL-colME
por: Stankovic, Nikola
Publicado: (2026)
por: Stankovic, Nikola
Publicado: (2026)
AI-focused HPC Data Centers Can Provide More Power Grid Flexibility and at Lower Cost
por: Zhou, Yihong, et al.
Publicado: (2024)
por: Zhou, Yihong, et al.
Publicado: (2024)
Ejemplares similares
-
On-demand Cold Start Frequency Reduction with Off-Policy Reinforcement Learning in Serverless Computing
por: Agarwal, Siddharth, et al.
Publicado: (2023) -
A Case Study of API Design for Interoperability and Security of the Internet of Things
por: Kim, Dongha, et al.
Publicado: (2024) -
AdapTBF: Decentralized Bandwidth Control via Adaptive Token Borrowing for HPC Storage
por: Rashid, Md Hasanur, et al.
Publicado: (2026) -
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
por: Lee, Jin, et al.
Publicado: (2026) -
Efficient Coordination for Distributed Discrete-Event Systems
por: Jun, Byeonggil, et al.
Publicado: (2024)