LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Seo, Eunil, Nguyen, Chanh, Elmroth, Erik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Speculative Policy Orchestration: A Latency-Resilient Framework for Cloud-Robotic Manipulation
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026)
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026)
Taming Cold Starts: Proactive Serverless Scheduling with Model Predictive Control
von: Nguyen, Chanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Chanh, et al.
Veröffentlicht: (2025)
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026)
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026)
Silent Failures in Stateless Systems: Rethinking Anomaly Detection for Serverless Computing
von: Nguyen, Chanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Chanh, et al.
Veröffentlicht: (2025)
Shared Memory-Aware Latency-Sensitive Message Aggregation for Fine-Grained Communication
von: Chandrasekar, Kavitha, et al.
Veröffentlicht: (2024)
von: Chandrasekar, Kavitha, et al.
Veröffentlicht: (2024)
Predictive Autoscaling for Node.js on Kubernetes: Lower Latency, Right-Sized Capacity
von: Tymoshenko, Ivan, et al.
Veröffentlicht: (2026)
von: Tymoshenko, Ivan, et al.
Veröffentlicht: (2026)
Low-Latency Layer-Aware Proactive and Passive Container Migration in Meta Computing
von: Liu, Mengjie, et al.
Veröffentlicht: (2024)
von: Liu, Mengjie, et al.
Veröffentlicht: (2024)
Proactive and Reactive Autoscaling Techniques for Edge Computing
von: Gupta, Suhrid, et al.
Veröffentlicht: (2025)
von: Gupta, Suhrid, et al.
Veröffentlicht: (2025)
Action Deviation-Aware Inference for Low-Latency Wireless Robots
von: Park, Jeyoung, et al.
Veröffentlicht: (2025)
von: Park, Jeyoung, et al.
Veröffentlicht: (2025)
HE2C: A Holistic Approach for Allocating Latency-Sensitive AI Tasks across Edge-Cloud
von: Kim, Minseo, et al.
Veröffentlicht: (2024)
von: Kim, Minseo, et al.
Veröffentlicht: (2024)
TailBench++: Flexible Multi-Client, Multi-Server Benchmarking for Latency-Critical Workloads
von: Li, Zhilin, et al.
Veröffentlicht: (2025)
von: Li, Zhilin, et al.
Veröffentlicht: (2025)
A New Approach for Evaluating the Performance of Distributed Latency-Sensitive Services
von: Theodoropoulos, Theodoros, et al.
Veröffentlicht: (2024)
von: Theodoropoulos, Theodoros, et al.
Veröffentlicht: (2024)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
von: Jeffery, Andrew, et al.
Veröffentlicht: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
von: Lou, Chiheng, et al.
Veröffentlicht: (2025)
Modeling Anomaly Detection in Cloud Services: Analysis of the Properties that Impact Latency and Resource Consumption
von: Grabher, Gabriel Job Antunes, et al.
Veröffentlicht: (2025)
von: Grabher, Gabriel Job Antunes, et al.
Veröffentlicht: (2025)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
von: Qi, S., et al.
Veröffentlicht: (2024)
von: Qi, S., et al.
Veröffentlicht: (2024)
Workload Buoyancy: Keeping Apps Afloat by Identifying Shared Resource Bottlenecks
von: Larsson, Oliver, et al.
Veröffentlicht: (2026)
von: Larsson, Oliver, et al.
Veröffentlicht: (2026)
Scene-Aware Latency Estimation for Microservices via Multi-Scale Graph Fusion
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
Asynchronous Latency and Fast Atomic Snapshot
von: Bezerra, João Paulo, et al.
Veröffentlicht: (2024)
von: Bezerra, João Paulo, et al.
Veröffentlicht: (2024)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
Optimising Virtual Resource Mapping in Multi-Level NUMA Disaggregated Systems
von: Lakew, Ewnetu Bayuh, et al.
Veröffentlicht: (2025)
von: Lakew, Ewnetu Bayuh, et al.
Veröffentlicht: (2025)
DAG it off: Latency Prefers No Common Coins
von: Amores-Sesar, Ignacio, et al.
Veröffentlicht: (2025)
von: Amores-Sesar, Ignacio, et al.
Veröffentlicht: (2025)
Methodology for GPU Frequency Switching Latency Measurement
von: Velicka, Daniel, et al.
Veröffentlicht: (2025)
von: Velicka, Daniel, et al.
Veröffentlicht: (2025)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
HGraphScale: Hierarchical Graph Learning for Autoscaling Microservice Applications in Container-based Cloud Computing
von: Fang, Zhengxin, et al.
Veröffentlicht: (2025)
von: Fang, Zhengxin, et al.
Veröffentlicht: (2025)
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
von: Yang, Yuting, et al.
Veröffentlicht: (2024)
von: Yang, Yuting, et al.
Veröffentlicht: (2024)
Areon: Latency-Friendly and Resilient Multi-Proposer Consensus
von: Castro-Castilla, Álvaro, et al.
Veröffentlicht: (2025)
von: Castro-Castilla, Álvaro, et al.
Veröffentlicht: (2025)
Hiding Latencies in Network-Based Image Loading for Deep Learning
von: Versaci, Francesco, et al.
Veröffentlicht: (2025)
von: Versaci, Francesco, et al.
Veröffentlicht: (2025)
Low Latency, High Bandwidth Streaming of Experimental Data with EJFAT
von: Baldin, Ilya, et al.
Veröffentlicht: (2025)
von: Baldin, Ilya, et al.
Veröffentlicht: (2025)
Revisiting Speculative Leaderless Protocols for Low-Latency BFT Replication
von: Qian, Daniel, et al.
Veröffentlicht: (2026)
von: Qian, Daniel, et al.
Veröffentlicht: (2026)
FogROS2-PLR: Probabilistic Latency-Reliability For Cloud Robotics
von: Chen, Kaiyuan, et al.
Veröffentlicht: (2024)
von: Chen, Kaiyuan, et al.
Veröffentlicht: (2024)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
Self-adaptive, Requirements-driven Autoscaling of Microservices
von: Nunes, João Paulo Karol Santos, et al.
Veröffentlicht: (2024)
von: Nunes, João Paulo Karol Santos, et al.
Veröffentlicht: (2024)
Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
Falcon: Advancing Asynchronous BFT Consensus for Lower Latency and Enhanced Throughput
von: Dai, Xiaohai, et al.
Veröffentlicht: (2025)
von: Dai, Xiaohai, et al.
Veröffentlicht: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
CD-Raft: Reducing the Latency of Distributed Consensus in Cross-Domain Sites
von: Wang, Yangyang, et al.
Veröffentlicht: (2026)
von: Wang, Yangyang, et al.
Veröffentlicht: (2026)
Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge Platforms
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2024)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2024)
A 1024 RV-Cores Shared-L1 Cluster with High Bandwidth Memory Link for Low-Latency 6G-SDR
von: Zhang, Yichao, et al.
Veröffentlicht: (2024)
von: Zhang, Yichao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Speculative Policy Orchestration: A Latency-Resilient Framework for Cloud-Robotic Manipulation
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026) -
Taming Cold Starts: Proactive Serverless Scheduling with Model Predictive Control
von: Nguyen, Chanh, et al.
Veröffentlicht: (2025) -
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026) -
Silent Failures in Stateless Systems: Rethinking Anomaly Detection for Serverless Computing
von: Nguyen, Chanh, et al.
Veröffentlicht: (2025) -
Shared Memory-Aware Latency-Sensitive Message Aggregation for Fine-Grained Communication
von: Chandrasekar, Kavitha, et al.
Veröffentlicht: (2024)