Hiding Latencies in Network-Based Image Loading for Deep Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Versaci, Francesco, Busonera, Giovanni |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
by: Wang, Zhibin, et al.
Published: (2025)
by: Wang, Zhibin, et al.
Published: (2025)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
by: Yuan, Yitao, et al.
Published: (2025)
by: Yuan, Yitao, et al.
Published: (2025)
Load Balancing with Network Latencies via Distributed Gradient Descent
by: Balseiro, Santiago R., et al.
Published: (2025)
by: Balseiro, Santiago R., et al.
Published: (2025)
Strong and Hiding Distributed Certification of Bipartiteness
by: Jauregui, Benjamin, et al.
Published: (2025)
by: Jauregui, Benjamin, et al.
Published: (2025)
Low-Latency Federated Fine-Tuning for Large Language Models Over Wireless Networks
by: Pang, Zhiwen, et al.
Published: (2026)
by: Pang, Zhiwen, et al.
Published: (2026)
Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution
by: Wang, Haiquan, et al.
Published: (2024)
by: Wang, Haiquan, et al.
Published: (2024)
Asynchronous Latency and Fast Atomic Snapshot
by: Bezerra, João Paulo, et al.
Published: (2024)
by: Bezerra, João Paulo, et al.
Published: (2024)
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
by: Seo, Eunil, et al.
Published: (2025)
by: Seo, Eunil, et al.
Published: (2025)
POSEIDON : Efficient Function Placement at the Edge using Deep Reinforcement Learning
by: Jain, Prakhar, et al.
Published: (2024)
by: Jain, Prakhar, et al.
Published: (2024)
Workload Distribution with Rateless Encoding: A Low-Latency Computation Offloading Method within Edge Networks
by: Guo, Zhongfu, et al.
Published: (2023)
by: Guo, Zhongfu, et al.
Published: (2023)
Formal Specification for Fast ACS: Low-Latency File-Based Ordered Message Delivery at Scale
by: Gupta, Sushant Kumar, et al.
Published: (2025)
by: Gupta, Sushant Kumar, et al.
Published: (2025)
DAG it off: Latency Prefers No Common Coins
by: Amores-Sesar, Ignacio, et al.
Published: (2025)
by: Amores-Sesar, Ignacio, et al.
Published: (2025)
Methodology for GPU Frequency Switching Latency Measurement
by: Velicka, Daniel, et al.
Published: (2025)
by: Velicka, Daniel, et al.
Published: (2025)
Efficient Graph-Based Approximate Nearest Neighbor Search Achieving: Low Latency Without Throughput Loss
by: Luo, Jingjia, et al.
Published: (2025)
by: Luo, Jingjia, et al.
Published: (2025)
Areon: Latency-Friendly and Resilient Multi-Proposer Consensus
by: Castro-Castilla, Álvaro, et al.
Published: (2025)
by: Castro-Castilla, Álvaro, et al.
Published: (2025)
Inference Load-Aware Orchestration for Hierarchical Federated Learning
by: Lackinger, Anna, et al.
Published: (2024)
by: Lackinger, Anna, et al.
Published: (2024)
Low Latency, High Bandwidth Streaming of Experimental Data with EJFAT
by: Baldin, Ilya, et al.
Published: (2025)
by: Baldin, Ilya, et al.
Published: (2025)
Revisiting Speculative Leaderless Protocols for Low-Latency BFT Replication
by: Qian, Daniel, et al.
Published: (2026)
by: Qian, Daniel, et al.
Published: (2026)
Falcon: Advancing Asynchronous BFT Consensus for Lower Latency and Enhanced Throughput
by: Dai, Xiaohai, et al.
Published: (2025)
by: Dai, Xiaohai, et al.
Published: (2025)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
by: Ma, Chenxiang, et al.
Published: (2025)
by: Ma, Chenxiang, et al.
Published: (2025)
CD-Raft: Reducing the Latency of Distributed Consensus in Cross-Domain Sites
by: Wang, Yangyang, et al.
Published: (2026)
by: Wang, Yangyang, et al.
Published: (2026)
A New Approach for Evaluating the Performance of Distributed Latency-Sensitive Services
by: Theodoropoulos, Theodoros, et al.
Published: (2024)
by: Theodoropoulos, Theodoros, et al.
Published: (2024)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
by: Wu, Siyu, et al.
Published: (2025)
by: Wu, Siyu, et al.
Published: (2025)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
by: Hidayetoglu, Mert, et al.
Published: (2025)
by: Hidayetoglu, Mert, et al.
Published: (2025)
Shared Memory-Aware Latency-Sensitive Message Aggregation for Fine-Grained Communication
by: Chandrasekar, Kavitha, et al.
Published: (2024)
by: Chandrasekar, Kavitha, et al.
Published: (2024)
BBCA-CHAIN: Low Latency, High Throughput BFT Consensus on a DAG
by: Malkhi, Dahlia, et al.
Published: (2023)
by: Malkhi, Dahlia, et al.
Published: (2023)
Scene-Aware Latency Estimation for Microservices via Multi-Scale Graph Fusion
by: Sun, Zhichao, et al.
Published: (2026)
by: Sun, Zhichao, et al.
Published: (2026)
Low-Latency Layer-Aware Proactive and Passive Container Migration in Meta Computing
by: Liu, Mengjie, et al.
Published: (2024)
by: Liu, Mengjie, et al.
Published: (2024)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
by: Yu, Minchen, et al.
Published: (2023)
by: Yu, Minchen, et al.
Published: (2023)
Chasing the Speed of Light: Low-Latency Planetary-Scale Adaptive Byzantine Consensus
by: Berger, Christian, et al.
Published: (2023)
by: Berger, Christian, et al.
Published: (2023)
On the Solvability of Byzantine-tolerant Reliable Communication in Dynamic Networks
by: Bonomi, Silvia, et al.
Published: (2025)
by: Bonomi, Silvia, et al.
Published: (2025)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
by: Khoshsirat, Aria, et al.
Published: (2024)
by: Khoshsirat, Aria, et al.
Published: (2024)
Universal Finite-State and Self-Stabilizing Computation in Anonymous Dynamic Networks
by: Di Luna, Giuseppe A., et al.
Published: (2024)
by: Di Luna, Giuseppe A., et al.
Published: (2024)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
by: Lou, Chiheng, et al.
Published: (2025)
by: Lou, Chiheng, et al.
Published: (2025)
TailBench++: Flexible Multi-Client, Multi-Server Benchmarking for Latency-Critical Workloads
by: Li, Zhilin, et al.
Published: (2025)
by: Li, Zhilin, et al.
Published: (2025)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
by: Wu, Tian, et al.
Published: (2025)
by: Wu, Tian, et al.
Published: (2025)
Modeling Anomaly Detection in Cloud Services: Analysis of the Properties that Impact Latency and Resource Consumption
by: Grabher, Gabriel Job Antunes, et al.
Published: (2025)
by: Grabher, Gabriel Job Antunes, et al.
Published: (2025)
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
by: Wang, Jun, et al.
Published: (2025)
by: Wang, Jun, et al.
Published: (2025)
EMLIO: Minimizing I/O Latency and Energy Consumption for Large-Scale AI Training
by: Jamil, Hasibul, et al.
Published: (2025)
by: Jamil, Hasibul, et al.
Published: (2025)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
by: Wang, Wenfeng, et al.
Published: (2026)
by: Wang, Wenfeng, et al.
Published: (2026)
Similar Items
-
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
by: Wang, Zhibin, et al.
Published: (2025) -
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
by: Yuan, Yitao, et al.
Published: (2025) -
Load Balancing with Network Latencies via Distributed Gradient Descent
by: Balseiro, Santiago R., et al.
Published: (2025) -
Strong and Hiding Distributed Certification of Bipartiteness
by: Jauregui, Benjamin, et al.
Published: (2025) -
Low-Latency Federated Fine-Tuning for Large Language Models Over Wireless Networks
by: Pang, Zhiwen, et al.
Published: (2026)