Shared Memory-Aware Latency-Sensitive Message Aggregation for Fine-Grained Communication
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chandrasekar, Kavitha, Kale, Laxmikant |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Communication-Aware Diffusion Load Balancing for Persistently Interacting Objects
von: Taylor, Maya, et al.
Veröffentlicht: (2026)
von: Taylor, Maya, et al.
Veröffentlicht: (2026)
An Elastic Job Scheduler for HPC Applications on the Cloud
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
An Adaptive Distributed Stencil Abstraction for GPUs
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
CkIO: Parallel File Input for Over-Decomposed Task-Based Systems
von: Jacob, Mathew, et al.
Veröffentlicht: (2024)
von: Jacob, Mathew, et al.
Veröffentlicht: (2024)
Efficient and Portable Support for Overdecomposition on Distributed Memory GPGPU Platforms
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
Towards an Adaptive Runtime System for Cloud-Native HPC
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
von: Seo, Eunil, et al.
Veröffentlicht: (2025)
von: Seo, Eunil, et al.
Veröffentlicht: (2025)
MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
von: Zhao, Lu, et al.
Veröffentlicht: (2025)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
Formal Specification for Fast ACS: Low-Latency File-Based Ordered Message Delivery at Scale
von: Gupta, Sushant Kumar, et al.
Veröffentlicht: (2025)
von: Gupta, Sushant Kumar, et al.
Veröffentlicht: (2025)
A 1024 RV-Cores Shared-L1 Cluster with High Bandwidth Memory Link for Low-Latency 6G-SDR
von: Zhang, Yichao, et al.
Veröffentlicht: (2024)
von: Zhang, Yichao, et al.
Veröffentlicht: (2024)
DRust: Language-Guided Distributed Shared Memory with Fine Granularity, Full Transparency, and Ultra Efficiency
von: Ma, Haoran, et al.
Veröffentlicht: (2024)
von: Ma, Haoran, et al.
Veröffentlicht: (2024)
A New Approach for Evaluating the Performance of Distributed Latency-Sensitive Services
von: Theodoropoulos, Theodoros, et al.
Veröffentlicht: (2024)
von: Theodoropoulos, Theodoros, et al.
Veröffentlicht: (2024)
A Communication- and Memory-Aware Model for Load Balancing Tasks
von: Lifflander, Jonathan, et al.
Veröffentlicht: (2024)
von: Lifflander, Jonathan, et al.
Veröffentlicht: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
von: Ma, Chenxiang, et al.
Veröffentlicht: (2025)
TeraNoC: A Multi-Channel 32-bit Fine-Grained, Hybrid Mesh-Crossbar NoC for Efficient Scale-up of 1000+ Core Shared-L1-Memory Clusters
von: Zhang, Yichao, et al.
Veröffentlicht: (2025)
von: Zhang, Yichao, et al.
Veröffentlicht: (2025)
Distributing Context-Aware Shared Memory Data Structures: A Case Study on Singly-Linked Lists
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2024)
von: Ravishankar, Raaghav, et al.
Veröffentlicht: (2024)
Low-Latency Federated Fine-Tuning for Large Language Models Over Wireless Networks
von: Pang, Zhiwen, et al.
Veröffentlicht: (2026)
von: Pang, Zhiwen, et al.
Veröffentlicht: (2026)
Modeling the Potential of Message-Free Communication via CXL.mem
von: Vanecek, Stepan, et al.
Veröffentlicht: (2025)
von: Vanecek, Stepan, et al.
Veröffentlicht: (2025)
Low-Latency Layer-Aware Proactive and Passive Container Migration in Meta Computing
von: Liu, Mengjie, et al.
Veröffentlicht: (2024)
von: Liu, Mengjie, et al.
Veröffentlicht: (2024)
Scene-Aware Latency Estimation for Microservices via Multi-Scale Graph Fusion
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
von: Sun, Zhichao, et al.
Veröffentlicht: (2026)
Do MPI Derived Datatypes Actually Help? A Single-Node Cross-Implementation Study on Shared-Memory Communication
von: Adefemi, Temitayo
Veröffentlicht: (2025)
von: Adefemi, Temitayo
Veröffentlicht: (2025)
PATSMA: Parameter Auto-tuning for Shared Memory Algorithms
von: Fernandes, Joao B., et al.
Veröffentlicht: (2024)
von: Fernandes, Joao B., et al.
Veröffentlicht: (2024)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
von: Murat, Antoine, et al.
Veröffentlicht: (2024)
von: Murat, Antoine, et al.
Veröffentlicht: (2024)
Byzantine-Tolerant Consensus in GPU-Inspired Shared Memory
von: Georgiou, Chryssis, et al.
Veröffentlicht: (2025)
von: Georgiou, Chryssis, et al.
Veröffentlicht: (2025)
HE2C: A Holistic Approach for Allocating Latency-Sensitive AI Tasks across Edge-Cloud
von: Kim, Minseo, et al.
Veröffentlicht: (2024)
von: Kim, Minseo, et al.
Veröffentlicht: (2024)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
Near-Optimal Communication Byzantine Reliable Broadcast under a Message Adversary
von: Albouy, Timothé, et al.
Veröffentlicht: (2023)
von: Albouy, Timothé, et al.
Veröffentlicht: (2023)
Optimizing Communication for Latency Sensitive HPC Applications on up to 48 FPGAs Using ACCL
von: Meyer, Marius, et al.
Veröffentlicht: (2024)
von: Meyer, Marius, et al.
Veröffentlicht: (2024)
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
Accelerating Latency-Critical Applications with AI-Powered Semi-Automatic Fine-Grained Parallelization on SMT Processors
von: Los, Denis, et al.
Veröffentlicht: (2025)
von: Los, Denis, et al.
Veröffentlicht: (2025)
FedFQ: Federated Learning with Fine-Grained Quantization
von: Li, Haowei, et al.
Veröffentlicht: (2024)
von: Li, Haowei, et al.
Veröffentlicht: (2024)
Robust Federated Fine-Tuning in Heterogeneous Networks with Unreliable Connections: An Aggregation View
von: Wang, Yanmeng, et al.
Veröffentlicht: (2025)
von: Wang, Yanmeng, et al.
Veröffentlicht: (2025)
Shared Virtual Memory: Its Design and Performance Implications for Diverse Applications
von: Cooper, Bennett, et al.
Veröffentlicht: (2024)
von: Cooper, Bennett, et al.
Veröffentlicht: (2024)
A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning
von: Lyu, Chenghao, et al.
Veröffentlicht: (2024)
von: Lyu, Chenghao, et al.
Veröffentlicht: (2024)
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
von: Jangda, Abhinav, et al.
Veröffentlicht: (2023)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2023)
Towards Fine-Grained Scalability for Stateful Stream Processing Systems
von: Qing, Yunfan, et al.
Veröffentlicht: (2025)
von: Qing, Yunfan, et al.
Veröffentlicht: (2025)
An Asynchronous Many-Task Algorithm for Unstructured $S_{N}$ Transport on Shared Memory Systems
von: Elwood, Alex, et al.
Veröffentlicht: (2025)
von: Elwood, Alex, et al.
Veröffentlicht: (2025)
Selection Guidelines for Geo-Replicated SMR Protocols: A Communication Pattern-based Latency Modeling Approach
von: Shiozaki, Kohya, et al.
Veröffentlicht: (2024)
von: Shiozaki, Kohya, et al.
Veröffentlicht: (2024)
Action Deviation-Aware Inference for Low-Latency Wireless Robots
von: Park, Jeyoung, et al.
Veröffentlicht: (2025)
von: Park, Jeyoung, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Communication-Aware Diffusion Load Balancing for Persistently Interacting Objects
von: Taylor, Maya, et al.
Veröffentlicht: (2026) -
An Elastic Job Scheduler for HPC Applications on the Cloud
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025) -
An Adaptive Distributed Stencil Abstraction for GPUs
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025) -
CkIO: Parallel File Input for Over-Decomposed Task-Based Systems
von: Jacob, Mathew, et al.
Veröffentlicht: (2024) -
Efficient and Portable Support for Overdecomposition on Distributed Memory GPGPU Platforms
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)