Computation-Bandwidth-Memory Trade-offs: A Unified Paradigm for AI Infrastructure
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Yuankai, Weng, Qizhen, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Space-Time Trade-off in Bounded Iterated Memory
von: Toyos-Marfurt, Guillermo, et al.
Veröffentlicht: (2025)
von: Toyos-Marfurt, Guillermo, et al.
Veröffentlicht: (2025)
Paradigm Shift in Infrastructure Inspection Technology: Leveraging High-performance Imaging and Advanced AI Analytics to Inspect Road Infrastructure
von: Wu, Du, et al.
Veröffentlicht: (2025)
von: Wu, Du, et al.
Veröffentlicht: (2025)
Performance Trade-offs of High Order Meshless Approximation on Distributed Memory Systems
von: Vehovar, Jon, et al.
Veröffentlicht: (2025)
von: Vehovar, Jon, et al.
Veröffentlicht: (2025)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
Trade-offs in Decentralized Agentic AI Discovery Across the Compute Continuum
von: Dazzi, Patrizio, et al.
Veröffentlicht: (2026)
von: Dazzi, Patrizio, et al.
Veröffentlicht: (2026)
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
von: Khalilov, Mikhail, et al.
Veröffentlicht: (2024)
von: Khalilov, Mikhail, et al.
Veröffentlicht: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
A Study on Messaging Trade-offs in Data Streaming for Scientific Workflows
von: George, Anjus, et al.
Veröffentlicht: (2025)
von: George, Anjus, et al.
Veröffentlicht: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
A 1024 RV-Cores Shared-L1 Cluster with High Bandwidth Memory Link for Low-Latency 6G-SDR
von: Zhang, Yichao, et al.
Veröffentlicht: (2024)
von: Zhang, Yichao, et al.
Veröffentlicht: (2024)
6G Infrastructures for Edge AI: An Analytical Perspective
von: Horvath, Kurt, et al.
Veröffentlicht: (2025)
von: Horvath, Kurt, et al.
Veröffentlicht: (2025)
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
CHIRON: Accelerating Node Synchronization without Security Trade-offs in Distributed Ledgers
von: Neiheiser, Ray, et al.
Veröffentlicht: (2024)
von: Neiheiser, Ray, et al.
Veröffentlicht: (2024)
A Space-Time Trade-off for Fast Self-Stabilizing Leader Election in Population Protocols
von: Austin, Henry, et al.
Veröffentlicht: (2025)
von: Austin, Henry, et al.
Veröffentlicht: (2025)
EPIC: An Energy-Efficient, High-Performance GPGPU Computing Research Infrastructure
von: Själander, Magnus, et al.
Veröffentlicht: (2019)
von: Själander, Magnus, et al.
Veröffentlicht: (2019)
Federated Single Sign-On and Zero Trust Co-design for AI and HPC Digital Research Infrastructures
von: Alam, Sadaf R., et al.
Veröffentlicht: (2024)
von: Alam, Sadaf R., et al.
Veröffentlicht: (2024)
Tutorial: Object as a Service (OaaS) Serverless Cloud Computing Paradigm
von: Lertpongrujikorn, Pawissanutt, et al.
Veröffentlicht: (2024)
von: Lertpongrujikorn, Pawissanutt, et al.
Veröffentlicht: (2024)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
von: Lu, Yao, et al.
Veröffentlicht: (2026)
von: Lu, Yao, et al.
Veröffentlicht: (2026)
PUSHtap: PIM-based In-Memory HTAP with Unified Data Storage Format
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
von: Xiong, Yifan, et al.
Veröffentlicht: (2024)
von: Xiong, Yifan, et al.
Veröffentlicht: (2024)
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
von: Suman, Shekhar, et al.
Veröffentlicht: (2026)
von: Suman, Shekhar, et al.
Veröffentlicht: (2026)
M$^2$-MFP: A Multi-Scale and Multi-Level Memory Failure Prediction Framework for Reliable Cloud Infrastructure
von: Xie, Hongyi, et al.
Veröffentlicht: (2025)
von: Xie, Hongyi, et al.
Veröffentlicht: (2025)
Bandwidth-Aware Network Topology Optimization for Decentralized Learning
von: Shen, Yipeng, et al.
Veröffentlicht: (2025)
von: Shen, Yipeng, et al.
Veröffentlicht: (2025)
The Carnot Bound: Limits and Possibilities for Bandwidth-Efficient Consensus
von: Lewis-Pye, Andrew, et al.
Veröffentlicht: (2026)
von: Lewis-Pye, Andrew, et al.
Veröffentlicht: (2026)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
RL over Commodity Networks: Overcoming the Bandwidth Barrier with Lossless Sparse Deltas
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2026)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2026)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
von: Li, Zixuan, et al.
Veröffentlicht: (2026)
von: Li, Zixuan, et al.
Veröffentlicht: (2026)
Low Latency, High Bandwidth Streaming of Experimental Data with EJFAT
von: Baldin, Ilya, et al.
Veröffentlicht: (2025)
von: Baldin, Ilya, et al.
Veröffentlicht: (2025)
Exploring Performance-Productivity Trade-offs in AMT Runtimes: A Task Bench Study of Itoyori, ItoyoriFBC, HPX, and MPI
von: Lahnor, Torben R., et al.
Veröffentlicht: (2026)
von: Lahnor, Torben R., et al.
Veröffentlicht: (2026)
Memory and Bandwidth are All You Need for Fully Sharded Data Parallel
von: Wang, Jiangtao, et al.
Veröffentlicht: (2025)
von: Wang, Jiangtao, et al.
Veröffentlicht: (2025)
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
von: Wu, Yongtong, et al.
Veröffentlicht: (2026)
von: Wu, Yongtong, et al.
Veröffentlicht: (2026)
A Unified Programming Model for Heterogeneous Computing with CPU and Accelerator Technologies
von: Xiong, Yuqing
Veröffentlicht: (2022)
von: Xiong, Yuqing
Veröffentlicht: (2022)
Repurposing of the Run 2 CMS High Level Trigger Infrastructure as a Cloud Resource for Offline Computing
von: Mascheroni, Marco, et al.
Veröffentlicht: (2024)
von: Mascheroni, Marco, et al.
Veröffentlicht: (2024)
Edge AI in Highly Volatile Environments: Is Fairness Worth the Accuracy Trade-off?
von: Zaland, Obaidullah, et al.
Veröffentlicht: (2025)
von: Zaland, Obaidullah, et al.
Veröffentlicht: (2025)
System-Level Performance Modeling of Photonic In-Memory Computing
von: Arockiaraj, Jebacyril, et al.
Veröffentlicht: (2026)
von: Arockiaraj, Jebacyril, et al.
Veröffentlicht: (2026)
Incremental GNN Embedding Computation on Streaming Graphs
von: Wang, Qiange, et al.
Veröffentlicht: (2026)
von: Wang, Qiange, et al.
Veröffentlicht: (2026)
WANify: Gauging and Balancing Runtime WAN Bandwidth for Geo-distributed Data Analytics
von: Mohapatra, Anshuman Das, et al.
Veröffentlicht: (2025)
von: Mohapatra, Anshuman Das, et al.
Veröffentlicht: (2025)
BSODiag: A Global Diagnosis Framework for Batch Servers Outage in Large-scale Cloud Infrastructure Systems
von: Duan, Tao, et al.
Veröffentlicht: (2025)
von: Duan, Tao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Space-Time Trade-off in Bounded Iterated Memory
von: Toyos-Marfurt, Guillermo, et al.
Veröffentlicht: (2025) -
Paradigm Shift in Infrastructure Inspection Technology: Leveraging High-performance Imaging and Advanced AI Analytics to Inspect Road Infrastructure
von: Wu, Du, et al.
Veröffentlicht: (2025) -
Performance Trade-offs of High Order Meshless Approximation on Distributed Memory Systems
von: Vehovar, Jon, et al.
Veröffentlicht: (2025) -
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024) -
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
von: Wang, Tianze, et al.
Veröffentlicht: (2025)