ZipFlow: a Compiler-based Framework to Unleash Compressed Data Movement for Modern GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yeo, Gwangoo, Shen, Zhiyang, Cui, Wei, Interlandi, Matteo, Sen, Rathijit, Ding, Bailu, Chen, Qi, Rhu, Minsoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
von: Yeo, Gwangoo, et al.
Veröffentlicht: (2024)
von: Yeo, Gwangoo, et al.
Veröffentlicht: (2024)
Kitsune: Enabling Dataflow Execution on GPUs
von: Davies, Michael, et al.
Veröffentlicht: (2025)
von: Davies, Michael, et al.
Veröffentlicht: (2025)
HetGPU: The pursuit of making binary compatibility towards GPUs
von: Yang, Yiwei, et al.
Veröffentlicht: (2025)
von: Yang, Yiwei, et al.
Veröffentlicht: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)
Hierarchical Resource Partitioning on Modern GPUs: A Reinforcement Learning Approach
von: Saroliya, Urvij, et al.
Veröffentlicht: (2024)
von: Saroliya, Urvij, et al.
Veröffentlicht: (2024)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
von: Shi, Tianyao, et al.
Veröffentlicht: (2024)
Exploration of Cryptocurrency Mining-Specific GPUs in AI Applications: A Case Study of CMP 170HX
von: Kangwei, Xing
Veröffentlicht: (2025)
von: Kangwei, Xing
Veröffentlicht: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
A Modern Primer on Processing in Memory
von: Mutlu, Onur, et al.
Veröffentlicht: (2020)
von: Mutlu, Onur, et al.
Veröffentlicht: (2020)
UPMEM Unleashed: Software Secrets for Speed
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing
von: Zou, Cheng, et al.
Veröffentlicht: (2026)
von: Zou, Cheng, et al.
Veröffentlicht: (2026)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
von: Deng, Yunhao, et al.
Veröffentlicht: (2025)
von: Deng, Yunhao, et al.
Veröffentlicht: (2025)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
Debunking the CUDA Myth Towards GPU-based AI Systems
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
von: Kong, Fanchen, et al.
Veröffentlicht: (2025)
Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -- A Critical Review
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
Implementation and Evaluation of GBDI Memory Compression Algorithm Using C/C++ on a Broader Range of Workloads
von: Aina, Adeyemi
Veröffentlicht: (2025)
von: Aina, Adeyemi
Veröffentlicht: (2025)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
von: Qiu, Tong Dong, et al.
Veröffentlicht: (2023)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
von: Kubwimana, Benjamin, et al.
Veröffentlicht: (2025)
von: Kubwimana, Benjamin, et al.
Veröffentlicht: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
von: li, Fei, et al.
Veröffentlicht: (2026)
von: li, Fei, et al.
Veröffentlicht: (2026)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
von: Canakci, Burcu, et al.
Veröffentlicht: (2025)
von: Canakci, Burcu, et al.
Veröffentlicht: (2025)
A Reliable, Time-Predictable Heterogeneous SoC for AI-Enhanced Mixed-Criticality Edge Applications
von: Garofalo, Angelo, et al.
Veröffentlicht: (2025)
von: Garofalo, Angelo, et al.
Veröffentlicht: (2025)
Atomique: A Quantum Compiler for Reconfigurable Neutral Atom Arrays
von: Wang, Hanrui, et al.
Veröffentlicht: (2023)
von: Wang, Hanrui, et al.
Veröffentlicht: (2023)
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
von: Vo, Huynh Q. N., et al.
Veröffentlicht: (2025)
von: Vo, Huynh Q. N., et al.
Veröffentlicht: (2025)
ZKProphet: Understanding Performance of Zero-Knowledge Proofs on GPUs
von: Verma, Tarunesh, et al.
Veröffentlicht: (2025)
von: Verma, Tarunesh, et al.
Veröffentlicht: (2025)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
von: Liu, Yiqi, et al.
Veröffentlicht: (2025)
von: Liu, Yiqi, et al.
Veröffentlicht: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
von: Kreuzer, Anke, et al.
Veröffentlicht: (2019)
von: Kreuzer, Anke, et al.
Veröffentlicht: (2019)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
von: Negi, Shubham, et al.
Veröffentlicht: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
von: Kubo, Tatsuya, et al.
Veröffentlicht: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
von: Liu, Lian, et al.
Veröffentlicht: (2026)
von: Liu, Lian, et al.
Veröffentlicht: (2026)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
von: Kurzynski, Marco, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026) -
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
von: Yeo, Gwangoo, et al.
Veröffentlicht: (2024) -
Kitsune: Enabling Dataflow Execution on GPUs
von: Davies, Michael, et al.
Veröffentlicht: (2025) -
HetGPU: The pursuit of making binary compatibility towards GPUs
von: Yang, Yiwei, et al.
Veröffentlicht: (2025) -
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
von: Zhang, Qijun, et al.
Veröffentlicht: (2026)