The DMA Streaming Framework: Kernel-Level Buffer Orchestration for High-Performance AI Data Paths
Fuente:
arXiv
Guardado en:
| Autor principal: | Graziano, Marco |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
por: Pati, Suchita, et al.
Publicado: (2025)
por: Pati, Suchita, et al.
Publicado: (2025)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
por: Deng, Yunhao, et al.
Publicado: (2025)
por: Deng, Yunhao, et al.
Publicado: (2025)
Intent-Driven Storage Systems: From Low-Level Tuning to High-Level Understanding
por: Bergman, Shai, et al.
Publicado: (2025)
por: Bergman, Shai, et al.
Publicado: (2025)
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
por: Zhou, Zhongchun, et al.
Publicado: (2025)
por: Zhou, Zhongchun, et al.
Publicado: (2025)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
por: Agrawal, Anirudha, et al.
Publicado: (2024)
por: Agrawal, Anirudha, et al.
Publicado: (2024)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
por: Stojkovic, Jovan, et al.
Publicado: (2025)
por: Stojkovic, Jovan, et al.
Publicado: (2025)
IOMMU Support for Virtual-Address Remote DMA in an ARMv8 environment
por: Psistakis, Antonis
Publicado: (2025)
por: Psistakis, Antonis
Publicado: (2025)
Improving AI Efficiency in Data Centres by Power Dynamic Response
por: Marinoni, Andrea, et al.
Publicado: (2025)
por: Marinoni, Andrea, et al.
Publicado: (2025)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
por: Li, Jonathan, et al.
Publicado: (2025)
por: Li, Jonathan, et al.
Publicado: (2025)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
por: Kong, Fanchen, et al.
Publicado: (2025)
por: Kong, Fanchen, et al.
Publicado: (2025)
Exploring energy consumption of AI frameworks on a 64-core RV64 Server CPU
por: Malenza, Giulio, et al.
Publicado: (2025)
por: Malenza, Giulio, et al.
Publicado: (2025)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
por: Li, Jiamin, et al.
Publicado: (2025)
por: Li, Jiamin, et al.
Publicado: (2025)
Power Stabilization for AI Training Datacenters
por: Choukse, Esha, et al.
Publicado: (2025)
por: Choukse, Esha, et al.
Publicado: (2025)
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
por: Zhao, Yiren, et al.
Publicado: (2026)
por: Zhao, Yiren, et al.
Publicado: (2026)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
por: Zhao, Dan, et al.
Publicado: (2024)
por: Zhao, Dan, et al.
Publicado: (2024)
Debunking the CUDA Myth Towards GPU-based AI Systems
por: Lee, Yunjae, et al.
Publicado: (2024)
por: Lee, Yunjae, et al.
Publicado: (2024)
Strict Partitioning for Sporadic Rigid Gang Tasks
por: Sun, Binqi, et al.
Publicado: (2024)
por: Sun, Binqi, et al.
Publicado: (2024)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
por: Stojkovic, Jovan, et al.
Publicado: (2024)
por: Stojkovic, Jovan, et al.
Publicado: (2024)
Modernizing Amdahl's Law: How AI Scaling Laws Shape Computer Architecture
por: Lu, Chien-Ping
Publicado: (2026)
por: Lu, Chien-Ping
Publicado: (2026)
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
por: Peccia, Federico Nicolas, et al.
Publicado: (2024)
por: Peccia, Federico Nicolas, et al.
Publicado: (2024)
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
por: Canakci, Burcu, et al.
Publicado: (2025)
por: Canakci, Burcu, et al.
Publicado: (2025)
RedFuser: An Automatic Operator Fusion Framework for Cascaded Reductions on AI Accelerators
por: Tang, Xinsheng, et al.
Publicado: (2026)
por: Tang, Xinsheng, et al.
Publicado: (2026)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
por: Pan, Yudong, et al.
Publicado: (2026)
por: Pan, Yudong, et al.
Publicado: (2026)
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
por: Cao, Yingqi, et al.
Publicado: (2024)
por: Cao, Yingqi, et al.
Publicado: (2024)
CELLO: Co-designing Schedule and Hybrid Implicit/Explicit Buffer for Complex Tensor Reuse
por: Garg, Raveesh, et al.
Publicado: (2023)
por: Garg, Raveesh, et al.
Publicado: (2023)
Can Asymmetric Tile Buffering Be Beneficial?
por: Wang, Chengyue, et al.
Publicado: (2025)
por: Wang, Chengyue, et al.
Publicado: (2025)
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
por: Chen, Jiesong, et al.
Publicado: (2026)
por: Chen, Jiesong, et al.
Publicado: (2026)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
por: Liu, Fangxin, et al.
Publicado: (2026)
por: Liu, Fangxin, et al.
Publicado: (2026)
ODIN-Based CPU-GPU Architecture with Replay-Driven Simulation and Emulation
por: Dorairaj, Nij, et al.
Publicado: (2026)
por: Dorairaj, Nij, et al.
Publicado: (2026)
NPU Design for Diffusion Language Model Inference
por: Lou, Binglei, et al.
Publicado: (2026)
por: Lou, Binglei, et al.
Publicado: (2026)
Forge-UGC: FX optimization and register-graph engine for universal graph compiler
por: Kumar, Satyam, et al.
Publicado: (2026)
por: Kumar, Satyam, et al.
Publicado: (2026)
PhD Thesis Summary: Methods for Reliability Assessment and Enhancement of Deep Neural Network Hardware Accelerators
por: Taheri, Mahdi
Publicado: (2026)
por: Taheri, Mahdi
Publicado: (2026)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
por: DeBole, Michael V., et al.
Publicado: (2025)
por: DeBole, Michael V., et al.
Publicado: (2025)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
por: Qin, Ruoyu, et al.
Publicado: (2024)
por: Qin, Ruoyu, et al.
Publicado: (2024)
PiKV: KV Cache Management System for Mixture of Experts
por: Liu, Dong, et al.
Publicado: (2025)
por: Liu, Dong, et al.
Publicado: (2025)
Co-design of a novel CMOS highly parallel, low-power, multi-chip neural network accelerator
por: Hokenmaier, W, et al.
Publicado: (2024)
por: Hokenmaier, W, et al.
Publicado: (2024)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025)
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025)
Investigating Memory Failure Prediction Across CPU Architectures
por: Yu, Qiao, et al.
Publicado: (2024)
por: Yu, Qiao, et al.
Publicado: (2024)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
por: Kubwimana, Benjamin, et al.
Publicado: (2025)
por: Kubwimana, Benjamin, et al.
Publicado: (2025)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
por: Zhu, Wenbin, et al.
Publicado: (2025)
por: Zhu, Wenbin, et al.
Publicado: (2025)
Ejemplares similares
-
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
por: Pati, Suchita, et al.
Publicado: (2025) -
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
por: Deng, Yunhao, et al.
Publicado: (2025) -
Intent-Driven Storage Systems: From Low-Level Tuning to High-Level Understanding
por: Bergman, Shai, et al.
Publicado: (2025) -
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
por: Zhou, Zhongchun, et al.
Publicado: (2025) -
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
por: Agrawal, Anirudha, et al.
Publicado: (2024)