Towards Performance-Aware Allocation for Accelerated Machine Learning on GPU-SSD Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gundawar, Ayush, Chung, Euijun, Kim, Hyesoon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
von: Chung, Euijun, et al.
Veröffentlicht: (2026)
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
von: Pu, Huanzhi, et al.
Veröffentlicht: (2025)
von: Pu, Huanzhi, et al.
Veröffentlicht: (2025)
CuLifter: Lifting GPU Binaries to Typed IR
von: Zhao, Jisheng, et al.
Veröffentlicht: (2026)
von: Zhao, Jisheng, et al.
Veröffentlicht: (2026)
CXL Topology-Aware and Expander-Driven Prefetching: Unlocking SSD Performance
von: Oh, Dongsuk, et al.
Veröffentlicht: (2025)
von: Oh, Dongsuk, et al.
Veröffentlicht: (2025)
A Host-SSD Collaborative Write Accelerator for LSM-Tree-Based Key-Value Stores
von: Kim, KiHwan, et al.
Veröffentlicht: (2024)
von: Kim, KiHwan, et al.
Veröffentlicht: (2024)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
DCI: A Coordinated Allocation and Filling Workload-Aware Dual-Cache Allocation GNN Inference Acceleration System
von: Luo, Yi, et al.
Veröffentlicht: (2025)
von: Luo, Yi, et al.
Veröffentlicht: (2025)
Block-SSD: A New Block-Based Blocking SSD Architecture
von: Wong, Ryan, et al.
Veröffentlicht: (2024)
von: Wong, Ryan, et al.
Veröffentlicht: (2024)
Containerized In-Storage Processing and Computing-Enabled SSD Disaggregation
von: Kwon, Miryeong, et al.
Veröffentlicht: (2025)
von: Kwon, Miryeong, et al.
Veröffentlicht: (2025)
GAP-LA: GPU-Accelerated Performance-Driven Layer Assignment
von: Zhao, Chunyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Chunyuan, et al.
Veröffentlicht: (2025)
A Full-System Simulation Framework for CXL-Based SSD Memory System
von: Wang, Yaohui, et al.
Veröffentlicht: (2025)
von: Wang, Yaohui, et al.
Veröffentlicht: (2025)
Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
von: Majeed, Ashiyana Abdul, et al.
Veröffentlicht: (2025)
GPU-Accelerated Simulated Oscillator Ising/Potts Machine Solving Combinatorial Optimization Problems
von: Gonul, Yilmaz Ege, et al.
Veröffentlicht: (2025)
von: Gonul, Yilmaz Ege, et al.
Veröffentlicht: (2025)
TCAM-SSD: A Framework for Search-Based Computing in Solid-State Drives
von: Wong, Ryan, et al.
Veröffentlicht: (2024)
von: Wong, Ryan, et al.
Veröffentlicht: (2024)
RoboGPU: Accelerating GPU Collision Detection for Robotics
von: Liu, Lufei, et al.
Veröffentlicht: (2026)
von: Liu, Lufei, et al.
Veröffentlicht: (2026)
Search-in-Memory (SiM): Reliable, Versatile, and Efficient Data Matching in SSD's NAND Flash Memory Chip for Data Indexing Acceleration
von: Chen, Yun-Chih, et al.
Veröffentlicht: (2024)
von: Chen, Yun-Chih, et al.
Veröffentlicht: (2024)
Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
von: Kyung, Kwanhee, et al.
Veröffentlicht: (2025)
von: Kyung, Kwanhee, et al.
Veröffentlicht: (2025)
DG-RePlAce: A Dataflow-Driven GPU-Accelerated Analytical Global Placement Framework for Machine Learning Accelerators
von: Kahng, Andrew B., et al.
Veröffentlicht: (2024)
von: Kahng, Andrew B., et al.
Veröffentlicht: (2024)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node
von: Latif, Imran, et al.
Veröffentlicht: (2024)
von: Latif, Imran, et al.
Veröffentlicht: (2024)
SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-design
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
Energy-Aware Heterogeneous Federated Learning via Approximate DNN Accelerators
von: Pfeiffer, Kilian, et al.
Veröffentlicht: (2024)
von: Pfeiffer, Kilian, et al.
Veröffentlicht: (2024)
Adapting Atmospheric Chemistry Components for Efficient GPU Accelerators
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
HillInfer: Efficient Long-Context LLM Inference on the Edge with Hierarchical KV Eviction using SmartSSD
von: Sun, He, et al.
Veröffentlicht: (2026)
von: Sun, He, et al.
Veröffentlicht: (2026)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
von: Zhang, Kaixuan, et al.
Veröffentlicht: (2026)
von: Zhang, Kaixuan, et al.
Veröffentlicht: (2026)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
von: Seo, Minseok, et al.
Veröffentlicht: (2024)
von: Seo, Minseok, et al.
Veröffentlicht: (2024)
GME: GPU-based Microarchitectural Extensions to Accelerate Homomorphic Encryption
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2023)
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2023)
The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling Guidance
von: Tung, Chung-Hsuan, et al.
Veröffentlicht: (2026)
von: Tung, Chung-Hsuan, et al.
Veröffentlicht: (2026)
STAR: Improving Lifetime and Performance of High-Capacity Modern SSDs Using State-Aware Randomizer
von: Kwon, Omin, et al.
Veröffentlicht: (2025)
von: Kwon, Omin, et al.
Veröffentlicht: (2025)
LLM-PRISM: Characterizing Silent Data Corruption from Permanent GPU Faults in LLM Training
von: Tyagi, Abhishek, et al.
Veröffentlicht: (2026)
von: Tyagi, Abhishek, et al.
Veröffentlicht: (2026)
Hardware-Aware Neural Network Compilation with Learned Optimization: A RISC-V Accelerator Approach
von: Ganti, Ravindra, et al.
Veröffentlicht: (2025)
von: Ganti, Ravindra, et al.
Veröffentlicht: (2025)
LLC Intra-set Write Balancing
von: Krishna, Keshav, et al.
Veröffentlicht: (2024)
von: Krishna, Keshav, et al.
Veröffentlicht: (2024)
Stream-HLS: Towards Automatic Dataflow Acceleration
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
von: Basalama, Suhail, et al.
Veröffentlicht: (2025)
FlexNeRFer: A Multi-Dataflow, Adaptive Sparsity-Aware Accelerator for On-Device NeRF Rendering
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
von: Fan, Zehao, et al.
Veröffentlicht: (2025)
von: Fan, Zehao, et al.
Veröffentlicht: (2025)
Modeling and Optimizing Performance Bottlenecks for Neuromorphic Accelerators
von: Yik, Jason, et al.
Veröffentlicht: (2025)
von: Yik, Jason, et al.
Veröffentlicht: (2025)
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
von: Chung, Euijun, et al.
Veröffentlicht: (2026) -
Hardware vs. Software Implementation of Warp-Level Features in Vortex RISC-V GPU
von: Pu, Huanzhi, et al.
Veröffentlicht: (2025) -
CuLifter: Lifting GPU Binaries to Typed IR
von: Zhao, Jisheng, et al.
Veröffentlicht: (2026) -
CXL Topology-Aware and Expander-Driven Prefetching: Unlocking SSD Performance
von: Oh, Dongsuk, et al.
Veröffentlicht: (2025) -
A Host-SSD Collaborative Write Accelerator for LSM-Tree-Based Key-Value Stores
von: Kim, KiHwan, et al.
Veröffentlicht: (2024)