Vector-Centric Machine Learning Systems: A Cross-Stack Approach
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Jiang, Wenqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
von: Yang, Peiming, et al.
Veröffentlicht: (2025)
von: Yang, Peiming, et al.
Veröffentlicht: (2025)
Experimental Assessment of Containers Running on Top of Virtual Machines
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
von: Wang, Xi, et al.
Veröffentlicht: (2024)
von: Wang, Xi, et al.
Veröffentlicht: (2024)
GigaAPI for GPU Parallelization
von: Suvarna, M., et al.
Veröffentlicht: (2025)
von: Suvarna, M., et al.
Veröffentlicht: (2025)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
UPMEM Unleashed: Software Secrets for Speed
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
von: Luo, Weile, et al.
Veröffentlicht: (2025)
von: Luo, Weile, et al.
Veröffentlicht: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
Parallelizing a modern GPU simulator
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
Can Asymmetric Tile Buffering Be Beneficial?
von: Wang, Chengyue, et al.
Veröffentlicht: (2025)
von: Wang, Chengyue, et al.
Veröffentlicht: (2025)
Exploiting long vectors with a CFD code: a co-design show case
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
Simopt -- Simulation pass for Speculative Optimisation of FPGA-CAD flow
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2024)
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2024)
Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry
von: Chen, Ziji, et al.
Veröffentlicht: (2025)
von: Chen, Ziji, et al.
Veröffentlicht: (2025)
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
von: Kumar, Deepak, et al.
Veröffentlicht: (2025)
von: Kumar, Deepak, et al.
Veröffentlicht: (2025)
PIMDAL: Mitigating the Memory Bottleneck in Data Analytics using a Real Processing-in-Memory System
von: Frouzakis, Manos, et al.
Veröffentlicht: (2025)
von: Frouzakis, Manos, et al.
Veröffentlicht: (2025)
NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing
von: Zou, Cheng, et al.
Veröffentlicht: (2026)
von: Zou, Cheng, et al.
Veröffentlicht: (2026)
Design and Implementation of an IoT Cluster with Raspberry Pi Powered by Solar Energy: A Theoretical Approach
von: Portillo, Noel
Veröffentlicht: (2025)
von: Portillo, Noel
Veröffentlicht: (2025)
Efficient Batch Search Algorithm for B+ Tree Index Structures with Level-Wise Traversal on FPGAs
von: Tzschoppe, Max, et al.
Veröffentlicht: (2026)
von: Tzschoppe, Max, et al.
Veröffentlicht: (2026)
ZipFlow: a Compiler-based Framework to Unleash Compressed Data Movement for Modern GPUs
von: Yeo, Gwangoo, et al.
Veröffentlicht: (2026)
von: Yeo, Gwangoo, et al.
Veröffentlicht: (2026)
ARCAS: Adaptive Runtime System for Chiplet-Aware Scheduling
von: Fogli, Alessandro, et al.
Veröffentlicht: (2025)
von: Fogli, Alessandro, et al.
Veröffentlicht: (2025)
The Hitchhiker's Guide to Programming and Optimizing Cache Coherent Heterogeneous Systems: CXL, NVLink-C2C, and AMD Infinity Fabric
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
Massimult: A Novel Parallel CPU Architecture Based on Combinator Reduction
von: Nicklisch-Franken, Jurgen, et al.
Veröffentlicht: (2024)
von: Nicklisch-Franken, Jurgen, et al.
Veröffentlicht: (2024)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
TAPA: A Scalable Task-Parallel Dataflow Programming Framework for Modern FPGAs with Co-Optimization of HLS and Physical Design
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
von: Guo, Licheng, et al.
Veröffentlicht: (2022)
Spira: Exploiting Voxel Data Structural Properties for Efficient Sparse Convolution in Point Cloud Networks
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
Random Adaptive Cache Placement Policy
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using $\mathbb{F}_2$
von: Zhou, Keren, et al.
Veröffentlicht: (2025)
von: Zhou, Keren, et al.
Veröffentlicht: (2025)
ZKProphet: Understanding Performance of Zero-Knowledge Proofs on GPUs
von: Verma, Tarunesh, et al.
Veröffentlicht: (2025)
von: Verma, Tarunesh, et al.
Veröffentlicht: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
von: Lei, Jianlong, et al.
Veröffentlicht: (2026)
von: Lei, Jianlong, et al.
Veröffentlicht: (2026)
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
von: Chen, Qian, et al.
Veröffentlicht: (2024)
von: Chen, Qian, et al.
Veröffentlicht: (2024)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
von: Ruiz, Christian Guzman, et al.
Veröffentlicht: (2024)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
von: Renney, Harri, et al.
Veröffentlicht: (2026)
von: Renney, Harri, et al.
Veröffentlicht: (2026)
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
von: Yang, Peiming, et al.
Veröffentlicht: (2025) -
Experimental Assessment of Containers Running on Top of Virtual Machines
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024) -
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
von: Wang, Xi, et al.
Veröffentlicht: (2024) -
GigaAPI for GPU Parallelization
von: Suvarna, M., et al.
Veröffentlicht: (2025) -
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)