Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Jain, Rishabh, Bhasi, Vivek M., Jog, Adwait, Sivasubramaniam, Anand, Kandemir, Mahmut T., Das, Chita R. |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
JSPIM: A Skew-Aware PIM Accelerator for High-Performance Databases Join and Select Operations
par: Tajdari, Sabiha, et autres
Publié: (2025)
par: Tajdari, Sabiha, et autres
Publié: (2025)
GreenMalloc: Allocator Optimisation for Industrial Workloads
par: Dakhama, Aidan, et autres
Publié: (2025)
par: Dakhama, Aidan, et autres
Publié: (2025)
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
par: Mukunoki, Daichi
Publié: (2025)
par: Mukunoki, Daichi
Publié: (2025)
Bridging the Gap: Physical PCI Device Integration Into SystemC-TLM Virtual Platforms
par: Bosbach, Nils, et autres
Publié: (2025)
par: Bosbach, Nils, et autres
Publié: (2025)
Selective Parallel Loading of Large-Scale Compressed Graphs with ParaGrapher
par: Esfahani, Mohsen Koohi, et autres
Publié: (2024)
par: Esfahani, Mohsen Koohi, et autres
Publié: (2024)
Don't Persist All : Efficient Persistent Data Structures
par: Mahapatra, Pratyush, et autres
Publié: (2019)
par: Mahapatra, Pratyush, et autres
Publié: (2019)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
par: Qararyah, Fareed, et autres
Publié: (2024)
par: Qararyah, Fareed, et autres
Publié: (2024)
Vector-Centric Machine Learning Systems: A Cross-Stack Approach
par: Jiang, Wenqi
Publié: (2025)
par: Jiang, Wenqi
Publié: (2025)
Toward Capturing Genetic Epistasis From Multivariate Genome-Wide Association Studies Using Mixed-Precision Kernel Ridge Regression
par: Ltaief, Hatem, et autres
Publié: (2024)
par: Ltaief, Hatem, et autres
Publié: (2024)
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
par: Wadle, Shayne, et autres
Publié: (2025)
par: Wadle, Shayne, et autres
Publié: (2025)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
par: Liu, Songze, et autres
Publié: (2025)
par: Liu, Songze, et autres
Publié: (2025)
Towards CPU Performance Prediction: New Challenge Benchmark Dataset and Novel Approach
par: Liu, Xiaoman
Publié: (2024)
par: Liu, Xiaoman
Publié: (2024)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
par: Mazzola, Sergio, et autres
Publié: (2025)
par: Mazzola, Sergio, et autres
Publié: (2025)
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
par: Zhang, Kaixuan, et autres
Publié: (2026)
par: Zhang, Kaixuan, et autres
Publié: (2026)
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation
par: Ke, Chih-Hua
Publié: (2026)
par: Ke, Chih-Hua
Publié: (2026)
ACALSim: A Scalable Parallel Simulation Framework for High-Performance System Design Space Exploration
par: Lin, Wei-Fen, et autres
Publié: (2026)
par: Lin, Wei-Fen, et autres
Publié: (2026)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
par: Patwari, Rajeev, et autres
Publié: (2025)
par: Patwari, Rajeev, et autres
Publié: (2025)
Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference
par: Javat, Abdurrahman, et autres
Publié: (2026)
par: Javat, Abdurrahman, et autres
Publié: (2026)
PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System Inferences
par: Huo, Pingyi, et autres
Publié: (2024)
par: Huo, Pingyi, et autres
Publié: (2024)
Characterizing and Understanding HGNN Training on GPUs
par: Han, Dengke, et autres
Publié: (2024)
par: Han, Dengke, et autres
Publié: (2024)
Optimized thread-block arrangement in a GPU implementation of a linear solver for atmospheric chemistry mechanisms
par: Ruiz, Christian Guzman, et autres
Publié: (2024)
par: Ruiz, Christian Guzman, et autres
Publié: (2024)
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
par: Zhang, Yongkang, et autres
Publié: (2024)
par: Zhang, Yongkang, et autres
Publié: (2024)
ZKProphet: Understanding Performance of Zero-Knowledge Proofs on GPUs
par: Verma, Tarunesh, et autres
Publié: (2025)
par: Verma, Tarunesh, et autres
Publié: (2025)
Salient Store: Enabling Smart Storage for Continuous Learning Edge Servers
par: Mishra, Cyan Subhra, et autres
Publié: (2024)
par: Mishra, Cyan Subhra, et autres
Publié: (2024)
Heterogeneous Memory Benchmarking Toolkit
par: Ghaemi, Golsana, et autres
Publié: (2025)
par: Ghaemi, Golsana, et autres
Publié: (2025)
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
par: Ali, Wajid, et autres
Publié: (2025)
par: Ali, Wajid, et autres
Publié: (2025)
Enhancing Instruction Prefetching via Cache and TLB Management
par: Jamet, Alexandre Valentin, et autres
Publié: (2026)
par: Jamet, Alexandre Valentin, et autres
Publié: (2026)
ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
par: Zuepke, Alexander, et autres
Publié: (2026)
par: Zuepke, Alexander, et autres
Publié: (2026)
Recurrent CircuitSAT Sampling for Sequential Circuits
par: Ardakani, Arash, et autres
Publié: (2025)
par: Ardakani, Arash, et autres
Publié: (2025)
Introducing the Arm-membench Throughput Benchmark
par: Burth, Cyrill, et autres
Publié: (2025)
par: Burth, Cyrill, et autres
Publié: (2025)
Enhancing software-hardware co-design for HEP by low-overhead profiling of single- and multi-threaded programs on diverse architectures with Adaptyst
par: Graczyk, Maksymilian, et autres
Publié: (2025)
par: Graczyk, Maksymilian, et autres
Publié: (2025)
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison
par: Li, Ruihao, et autres
Publié: (2026)
par: Li, Ruihao, et autres
Publié: (2026)
Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures
par: Anik, Shafayat Mowla, et autres
Publié: (2026)
par: Anik, Shafayat Mowla, et autres
Publié: (2026)
Makinote: An FPGA-Based HW/SW Platform for Pre-Silicon Emulation of RISC-V Designs
par: Perdomo, Elias, et autres
Publié: (2024)
par: Perdomo, Elias, et autres
Publié: (2024)
AI Load Dynamics--A Power Electronics Perspective
par: Li, Yuzhuo, et autres
Publié: (2025)
par: Li, Yuzhuo, et autres
Publié: (2025)
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
par: Ham, Hyungkyu, et autres
Publié: (2024)
par: Ham, Hyungkyu, et autres
Publié: (2024)
LightningSimV2: Faster and Scalable Simulation for High-Level Synthesis via Graph Compilation and Optimization
par: Sarkar, Rishov, et autres
Publié: (2024)
par: Sarkar, Rishov, et autres
Publié: (2024)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
par: Bergach, Mohamed Amine
Publié: (2026)
par: Bergach, Mohamed Amine
Publié: (2026)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
par: Mao, Shunyu, et autres
Publié: (2024)
par: Mao, Shunyu, et autres
Publié: (2024)
OPTIMA: Design-Space Exploration of Discharge-Based In-SRAM Computing: Quantifying Energy-Accuracy Trade-Offs
par: Seyedfaraji, Saeed, et autres
Publié: (2024)
par: Seyedfaraji, Saeed, et autres
Publié: (2024)
Documents similaires
-
JSPIM: A Skew-Aware PIM Accelerator for High-Performance Databases Join and Select Operations
par: Tajdari, Sabiha, et autres
Publié: (2025) -
GreenMalloc: Allocator Optimisation for Industrial Workloads
par: Dakhama, Aidan, et autres
Publié: (2025) -
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
par: Mukunoki, Daichi
Publié: (2025) -
Bridging the Gap: Physical PCI Device Integration Into SystemC-TLM Virtual Platforms
par: Bosbach, Nils, et autres
Publié: (2025) -
Selective Parallel Loading of Large-Scale Compressed Graphs with ParaGrapher
par: Esfahani, Mohsen Koohi, et autres
Publié: (2024)