Enabling full-speed random access to the entire memory on the A100 GPU
Fuente:
arXiv
Saved in:
| Main Author: | Walker, Alden |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Per-Access Upper Bound for Shared-Resource Interference in Direct-Mapped Multicore Architectures
by: Pedroni, Felipe T.
Published: (2026)
by: Pedroni, Felipe T.
Published: (2026)
HERMES: High-Performance RISC-V Memory Hierarchy for ML Workloads
by: Suryadevara, Pranav
Published: (2025)
by: Suryadevara, Pranav
Published: (2025)
CLIPGen: A Chiplet Link IP Modeling and Generation Framework for 2.5D Architecture Exploration
by: Zhu, Zhengping, et al.
Published: (2026)
by: Zhu, Zhengping, et al.
Published: (2026)
Megakernel vs Wavefront GPU Path Tracing
by: Padilla, Rafael, et al.
Published: (2026)
by: Padilla, Rafael, et al.
Published: (2026)
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
by: Palaniappan, Kathiravan
Published: (2026)
by: Palaniappan, Kathiravan
Published: (2026)
GDEV-AI: A Generalized Evaluation of Deep Learning Inference Scaling and Architectural Saturation
by: Palaniappan, Kathiravan
Published: (2026)
by: Palaniappan, Kathiravan
Published: (2026)
Arm DynamIQ Shared Unit and Real-Time: An Empirical Evaluation
by: Pradhan, Ashutosh, et al.
Published: (2025)
by: Pradhan, Ashutosh, et al.
Published: (2025)
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
by: Tahmasebi, Faraz, et al.
Published: (2025)
by: Tahmasebi, Faraz, et al.
Published: (2025)
RowHammer Vulnerability Counter (RVC): Redefining RowHammer Detection with Victim-Centric Tracking
by: Jain, Lavi, et al.
Published: (2026)
by: Jain, Lavi, et al.
Published: (2026)
Benchmarking NIST-Standardised ML-KEM and ML-DSA on ARM Cortex-M0+: Performance, Memory, and Energy on the RP2040
by: Chhetri, Rojin
Published: (2026)
by: Chhetri, Rojin
Published: (2026)
AXI4MLIR: User-Driven Automatic Host Code Generation for Custom AXI-Based Accelerators
by: Agostini, Nicolas Bohm, et al.
Published: (2023)
by: Agostini, Nicolas Bohm, et al.
Published: (2023)
ASTER: Attention-based Spiking Transformer Engine for Event-driven Reasoning
by: Das, Tamoghno, et al.
Published: (2025)
by: Das, Tamoghno, et al.
Published: (2025)
SPEC CPU: The Next Generation
by: Madhav, Mahesh, et al.
Published: (2026)
by: Madhav, Mahesh, et al.
Published: (2026)
Fast NF4 Dequantization Kernels for Large Language Model Inference
by: Qi, Xiangbo, et al.
Published: (2026)
by: Qi, Xiangbo, et al.
Published: (2026)
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
by: Liu, Lian, et al.
Published: (2025)
by: Liu, Lian, et al.
Published: (2025)
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
by: Zhang, Kaixuan, et al.
Published: (2026)
by: Zhang, Kaixuan, et al.
Published: (2026)
Hardware/Algorithm Co-design for Real-Time I/O Control with Improved Timing Accuracy and Robustness
by: Jiang, Zhe, et al.
Published: (2024)
by: Jiang, Zhe, et al.
Published: (2024)
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs
by: Guan, Jiapeng, et al.
Published: (2024)
by: Guan, Jiapeng, et al.
Published: (2024)
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
by: Zhang, Kaixuan, et al.
Published: (2026)
by: Zhang, Kaixuan, et al.
Published: (2026)
Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures
by: Anik, Shafayat Mowla, et al.
Published: (2026)
by: Anik, Shafayat Mowla, et al.
Published: (2026)
Deep Recommender Models Inference: Automatic Asymmetric Data Flow Optimization
by: Ruggeri, Giuseppe, et al.
Published: (2025)
by: Ruggeri, Giuseppe, et al.
Published: (2025)
An Integrated UVM-TLM Co-Simulation Framework for RISC-V Functional Verification and Performance Evaluation
by: Qiu, Ruizhi, et al.
Published: (2025)
by: Qiu, Ruizhi, et al.
Published: (2025)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024)
by: Mo, Zhiwen, et al.
Published: (2024)
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
by: Li, Zhuoran, et al.
Published: (2026)
by: Li, Zhuoran, et al.
Published: (2026)
Efficient and Scalable Architecture for Multiple-chip Implementation of Simulated Bifurcation Machines
by: Kashimata, Tomoya, et al.
Published: (2023)
by: Kashimata, Tomoya, et al.
Published: (2023)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
A Compilation Framework for Quantum Circuits with Mid-Circuit Measurement Error Awareness
by: Zhong, Ming, et al.
Published: (2025)
by: Zhong, Ming, et al.
Published: (2025)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
by: De Sensi, Daniele, et al.
Published: (2024)
by: De Sensi, Daniele, et al.
Published: (2024)
EStacker: Explaining Battery-Less IoT System Performance with Energy Stacks
by: Liedtke, Lukas, et al.
Published: (2025)
by: Liedtke, Lukas, et al.
Published: (2025)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
FREESS: A Web-Based Educational Simulator for a RISC-V-Inspired Superscalar Processor with Tomasulo-Style Dynamic Scheduling
by: Giorgi, Roberto, et al.
Published: (2026)
by: Giorgi, Roberto, et al.
Published: (2026)
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
by: Liu, Qunyou, et al.
Published: (2025)
by: Liu, Qunyou, et al.
Published: (2025)
Taming Wild Branches: Overcoming Hard-to-Predict Branches using the Bullseye Predictor
by: Behrendt, Emet, et al.
Published: (2025)
by: Behrendt, Emet, et al.
Published: (2025)
ArchAgent: Agentic AI-driven Computer Architecture Discovery
by: Gupta, Raghav, et al.
Published: (2026)
by: Gupta, Raghav, et al.
Published: (2026)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
Revelator: Rapid Data Fetching via OS-Driven Hash-based Speculative Address Translation
by: Kanellopoulos, Konstantinos, et al.
Published: (2025)
by: Kanellopoulos, Konstantinos, et al.
Published: (2025)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
by: Abraham, Ojima, et al.
Published: (2026)
by: Abraham, Ojima, et al.
Published: (2026)
MEDEA: A Design-Time Multi-Objective Manager for Energy-Efficient DNN Inference on Heterogeneous Ultra-Low Power Platforms
by: Taji, Hossein, et al.
Published: (2025)
by: Taji, Hossein, et al.
Published: (2025)
AES-RV: Hardware-Efficient RISC-V Accelerator with Low-Latency AES Instruction Extension for IoT Security
by: Nguyen, Van Tinh, et al.
Published: (2025)
by: Nguyen, Van Tinh, et al.
Published: (2025)
pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup Tables
by: Ferreira, João Dinis, et al.
Published: (2021)
by: Ferreira, João Dinis, et al.
Published: (2021)
Similar Items
-
A Per-Access Upper Bound for Shared-Resource Interference in Direct-Mapped Multicore Architectures
by: Pedroni, Felipe T.
Published: (2026) -
HERMES: High-Performance RISC-V Memory Hierarchy for ML Workloads
by: Suryadevara, Pranav
Published: (2025) -
CLIPGen: A Chiplet Link IP Modeling and Generation Framework for 2.5D Architecture Exploration
by: Zhu, Zhengping, et al.
Published: (2026) -
Megakernel vs Wavefront GPU Path Tracing
by: Padilla, Rafael, et al.
Published: (2026) -
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
by: Palaniappan, Kathiravan
Published: (2026)