DX100: A Programmable Data Access Accelerator for Indirection
Fuente:
arXiv
Saved in:
| Main Authors: | Khadem, Alireza, Kamalakkannan, Kamalavasan, Zhu, Zhenyan, Poptani, Akash, Gu, Yufeng, Dominguez-Trujillo, Jered Benjamin, Talati, Nishil, Fujiki, Daichi, Mahlke, Scott, Shipman, Galen, Das, Reetuparna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Dimensional Vector ISA Extension for Mobile In-Cache Computing
by: Khadem, Alireza, et al.
Published: (2025)
by: Khadem, Alireza, et al.
Published: (2025)
Characterizing Adaptive Mesh Refinement on Heterogeneous Platforms with Parthenon-VIBE
by: Poptani, Akash, et al.
Published: (2025)
by: Poptani, Akash, et al.
Published: (2025)
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference
by: Gu, Yufeng, et al.
Published: (2025)
by: Gu, Yufeng, et al.
Published: (2025)
Toward Cross-Layer Energy Optimizations in AI Systems
by: Chung, Jae-Won, et al.
Published: (2024)
by: Chung, Jae-Won, et al.
Published: (2024)
From Legacy Fortran to Portable Kokkos: An Autonomous Agentic AI Workflow
by: Gupta, Sparsh, et al.
Published: (2025)
by: Gupta, Sparsh, et al.
Published: (2025)
Bridging Simulation and Silicon: A Study of RISC-V Hardware and FireSim Simulation
by: Barai, Atanu, et al.
Published: (2025)
by: Barai, Atanu, et al.
Published: (2025)
ZKProphet: Understanding Performance of Zero-Knowledge Proofs on GPUs
by: Verma, Tarunesh, et al.
Published: (2025)
by: Verma, Tarunesh, et al.
Published: (2025)
LAAFD: LLM-based Agents for Accelerated FPGA Design
by: Moraru, Maxim, et al.
Published: (2026)
by: Moraru, Maxim, et al.
Published: (2026)
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization
by: Matsushima, Kosuke, et al.
Published: (2026)
by: Matsushima, Kosuke, et al.
Published: (2026)
A Customized Memory-aware Architecture for Biological Sequence Alignment
by: Akbari, Nasrin, et al.
Published: (2025)
by: Akbari, Nasrin, et al.
Published: (2025)
Palermo: Improving the Performance of Oblivious Memory using Protocol-Hardware Co-Design
by: Ye, Haojie, et al.
Published: (2024)
by: Ye, Haojie, et al.
Published: (2024)
NMP-PaK: Near-Memory Processing Acceleration of Scalable De Novo Genome Assembly
by: Kim, Heewoo, et al.
Published: (2025)
by: Kim, Heewoo, et al.
Published: (2025)
SNIP: An Adaptive Mixed Precision Framework for Subbyte Large Language Model Training
by: Pan, Yunjie, et al.
Published: (2026)
by: Pan, Yunjie, et al.
Published: (2026)
Pickle Prefetcher: Programmable and Scalable Last-Level Cache Prefetcher
by: Nguyen, Hoa, et al.
Published: (2025)
by: Nguyen, Hoa, et al.
Published: (2025)
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
by: Mukunoki, Daichi
Published: (2025)
by: Mukunoki, Daichi
Published: (2025)
Resilient and Secure Programmable System-on-Chip Accelerator Offload
by: Gouveia, Inês Pinto, et al.
Published: (2024)
by: Gouveia, Inês Pinto, et al.
Published: (2024)
Fletch: File-System Metadata Caching in Programmable Switches
by: Liu, Qingxiu, et al.
Published: (2025)
by: Liu, Qingxiu, et al.
Published: (2025)
Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory Access
by: Wang, Luming, et al.
Published: (2024)
by: Wang, Luming, et al.
Published: (2024)
Towards Generalized On-Chip Communication for Programmable Accelerators in Heterogeneous Architectures
by: Zuckerman, Joseph, et al.
Published: (2024)
by: Zuckerman, Joseph, et al.
Published: (2024)
Scalable and RISC-V Programmable Near-Memory Computing Architectures for Edge Nodes
by: Caon, Michele, et al.
Published: (2024)
by: Caon, Michele, et al.
Published: (2024)
Field-Programmable Gate Array Architecture for Deep Learning: Survey & Future Directions
by: Boutros, Andrew, et al.
Published: (2024)
by: Boutros, Andrew, et al.
Published: (2024)
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
by: Wu, Jiajun, et al.
Published: (2024)
by: Wu, Jiajun, et al.
Published: (2024)
UpDown: Programmable fine-grained Events for Scalable Performance on Irregular Applications
by: Rajasukumar, Andronicus, et al.
Published: (2024)
by: Rajasukumar, Andronicus, et al.
Published: (2024)
Accelerating PageRank Algorithmic Tasks with a new Programmable Hardware Architecture
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2024)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2024)
MCEL: Margin-Based Cross-Entropy Loss for Error-Tolerant Quantized Neural Networks
by: Yayla, Mikail, et al.
Published: (2026)
by: Yayla, Mikail, et al.
Published: (2026)
NX-CGRA: A Programmable Hardware Accelerator for Core Transformer Algorithms on Edge Devices
by: Prasad, Rohit
Published: (2025)
by: Prasad, Rohit
Published: (2025)
COOK Access Control on an embedded Volta GPU
by: Lesage, Benjamin, et al.
Published: (2024)
by: Lesage, Benjamin, et al.
Published: (2024)
FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
by: Rahoof, Abdul, et al.
Published: (2025)
by: Rahoof, Abdul, et al.
Published: (2025)
SPPAM: Signature Pattern Prediction and Access-Map Prefetcher
by: Merrell, Maccoy, et al.
Published: (2026)
by: Merrell, Maccoy, et al.
Published: (2026)
Towards Forever Access for Implanted Brain-Computer Interfaces
by: Ugur, Muhammed, et al.
Published: (2024)
by: Ugur, Muhammed, et al.
Published: (2024)
ORAP: Optimized Row Access Prefetching for Rowhammer-mitigated Memory
by: Merrell, Maccoy, et al.
Published: (2026)
by: Merrell, Maccoy, et al.
Published: (2026)
ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory Accesses
by: Li, Mengming, et al.
Published: (2026)
by: Li, Mengming, et al.
Published: (2026)
Decoupled Control Flow and Data Access in RISC-V GPGPUs
by: Sarda, Giuseppe M., et al.
Published: (2025)
by: Sarda, Giuseppe M., et al.
Published: (2025)
Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
by: Nguyen, Hoa, et al.
Published: (2025)
by: Nguyen, Hoa, et al.
Published: (2025)
GRAU: Generic Reconfigurable Activation Unit Design for Neural Network Hardware Accelerators
by: Liu, Yuhao, et al.
Published: (2026)
by: Liu, Yuhao, et al.
Published: (2026)
Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators
by: Liu, Yuhao, et al.
Published: (2026)
by: Liu, Yuhao, et al.
Published: (2026)
BiKA: Kolmogorov-Arnold-Network-inspired Ultra Lightweight Neural Network Hardware Accelerator
by: Liu, Yuhao, et al.
Published: (2026)
by: Liu, Yuhao, et al.
Published: (2026)
SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
by: Scheffler, Paul, et al.
Published: (2024)
by: Scheffler, Paul, et al.
Published: (2024)
RoMe: Row Granularity Access Memory System for Large Language Models
by: Nam, Hwayong, et al.
Published: (2025)
by: Nam, Hwayong, et al.
Published: (2025)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Similar Items
-
Multi-Dimensional Vector ISA Extension for Mobile In-Cache Computing
by: Khadem, Alireza, et al.
Published: (2025) -
Characterizing Adaptive Mesh Refinement on Heterogeneous Platforms with Parthenon-VIBE
by: Poptani, Akash, et al.
Published: (2025) -
PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference
by: Gu, Yufeng, et al.
Published: (2025) -
Toward Cross-Layer Energy Optimizations in AI Systems
by: Chung, Jae-Won, et al.
Published: (2024) -
From Legacy Fortran to Portable Kokkos: An Autonomous Agentic AI Workflow
by: Gupta, Sparsh, et al.
Published: (2025)