D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
Fuente:
arXiv
Saved in:
| Main Authors: | Tahmasebi, Faraz, Pelluer, Michael, Kwon, Hyoukjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024)
by: Karami, Rachid, et al.
Published: (2024)
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
by: Palaniappan, Kathiravan
Published: (2026)
by: Palaniappan, Kathiravan
Published: (2026)
Fast NF4 Dequantization Kernels for Large Language Model Inference
by: Qi, Xiangbo, et al.
Published: (2026)
by: Qi, Xiangbo, et al.
Published: (2026)
Design Space Exploration of Approximate Computing Techniques with a Reinforcement Learning Approach
by: Saeedi, Sepide, et al.
Published: (2023)
by: Saeedi, Sepide, et al.
Published: (2023)
A Per-Access Upper Bound for Shared-Resource Interference in Direct-Mapped Multicore Architectures
by: Pedroni, Felipe T.
Published: (2026)
by: Pedroni, Felipe T.
Published: (2026)
Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer
by: Li, Richie, et al.
Published: (2025)
by: Li, Richie, et al.
Published: (2025)
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
by: Odema, Mohanad, et al.
Published: (2024)
by: Odema, Mohanad, et al.
Published: (2024)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
by: Peng, Hongwu, et al.
Published: (2023)
by: Peng, Hongwu, et al.
Published: (2023)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
by: Rout, Nikhil, et al.
Published: (2025)
by: Rout, Nikhil, et al.
Published: (2025)
Factor Machine: Mixed-signal Architecture for Fine-Grained Graph-Based Computing
by: Dudek, Piotr
Published: (2024)
by: Dudek, Piotr
Published: (2024)
GDEV-AI: A Generalized Evaluation of Deep Learning Inference Scaling and Architectural Saturation
by: Palaniappan, Kathiravan
Published: (2026)
by: Palaniappan, Kathiravan
Published: (2026)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
by: Xu, Haocheng, et al.
Published: (2024)
by: Xu, Haocheng, et al.
Published: (2024)
Taming Wild Branches: Overcoming Hard-to-Predict Branches using the Bullseye Predictor
by: Behrendt, Emet, et al.
Published: (2025)
by: Behrendt, Emet, et al.
Published: (2025)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024)
by: Sanovar, Rya, et al.
Published: (2024)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
by: Zhou, Cyrus, et al.
Published: (2023)
by: Zhou, Cyrus, et al.
Published: (2023)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
Dataflow & Tiling Strategies in Edge-AI FPGA Accelerators: A Comprehensive Literature Review
by: Li, Richie
Published: (2025)
by: Li, Richie
Published: (2025)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
by: Saha, Rappy, et al.
Published: (2026)
by: Saha, Rappy, et al.
Published: (2026)
Enabling full-speed random access to the entire memory on the A100 GPU
by: Walker, Alden
Published: (2024)
by: Walker, Alden
Published: (2024)
FREESS: A Web-Based Educational Simulator for a RISC-V-Inspired Superscalar Processor with Tomasulo-Style Dynamic Scheduling
by: Giorgi, Roberto, et al.
Published: (2026)
by: Giorgi, Roberto, et al.
Published: (2026)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
by: Müller, Mika Markus, et al.
Published: (2025)
by: Müller, Mika Markus, et al.
Published: (2025)
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
by: Yuan, Aojie, et al.
Published: (2026)
by: Yuan, Aojie, et al.
Published: (2026)
USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks
by: Ibrahim, Muhammad Sohail, et al.
Published: (2024)
by: Ibrahim, Muhammad Sohail, et al.
Published: (2024)
Toward A Formalized Approach for Spike Sorting Algorithms and Hardware Evaluation
by: Zhang, Tim, et al.
Published: (2022)
by: Zhang, Tim, et al.
Published: (2022)
Search Your Block Floating Point Scales!
by: Gupta, Tanmaey, et al.
Published: (2026)
by: Gupta, Tanmaey, et al.
Published: (2026)
A2Q+: Improving Accumulator-Aware Weight Quantization
by: Colbert, Ian, et al.
Published: (2024)
by: Colbert, Ian, et al.
Published: (2024)
GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
by: Atmer, Hannah, et al.
Published: (2025)
by: Atmer, Hannah, et al.
Published: (2025)
Graph neural networks with configuration cross-attention for tensor compilers
by: Khizbullin, Dmitrii, et al.
Published: (2024)
by: Khizbullin, Dmitrii, et al.
Published: (2024)
Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
by: Siddiqui, Mohammed Humaid, et al.
Published: (2025)
by: Siddiqui, Mohammed Humaid, et al.
Published: (2025)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
by: Lübeck, Konstantin, et al.
Published: (2024)
by: Lübeck, Konstantin, et al.
Published: (2024)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
by: AbouElhamayed, Ahmed F., et al.
Published: (2025)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
by: Jung, Alexander Louis-Ferdinand, et al.
Published: (2024)
Lincoln AI Computing Survey (LAICS) and Trends
by: Reuther, Albert, et al.
Published: (2025)
by: Reuther, Albert, et al.
Published: (2025)
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
by: Liu, Qunyou, et al.
Published: (2025)
by: Liu, Qunyou, et al.
Published: (2025)
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
by: Odema, Mohanad, et al.
Published: (2024)
by: Odema, Mohanad, et al.
Published: (2024)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
by: Qiao, Ye, et al.
Published: (2026)
by: Qiao, Ye, et al.
Published: (2026)
Similar Items
-
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024) -
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024) -
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
by: Palaniappan, Kathiravan
Published: (2026) -
Fast NF4 Dequantization Kernels for Large Language Model Inference
by: Qi, Xiangbo, et al.
Published: (2026) -
Design Space Exploration of Approximate Computing Techniques with a Reinforcement Learning Approach
by: Saeedi, Sepide, et al.
Published: (2023)