Strassen Multisystolic Array Hardware Architectures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pogue, Trevor E., Nicolici, Nicola |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Karatsuba Matrix Multiplication and its Efficient Custom Hardware Implementations
von: Pogue, Trevor E., et al.
Veröffentlicht: (2025)
von: Pogue, Trevor E., et al.
Veröffentlicht: (2025)
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
von: Vishwanathan, Manoj, et al.
Veröffentlicht: (2026)
von: Vishwanathan, Manoj, et al.
Veröffentlicht: (2026)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
von: Müller, Mika Markus, et al.
Veröffentlicht: (2025)
von: Müller, Mika Markus, et al.
Veröffentlicht: (2025)
Assessing Tenstorrent's RISC-V MatMul Acceleration Capabilities
von: Cavagna, Hiari Pizzini, et al.
Veröffentlicht: (2025)
von: Cavagna, Hiari Pizzini, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
von: Fang, Yunhua, et al.
Veröffentlicht: (2025)
von: Fang, Yunhua, et al.
Veröffentlicht: (2025)
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
von: Chen, Yao, et al.
Veröffentlicht: (2022)
von: Chen, Yao, et al.
Veröffentlicht: (2022)
Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference
von: Javat, Abdurrahman, et al.
Veröffentlicht: (2026)
von: Javat, Abdurrahman, et al.
Veröffentlicht: (2026)
KWT-Tiny: RISC-V Accelerated, Embedded Keyword Spotting Transformer
von: Al-Qawlaq, Aness, et al.
Veröffentlicht: (2024)
von: Al-Qawlaq, Aness, et al.
Veröffentlicht: (2024)
LLM-Driven Design Space Exploration of FPGA-based Accelerators
von: Sharma, Vinamra, et al.
Veröffentlicht: (2026)
von: Sharma, Vinamra, et al.
Veröffentlicht: (2026)
DRAGON (Differentiable Graph Execution) : A suite of Hardware Simulation and Optimization tools for Modern AI/Non-AI Workloads
von: Sethi, Khushal
Veröffentlicht: (2022)
von: Sethi, Khushal
Veröffentlicht: (2022)
OISMA: On-the-fly In-memory Stochastic Multiplication Architecture for Matrix-Multiplication Workloads
von: Agwa, Shady, et al.
Veröffentlicht: (2025)
von: Agwa, Shady, et al.
Veröffentlicht: (2025)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
von: Yang, Hanchen, et al.
Veröffentlicht: (2025)
von: Yang, Hanchen, et al.
Veröffentlicht: (2025)
DISCA: A Digital In-memory Stochastic Computing Architecture Using A Compressed Bent-Pyramid Format
von: Agwa, Shady, et al.
Veröffentlicht: (2025)
von: Agwa, Shady, et al.
Veröffentlicht: (2025)
GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS
von: Kazemkhani, Saman, et al.
Veröffentlicht: (2024)
von: Kazemkhani, Saman, et al.
Veröffentlicht: (2024)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
von: Renney, Harri, et al.
Veröffentlicht: (2026)
von: Renney, Harri, et al.
Veröffentlicht: (2026)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
von: Mazzola, Sergio, et al.
Veröffentlicht: (2025)
OmniSim: Simulating Hardware with C Speed and RTL Accuracy for High-Level Synthesis Designs
von: Sarkar, Rishov, et al.
Veröffentlicht: (2025)
von: Sarkar, Rishov, et al.
Veröffentlicht: (2025)
Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
von: Liu, Yuhao, et al.
Veröffentlicht: (2026)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
von: Du, Dayou, et al.
Veröffentlicht: (2025)
von: Du, Dayou, et al.
Veröffentlicht: (2025)
Optimization of Armv9 architecture general large language model inference performance based on Llama.cpp
von: Chen, Longhao, et al.
Veröffentlicht: (2024)
von: Chen, Longhao, et al.
Veröffentlicht: (2024)
The Unseen AI Disruptions for Power Grids: LLM-Induced Transients
von: Li, Yuzhuo, et al.
Veröffentlicht: (2024)
von: Li, Yuzhuo, et al.
Veröffentlicht: (2024)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
von: Lübeck, Konstantin, et al.
Veröffentlicht: (2024)
von: Lübeck, Konstantin, et al.
Veröffentlicht: (2024)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
von: Bi, Zhen, et al.
Veröffentlicht: (2026)
von: Bi, Zhen, et al.
Veröffentlicht: (2026)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
von: Chhugani, Jatin, et al.
Veröffentlicht: (2026)
von: Chhugani, Jatin, et al.
Veröffentlicht: (2026)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
von: Jung, Alexander Louis-Ferdinand, et al.
Veröffentlicht: (2024)
von: Jung, Alexander Louis-Ferdinand, et al.
Veröffentlicht: (2024)
Characterizing and Understanding HGNN Training on GPUs
von: Han, Dengke, et al.
Veröffentlicht: (2024)
von: Han, Dengke, et al.
Veröffentlicht: (2024)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
von: Ali, Wajid, et al.
Veröffentlicht: (2025)
von: Ali, Wajid, et al.
Veröffentlicht: (2025)
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation
von: Ke, Chih-Hua
Veröffentlicht: (2026)
von: Ke, Chih-Hua
Veröffentlicht: (2026)
In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
von: Hoefler, Torsten, et al.
Veröffentlicht: (2026)
von: Hoefler, Torsten, et al.
Veröffentlicht: (2026)
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
von: Du, Dayou, et al.
Veröffentlicht: (2024)
von: Du, Dayou, et al.
Veröffentlicht: (2024)
Toward A Formalized Approach for Spike Sorting Algorithms and Hardware Evaluation
von: Zhang, Tim, et al.
Veröffentlicht: (2022)
von: Zhang, Tim, et al.
Veröffentlicht: (2022)
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
von: Panigrahy, Deepak, et al.
Veröffentlicht: (2026)
von: Panigrahy, Deepak, et al.
Veröffentlicht: (2026)
Quantum Hardware Roofline: Evaluating the Impact of Gate Expressivity on Quantum Processor Design
von: Kalloor, Justin, et al.
Veröffentlicht: (2024)
von: Kalloor, Justin, et al.
Veröffentlicht: (2024)
Using the Abstract Computer Architecture Description Language to Model AI Hardware Accelerators
von: Müller, Mika Markus, et al.
Veröffentlicht: (2024)
von: Müller, Mika Markus, et al.
Veröffentlicht: (2024)
Towards Efficient Neuro-Symbolic AI: From Workload Characterization to Hardware Architecture
von: Wan, Zishen, et al.
Veröffentlicht: (2024)
von: Wan, Zishen, et al.
Veröffentlicht: (2024)
CounterPoint: Using Hardware Event Counters to Refute and Refine Microarchitectural Assumptions (Extended Version)
von: Lindsay, Nick, et al.
Veröffentlicht: (2026)
von: Lindsay, Nick, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Karatsuba Matrix Multiplication and its Efficient Custom Hardware Implementations
von: Pogue, Trevor E., et al.
Veröffentlicht: (2025) -
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
von: Vishwanathan, Manoj, et al.
Veröffentlicht: (2026) -
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025) -
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
von: Müller, Mika Markus, et al.
Veröffentlicht: (2025) -
Assessing Tenstorrent's RISC-V MatMul Acceleration Capabilities
von: Cavagna, Hiari Pizzini, et al.
Veröffentlicht: (2025)