A Flexible Instruction Set Architecture for Efficient GEMMs
Fuente:
arXiv
Saved in:
| Main Authors: | Santana, Alexandre de Limas, Armejach, Adrià, Martinez, Francesc, Focht, Erich, Casas, Marc |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Comparative Analysis of ARM and x86-64 Laptop-Class Processors: Architecture, Assembly-Level Performance, and Energy Efficiency
by: Özyılmaz, Mustafa Mert
Published: (2026)
by: Özyılmaz, Mustafa Mert
Published: (2026)
How long can you sleep? Idle Time System Inefficiencies and Opportunities
by: Antoniou, Georgia, et al.
Published: (2025)
by: Antoniou, Georgia, et al.
Published: (2025)
Tekum: Balanced Ternary Tapered Precision Real Arithmetic
by: Hunhold, Laslo
Published: (2025)
by: Hunhold, Laslo
Published: (2025)
basic_RV32s: An Open-Source Microarchitectural Roadmap for RISC-V RV32I
by: Kang, Hyun Woo, et al.
Published: (2025)
by: Kang, Hyun Woo, et al.
Published: (2025)
RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures
by: Kang, Hyunwoo
Published: (2026)
by: Kang, Hyunwoo
Published: (2026)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024)
by: Mo, Zhiwen, et al.
Published: (2024)
Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures
by: Siracusa, Marco, et al.
Published: (2025)
by: Siracusa, Marco, et al.
Published: (2025)
RayFlex: An Open-Source RTL Implementation of the Hardware Ray Tracer Datapath
by: Shen, Fangjia, et al.
Published: (2024)
by: Shen, Fangjia, et al.
Published: (2024)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
Data Gravity and the Energy Limits of Computation
by: Lee, Wonsuk, et al.
Published: (2026)
by: Lee, Wonsuk, et al.
Published: (2026)
HeTraX: Energy Efficient 3D Heterogeneous Manycore Architecture for Transformer Acceleration
by: Dhingra, Pratyush, et al.
Published: (2024)
by: Dhingra, Pratyush, et al.
Published: (2024)
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
by: Prabhakar, Raghu, et al.
Published: (2024)
by: Prabhakar, Raghu, et al.
Published: (2024)
Streamlining SIMD ISA Extensions with Takum Arithmetic: A Case Study on Intel AVX10.2
by: Hunhold, Laslo
Published: (2025)
by: Hunhold, Laslo
Published: (2025)
Fast NF4 Dequantization Kernels for Large Language Model Inference
by: Qi, Xiangbo, et al.
Published: (2026)
by: Qi, Xiangbo, et al.
Published: (2026)
A Dense and Efficient Instruction Set Architecture Encoding
by: Maroun, Emad Jacob
Published: (2025)
by: Maroun, Emad Jacob
Published: (2025)
Application-Driven Exascale: The JUPITER Benchmark Suite
by: Herten, Andreas, et al.
Published: (2024)
by: Herten, Andreas, et al.
Published: (2024)
Taming Wild Branches: Overcoming Hard-to-Predict Branches using the Bullseye Predictor
by: Behrendt, Emet, et al.
Published: (2025)
by: Behrendt, Emet, et al.
Published: (2025)
FREESS: A Web-Based Educational Simulator for a RISC-V-Inspired Superscalar Processor with Tomasulo-Style Dynamic Scheduling
by: Giorgi, Roberto, et al.
Published: (2026)
by: Giorgi, Roberto, et al.
Published: (2026)
Squire: A General-Purpose Accelerator to Exploit Fine-Grain Parallelism on Dependency-Bound Kernels
by: Langarita, Rubén, et al.
Published: (2025)
by: Langarita, Rubén, et al.
Published: (2025)
Mestra: Exploring Migration on Virtualized CGRAs
by: Kyriazis, Agamemnon, et al.
Published: (2026)
by: Kyriazis, Agamemnon, et al.
Published: (2026)
Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache Resources
by: Kanellopoulos, Konstantinos, et al.
Published: (2023)
by: Kanellopoulos, Konstantinos, et al.
Published: (2023)
SpChar: Characterizing the Sparse Puzzle via Decision Trees
by: Sgherzi, Francesco, et al.
Published: (2023)
by: Sgherzi, Francesco, et al.
Published: (2023)
Factor Machine: Mixed-signal Architecture for Fine-Grained Graph-Based Computing
by: Dudek, Piotr
Published: (2024)
by: Dudek, Piotr
Published: (2024)
PG-MDP: Profile-Guided Memory Dependence Prediction for Area-Constrained Cores
by: Panayi, Luke, et al.
Published: (2026)
by: Panayi, Luke, et al.
Published: (2026)
Multi-diseases detection with memristive system on chip
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Improving Memory Dependence Prediction with Static Analysis
by: Panayi, Luke, et al.
Published: (2024)
by: Panayi, Luke, et al.
Published: (2024)
Silent Data Corruption by 10x Test Escapes Threatens Reliable Computing
by: Mitra, Subhasish, et al.
Published: (2025)
by: Mitra, Subhasish, et al.
Published: (2025)
Enhancing Instruction Prefetching via Cache and TLB Management
by: Jamet, Alexandre Valentin, et al.
Published: (2026)
by: Jamet, Alexandre Valentin, et al.
Published: (2026)
Design Space Exploration of Approximate Computing Techniques with a Reinforcement Learning Approach
by: Saeedi, Sepide, et al.
Published: (2023)
by: Saeedi, Sepide, et al.
Published: (2023)
An SMT Formalization of Mixed-Precision Matrix Multiplication: Modeling Three Generations of Tensor Cores
by: Valpey, Benjamin, et al.
Published: (2025)
by: Valpey, Benjamin, et al.
Published: (2025)
FlexiBit: Fully Flexible Precision Bit-parallel Accelerator Architecture for Arbitrary Mixed Precision AI
by: Tahmasebi, Faraz, et al.
Published: (2024)
by: Tahmasebi, Faraz, et al.
Published: (2024)
AMC: Access to Miss Correlation Prefetcher for Evolving Graph Analytics
by: Singh, Abhishek, et al.
Published: (2024)
by: Singh, Abhishek, et al.
Published: (2024)
FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction
by: Kulp, Gabriel, et al.
Published: (2024)
by: Kulp, Gabriel, et al.
Published: (2024)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
by: Zhao, Chenggang, et al.
Published: (2025)
by: Zhao, Chenggang, et al.
Published: (2025)
IPU: Flexible Hardware Introspection Units
by: McDougall, Ian, et al.
Published: (2023)
by: McDougall, Ian, et al.
Published: (2023)
A Scalable Architecture for Efficient Multi-bit Fully Homomorphic Encryption
by: Ma, Jiaao, et al.
Published: (2025)
by: Ma, Jiaao, et al.
Published: (2025)
D-com: Accelerating Iterative Processing to Enable Low-rank Decomposition of Activations
by: Tahmasebi, Faraz, et al.
Published: (2025)
by: Tahmasebi, Faraz, et al.
Published: (2025)
MEDEA: A Design-Time Multi-Objective Manager for Energy-Efficient DNN Inference on Heterogeneous Ultra-Low Power Platforms
by: Taji, Hossein, et al.
Published: (2025)
by: Taji, Hossein, et al.
Published: (2025)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
by: Rout, Nikhil, et al.
Published: (2025)
by: Rout, Nikhil, et al.
Published: (2025)
Fast and Practical Strassen's Matrix Multiplication using FPGAs
by: Ahmad, Afzal, et al.
Published: (2024)
by: Ahmad, Afzal, et al.
Published: (2024)
Similar Items
-
A Comparative Analysis of ARM and x86-64 Laptop-Class Processors: Architecture, Assembly-Level Performance, and Energy Efficiency
by: Özyılmaz, Mustafa Mert
Published: (2026) -
How long can you sleep? Idle Time System Inefficiencies and Opportunities
by: Antoniou, Georgia, et al.
Published: (2025) -
Tekum: Balanced Ternary Tapered Precision Real Arithmetic
by: Hunhold, Laslo
Published: (2025) -
basic_RV32s: An Open-Source Microarchitectural Roadmap for RISC-V RV32I
by: Kang, Hyun Woo, et al.
Published: (2025) -
RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures
by: Kang, Hyunwoo
Published: (2026)