Banked Memories for Soft SIMT Processors
Fuente:
arXiv
Saved in:
| Main Authors: | Langhammer, Martin, Constantinides, George A. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A 950 MHz SIMT Soft Processor
by: Langhammer, Martin, et al.
Published: (2025)
by: Langhammer, Martin, et al.
Published: (2025)
A Statically and Dynamically Scalable Soft GPGPU
by: Langhammer, Martin, et al.
Published: (2024)
by: Langhammer, Martin, et al.
Published: (2024)
Soft GPGPU versus IP cores: Quantifying and Reducing the Performance Gap
by: Langhammer, Martin, et al.
Published: (2024)
by: Langhammer, Martin, et al.
Published: (2024)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
by: Wang, Jiayi, et al.
Published: (2026)
by: Wang, Jiayi, et al.
Published: (2026)
ROVER: RTL Optimization via Verified E-Graph Rewriting
by: Coward, Samuel, et al.
Published: (2024)
by: Coward, Samuel, et al.
Published: (2024)
NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable Functions
by: Andronic, Marta, et al.
Published: (2024)
by: Andronic, Marta, et al.
Published: (2024)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
by: Andronic, Marta, et al.
Published: (2023)
by: Andronic, Marta, et al.
Published: (2023)
Combining Power and Arithmetic Optimization via Datapath Rewriting
by: Coward, Samuel, et al.
Published: (2024)
by: Coward, Samuel, et al.
Published: (2024)
StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign
by: Wu, Michael, et al.
Published: (2025)
by: Wu, Michael, et al.
Published: (2025)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
by: Andronic, Marta, et al.
Published: (2025)
by: Andronic, Marta, et al.
Published: (2025)
FPGA Resource-aware Structured Pruning for Real-Time Neural Networks
by: Ramhorst, Benjamin, et al.
Published: (2023)
by: Ramhorst, Benjamin, et al.
Published: (2023)
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
by: Biggs, Benjamin, et al.
Published: (2023)
by: Biggs, Benjamin, et al.
Published: (2023)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
by: Cheng, Jianyi, et al.
Published: (2023)
by: Cheng, Jianyi, et al.
Published: (2023)
CMAX-CAMEL: A Coarse-to-Fine Adaptive, Memory-Efficient, and Low-Power Edge Processor for Contrast Maximization
by: Min, Kyeongpil, et al.
Published: (2026)
by: Min, Kyeongpil, et al.
Published: (2026)
Large Processor Chip Model
by: Chang, Kaiyan, et al.
Published: (2025)
by: Chang, Kaiyan, et al.
Published: (2025)
Exploring FPGA designs for MX and beyond
by: Samson, Ebby, et al.
Published: (2024)
by: Samson, Ebby, et al.
Published: (2024)
ReducedLUT: Table Decomposition with "Don't Care" Conditions
by: Cassidy, Oliver, et al.
Published: (2024)
by: Cassidy, Oliver, et al.
Published: (2024)
FERIVer: An FPGA-assisted Emulated Framework for RTL Verification of RISC-V Processors
by: Qin, Kun, et al.
Published: (2025)
by: Qin, Kun, et al.
Published: (2025)
Neuromorphic Processor Employing FPGA Technology with Universal Interconnections
by: Harlikar, Pracheta, et al.
Published: (2025)
by: Harlikar, Pracheta, et al.
Published: (2025)
SAMIPS: A Synthesised Asynchronous Processor
by: Zhang, Qianyi, et al.
Published: (2024)
by: Zhang, Qianyi, et al.
Published: (2024)
BARD: Reducing Write Latency of DDR5 Memory by Exploiting Bank-Parallelism
by: Vittal, Suhas, et al.
Published: (2025)
by: Vittal, Suhas, et al.
Published: (2025)
Per-Bank Memory Bandwidth Regulation for Predictable and Performant Real-Time System
by: Sullivan, Connor Rudy, et al.
Published: (2026)
by: Sullivan, Connor Rudy, et al.
Published: (2026)
Hypervisor Extension for a RISC-V Processor
by: Gauchola, Jaume, et al.
Published: (2024)
by: Gauchola, Jaume, et al.
Published: (2024)
Implementation of Compute Intensive Algorithms on Software Configurable Processor
by: Ganesha, et al.
Published: (2025)
by: Ganesha, et al.
Published: (2025)
Duet: Creating Harmony between Processors and Embedded FPGAs
by: Li, Ang, et al.
Published: (2023)
by: Li, Ang, et al.
Published: (2023)
Web-Based Simulator of Superscalar RISC-V Processors
by: Jaros, Jiri, et al.
Published: (2024)
by: Jaros, Jiri, et al.
Published: (2024)
CVA6-VMRT: A Modular Approach Towards Time-Predictable Virtual Memory in a 64-bit Application Class RISC-V Processor
by: Reinwardt, Christopher, et al.
Published: (2025)
by: Reinwardt, Christopher, et al.
Published: (2025)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
by: Gimenes, Pedro, et al.
Published: (2025)
by: Gimenes, Pedro, et al.
Published: (2025)
Image processing Application Development on Software Configurable Processor Array
by: Prabhu, Ganesh, et al.
Published: (2025)
by: Prabhu, Ganesh, et al.
Published: (2025)
Functional ISS-Driven Verification of Superscalar RISC-V Processors
by: Galimberti, Andrea, et al.
Published: (2024)
by: Galimberti, Andrea, et al.
Published: (2024)
Loop Control Management in Tightly Coupled Processor Arrays (TCPAs)
by: Walter, Dominik, et al.
Published: (2026)
by: Walter, Dominik, et al.
Published: (2026)
Floating Point HUB Adder for RISC-V Sargantana Processor
by: Bandera, Gerardo, et al.
Published: (2024)
by: Bandera, Gerardo, et al.
Published: (2024)
Automatic Microarchitecture-Aware Custom Instruction Design for RISC-V Processors
by: Rezunov, Evgenii, et al.
Published: (2025)
by: Rezunov, Evgenii, et al.
Published: (2025)
TurboFuzz: FPGA Accelerated Hardware Fuzzing for Processor Agile Verification
by: Zhong, Yang, et al.
Published: (2025)
by: Zhong, Yang, et al.
Published: (2025)
Mapping and Execution of Nested Loops on Processor Arrays: CGRAs vs. TCPAs
by: Walter, Dominik, et al.
Published: (2025)
by: Walter, Dominik, et al.
Published: (2025)
Optimizing Structured-Sparse Matrix Multiplication in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
AGON: Automated Design Framework for Customizing Processors from ISA Documents
by: Li, Chongxiao, et al.
Published: (2024)
by: Li, Chongxiao, et al.
Published: (2024)
On the Impact of ISA Extension on Energy Consumption of I-Cache in Extensible Processors
by: Behboudi, Noushin, et al.
Published: (2024)
by: Behboudi, Noushin, et al.
Published: (2024)
Similar Items
-
A 950 MHz SIMT Soft Processor
by: Langhammer, Martin, et al.
Published: (2025) -
A Statically and Dynamically Scalable Soft GPGPU
by: Langhammer, Martin, et al.
Published: (2024) -
Soft GPGPU versus IP cores: Quantifying and Reducing the Performance Gap
by: Langhammer, Martin, et al.
Published: (2024) -
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
by: Yin, Ruokai, et al.
Published: (2025) -
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
by: Wang, Jiayi, et al.
Published: (2026)