A Statically and Dynamically Scalable Soft GPGPU
Fuente:
arXiv
Salvato in:
| Autori principali: | Langhammer, Martin, Constantinides, George A. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Soft GPGPU versus IP cores: Quantifying and Reducing the Performance Gap
di: Langhammer, Martin, et al.
Pubblicazione: (2024)
di: Langhammer, Martin, et al.
Pubblicazione: (2024)
Banked Memories for Soft SIMT Processors
di: Langhammer, Martin, et al.
Pubblicazione: (2025)
di: Langhammer, Martin, et al.
Pubblicazione: (2025)
A 950 MHz SIMT Soft Processor
di: Langhammer, Martin, et al.
Pubblicazione: (2025)
di: Langhammer, Martin, et al.
Pubblicazione: (2025)
ROVER: RTL Optimization via Verified E-Graph Rewriting
di: Coward, Samuel, et al.
Pubblicazione: (2024)
di: Coward, Samuel, et al.
Pubblicazione: (2024)
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
di: Sarda, Giuseppe M., et al.
Pubblicazione: (2024)
di: Sarda, Giuseppe M., et al.
Pubblicazione: (2024)
Sim-FA: A GPGPU Simulator Framework for Fine-Grained FlashAttention Pipeline Analysis
di: Zhou, Zhongchun, et al.
Pubblicazione: (2026)
di: Zhou, Zhongchun, et al.
Pubblicazione: (2026)
NeuraLUT: Hiding Neural Network Density in Boolean Synthesizable Functions
di: Andronic, Marta, et al.
Pubblicazione: (2024)
di: Andronic, Marta, et al.
Pubblicazione: (2024)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
di: Andronic, Marta, et al.
Pubblicazione: (2023)
di: Andronic, Marta, et al.
Pubblicazione: (2023)
Combining Power and Arithmetic Optimization via Datapath Rewriting
di: Coward, Samuel, et al.
Pubblicazione: (2024)
di: Coward, Samuel, et al.
Pubblicazione: (2024)
StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign
di: Wu, Michael, et al.
Pubblicazione: (2025)
di: Wu, Michael, et al.
Pubblicazione: (2025)
FPGA Resource-aware Structured Pruning for Real-Time Neural Networks
di: Ramhorst, Benjamin, et al.
Pubblicazione: (2023)
di: Ramhorst, Benjamin, et al.
Pubblicazione: (2023)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
di: Andronic, Marta, et al.
Pubblicazione: (2025)
di: Andronic, Marta, et al.
Pubblicazione: (2025)
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
di: Biggs, Benjamin, et al.
Pubblicazione: (2023)
di: Biggs, Benjamin, et al.
Pubblicazione: (2023)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
di: Cheng, Jianyi, et al.
Pubblicazione: (2023)
di: Cheng, Jianyi, et al.
Pubblicazione: (2023)
Exploring FPGA designs for MX and beyond
di: Samson, Ebby, et al.
Pubblicazione: (2024)
di: Samson, Ebby, et al.
Pubblicazione: (2024)
ReducedLUT: Table Decomposition with "Don't Care" Conditions
di: Cassidy, Oliver, et al.
Pubblicazione: (2024)
di: Cassidy, Oliver, et al.
Pubblicazione: (2024)
Table-Lookup MAC: Scalable Processing of Quantised Neural Networks in FPGA Soft Logic
di: Gerlinghoff, Daniel, et al.
Pubblicazione: (2024)
di: Gerlinghoff, Daniel, et al.
Pubblicazione: (2024)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
Beyond Static Policies: Exploring Dynamic Policy Selection for Single-Thread Performance Optimization
di: Zhang, Yanxin, et al.
Pubblicazione: (2026)
di: Zhang, Yanxin, et al.
Pubblicazione: (2026)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
Ten-Four: An Open-Source Fused Dot Product Unit for Mixed-Precision GPGPU Tensor Cores
di: Rout, Nikhil, et al.
Pubblicazione: (2025)
di: Rout, Nikhil, et al.
Pubblicazione: (2025)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
di: Xu, Weikai, et al.
Pubblicazione: (2025)
di: Xu, Weikai, et al.
Pubblicazione: (2025)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
di: Lee, Dongjae, et al.
Pubblicazione: (2025)
di: Lee, Dongjae, et al.
Pubblicazione: (2025)
Static Hardware Partitioning on RISC-V -- Shortcomings, Limitations, and Prospects
di: Ramsauer, Ralf, et al.
Pubblicazione: (2022)
di: Ramsauer, Ralf, et al.
Pubblicazione: (2022)
Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous Parallelism
di: Emami, Mahyar, et al.
Pubblicazione: (2023)
di: Emami, Mahyar, et al.
Pubblicazione: (2023)
3DGauCIM: Accelerating Static/Dynamic 3D Gaussian Splatting via Digital CIM for High Frame Rate Real-Time Edge Rendering
di: Huang, Wei-Hsing, et al.
Pubblicazione: (2025)
di: Huang, Wei-Hsing, et al.
Pubblicazione: (2025)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
di: Chen, Yuzong, et al.
Pubblicazione: (2024)
di: Chen, Yuzong, et al.
Pubblicazione: (2024)
Exploiting Control-flow Enforcement Technology for Sound and Precise Static Binary Disassembly
di: Zhao, Brian, et al.
Pubblicazione: (2025)
di: Zhao, Brian, et al.
Pubblicazione: (2025)
Global and Local Attention-based Inception U-Net for Static IR Drop Prediction
di: Chen, Yilu, et al.
Pubblicazione: (2024)
di: Chen, Yilu, et al.
Pubblicazione: (2024)
SCREME: A Scalable Framework for Resilient Memory Design
di: Li, Fan, et al.
Pubblicazione: (2025)
di: Li, Fan, et al.
Pubblicazione: (2025)
Runtime Energy Monitoring for RISC-V Soft-Cores
di: Scionti, Alberto, et al.
Pubblicazione: (2025)
di: Scionti, Alberto, et al.
Pubblicazione: (2025)
AssertMiner: Module-Level Spec Generation and Assertion Mining using Static Analysis Guided LLMs
di: Lyu, Hongqin, et al.
Pubblicazione: (2025)
di: Lyu, Hongqin, et al.
Pubblicazione: (2025)
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays
di: Wang, Jiayi, et al.
Pubblicazione: (2026)
di: Wang, Jiayi, et al.
Pubblicazione: (2026)
Soft Error Probability Estimation of Nano-scale Combinational Circuits
di: Jockar, Ali, et al.
Pubblicazione: (2025)
di: Jockar, Ali, et al.
Pubblicazione: (2025)
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
di: Zheng, Jianing, et al.
Pubblicazione: (2025)
di: Zheng, Jianing, et al.
Pubblicazione: (2025)
QED: Scalable Verification of Hardware Memory Consistency
di: Ravi, Gokulan, et al.
Pubblicazione: (2024)
di: Ravi, Gokulan, et al.
Pubblicazione: (2024)
Towards Reliable Systems: A Scalable Approach to AXI4 Transaction Monitoring
di: Liang, Chaoqun, et al.
Pubblicazione: (2025)
di: Liang, Chaoqun, et al.
Pubblicazione: (2025)
A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
di: Petropoulos, Anastasios, et al.
Pubblicazione: (2025)
di: Petropoulos, Anastasios, et al.
Pubblicazione: (2025)
CRYPTONITE: Scalable Accelerator Design for Cryptographic Primitives and Algorithms
di: Maheswaran, Karthikeya Sharma, et al.
Pubblicazione: (2025)
di: Maheswaran, Karthikeya Sharma, et al.
Pubblicazione: (2025)
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
di: Moon, Seungjae, et al.
Pubblicazione: (2024)
di: Moon, Seungjae, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Soft GPGPU versus IP cores: Quantifying and Reducing the Performance Gap
di: Langhammer, Martin, et al.
Pubblicazione: (2024) -
Banked Memories for Soft SIMT Processors
di: Langhammer, Martin, et al.
Pubblicazione: (2025) -
A 950 MHz SIMT Soft Processor
di: Langhammer, Martin, et al.
Pubblicazione: (2025) -
ROVER: RTL Optimization via Verified E-Graph Rewriting
di: Coward, Samuel, et al.
Pubblicazione: (2024) -
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
di: Sarda, Giuseppe M., et al.
Pubblicazione: (2024)