Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Hansung, Yan, Ruohan Richard, You, Joshua, Yang, Tieliang Vamber, Shao, Yakun Sophia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
von: Perotti, Matteo, et al.
Veröffentlicht: (2023)
von: Perotti, Matteo, et al.
Veröffentlicht: (2023)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024)
DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators
von: Hong, Charles, et al.
Veröffentlicht: (2025)
von: Hong, Charles, et al.
Veröffentlicht: (2025)
A Systematic Characterization of LLM Inference on GPUs
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
Control Flow Management in Modern GPUs
von: Shoushtary, Mojtaba Abaie, et al.
Veröffentlicht: (2024)
von: Shoushtary, Mojtaba Abaie, et al.
Veröffentlicht: (2024)
Privacy-Preserving Performance Profiling of In-The-Wild GPUs
von: McDougall, Ian, et al.
Veröffentlicht: (2025)
von: McDougall, Ian, et al.
Veröffentlicht: (2025)
An Energy-Efficient Approximate Posit Multiply-Divide Unit
von: Thotli, Rishi, et al.
Veröffentlicht: (2026)
von: Thotli, Rishi, et al.
Veröffentlicht: (2026)
Sparse-on-Dense: Area and Energy-Efficient Computing of Sparse Neural Networks on Dense Matrix Multiplication Accelerators
von: Yoon, Hyunsung, et al.
Veröffentlicht: (2026)
von: Yoon, Hyunsung, et al.
Veröffentlicht: (2026)
How to Increase Energy Efficiency with a Single Linux Command
von: Jelvani, Alborz, et al.
Veröffentlicht: (2025)
von: Jelvani, Alborz, et al.
Veröffentlicht: (2025)
D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2026)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2026)
Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs
von: Chowdhary, Sangeeta, et al.
Veröffentlicht: (2026)
von: Chowdhary, Sangeeta, et al.
Veröffentlicht: (2026)
Optimizing Scalable Multi-Cluster Architectures for Next-Generation Wireless Sensing and Communication
von: Riedel, Samuel, et al.
Veröffentlicht: (2025)
von: Riedel, Samuel, et al.
Veröffentlicht: (2025)
Optimizing Energy Efficiency in Subthreshold RISC-V Cores
von: Djupdal, Asbjørn, et al.
Veröffentlicht: (2025)
von: Djupdal, Asbjørn, et al.
Veröffentlicht: (2025)
CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs
von: Hu, Jiajun, et al.
Veröffentlicht: (2025)
von: Hu, Jiajun, et al.
Veröffentlicht: (2025)
A Scalable Resource Management Layer for FPGA SoCs in 6G Radio Units
von: Bartzoudis, Nikolaos, et al.
Veröffentlicht: (2025)
von: Bartzoudis, Nikolaos, et al.
Veröffentlicht: (2025)
Integrating Prefetcher Selection with Dynamic Request Allocation Improves Prefetching Efficiency
von: Li, Mengming, et al.
Veröffentlicht: (2025)
von: Li, Mengming, et al.
Veröffentlicht: (2025)
GPIR: Enabling Practical Private Information Retrieval with GPUs
von: Ji, Hyesung, et al.
Veröffentlicht: (2026)
von: Ji, Hyesung, et al.
Veröffentlicht: (2026)
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency
von: Kyung, Kwanhee, et al.
Veröffentlicht: (2025)
von: Kyung, Kwanhee, et al.
Veröffentlicht: (2025)
16 Years of SPEC Power: An Analysis of x86 Energy Efficiency Trends
von: Tröpgen, Hannes, et al.
Veröffentlicht: (2024)
von: Tröpgen, Hannes, et al.
Veröffentlicht: (2024)
Increasing the Energy-Efficiency of Wearables Using Low-Precision Posit Arithmetic with PHEE
von: Mallasén, David, et al.
Veröffentlicht: (2025)
von: Mallasén, David, et al.
Veröffentlicht: (2025)
Study on the Particle Sorting Performance for Reactor Monte Carlo Neutron Transport on Apple Unified Memory GPUs
von: Liu, Changyuan
Veröffentlicht: (2024)
von: Liu, Changyuan
Veröffentlicht: (2024)
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators
von: Hong, Charles, et al.
Veröffentlicht: (2025)
von: Hong, Charles, et al.
Veröffentlicht: (2025)
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
von: Perotti, Matteo, et al.
Veröffentlicht: (2024)
von: Perotti, Matteo, et al.
Veröffentlicht: (2024)
Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale Integration
von: Feng, Yinxiao, et al.
Veröffentlicht: (2024)
von: Feng, Yinxiao, et al.
Veröffentlicht: (2024)
LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference
von: Moon, Seungjae, et al.
Veröffentlicht: (2024)
von: Moon, Seungjae, et al.
Veröffentlicht: (2024)
Evaluation of Run-Time Energy Efficiency using Controlled Approximation in a RISC-V Core
von: Delavari, Arvin, et al.
Veröffentlicht: (2024)
von: Delavari, Arvin, et al.
Veröffentlicht: (2024)
Modeling PFAS in Semiconductor Manufacturing to Quantify Trade-offs in Energy Efficiency and Environmental Impact of Computing Systems
von: Elgamal, Mariam, et al.
Veröffentlicht: (2025)
von: Elgamal, Mariam, et al.
Veröffentlicht: (2025)
ISAAC: Intelligent, Scalable, Agile, and Accelerated CPU Verification via LLM-aided FPGA Parallelism
von: Sun, Jialin, et al.
Veröffentlicht: (2025)
von: Sun, Jialin, et al.
Veröffentlicht: (2025)
Instruction Scheduling in the Saturn Vector Unit
von: Zhao, Jerry, et al.
Veröffentlicht: (2024)
von: Zhao, Jerry, et al.
Veröffentlicht: (2024)
Ramping Up Open-Source RISC-V Cores: Assessing the Energy Efficiency of Superscalar, Out-of-Order Execution
von: Fu, Zexin, et al.
Veröffentlicht: (2025)
von: Fu, Zexin, et al.
Veröffentlicht: (2025)
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
A Quality-Aware Voltage Overscaling Framework to Improve the Energy Efficiency and Lifetime of TPUs based on Statistical Error Modeling
von: Senobari, Alireza, et al.
Veröffentlicht: (2024)
von: Senobari, Alireza, et al.
Veröffentlicht: (2024)
LFOC+: A Fair OS-level Cache-Clustering Policy for Commodity Multicore Systems
von: Saez, Juan Carlos, et al.
Veröffentlicht: (2024)
von: Saez, Juan Carlos, et al.
Veröffentlicht: (2024)
Fast, Scalable, Energy-Efficient Non-element-wise Matrix Multiplication on FPGA
von: Zhu, Xuqi, et al.
Veröffentlicht: (2024)
von: Zhu, Xuqi, et al.
Veröffentlicht: (2024)
RTGPU: Real-Time Computing with Graphics Processing Units
von: Gheibi-Fetrat, Atiyeh, et al.
Veröffentlicht: (2025)
von: Gheibi-Fetrat, Atiyeh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
von: Perotti, Matteo, et al.
Veröffentlicht: (2023) -
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025) -
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024) -
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
von: Hong, Jeongmin, et al.
Veröffentlicht: (2024) -
DOSA: Differentiable Model-Based One-Loop Search for DNN Accelerators
von: Hong, Charles, et al.
Veröffentlicht: (2025)