Guardado en:
| Autores principales: | Wu, Ruilong, Wang, Yisu, Kutscher, Dirk |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2408.15568 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
por: Zou, Jiaxiang, et al.
Publicado: (2026)
por: Zou, Jiaxiang, et al.
Publicado: (2026)
A System Development Kit for Big Data Applications on FPGA-based Clusters: The EVEREST Approach
por: Pilato, Christian, et al.
Publicado: (2024)
por: Pilato, Christian, et al.
Publicado: (2024)
When Small Variations Become Big Failures: Reliability Challenges in Compute-in-Memory Neural Accelerators
por: Qin, Yifan, et al.
Publicado: (2026)
por: Qin, Yifan, et al.
Publicado: (2026)
An Affordable Experimental Technique for SRAM Write Margin Characterization for Nanometer CMOS Technologies
por: Alorda, Bartomeu, et al.
Publicado: (2024)
por: Alorda, Bartomeu, et al.
Publicado: (2024)
Leveraging Recurrent Patterns in Graph Accelerators
por: Rahimi, Masoud, et al.
Publicado: (2025)
por: Rahimi, Masoud, et al.
Publicado: (2025)
Towards An Approach to Identify Divergences in Hardware Designs for HPC Workloads
por: Popovici, Doru Thom, et al.
Publicado: (2025)
por: Popovici, Doru Thom, et al.
Publicado: (2025)
Toward Open-Source Chiplets for HPC and AI: Occamy and Beyond
por: Scheffler, Paul, et al.
Publicado: (2025)
por: Scheffler, Paul, et al.
Publicado: (2025)
Big-PERCIVAL: Exploring the Native Use of 64-Bit Posit Arithmetic in Scientific Computing
por: Mallasén, David, et al.
Publicado: (2023)
por: Mallasén, David, et al.
Publicado: (2023)
Educating for Hardware Specialization in the Chiplet Era: A Path for the HPC Community
por: Yoshii, Kazutomo, et al.
Publicado: (2024)
por: Yoshii, Kazutomo, et al.
Publicado: (2024)
Characterization of Real Communication Patterns and Congestion Dynamics in HPC Interconnection Networks
por: de La Rosa, Miguel Sánchez, et al.
Publicado: (2026)
por: de La Rosa, Miguel Sánchez, et al.
Publicado: (2026)
Reconfigurable Computing Challenge: Real-Time Graph Neural Networks for Online Event Selection in Big Science
por: Neu, Marc, et al.
Publicado: (2026)
por: Neu, Marc, et al.
Publicado: (2026)
Calibrating DRAMPower Model for HPC: A Runtime Perspective from Real-Time Measurements
por: Shi, Xinyu, et al.
Publicado: (2024)
por: Shi, Xinyu, et al.
Publicado: (2024)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
por: Wang, Huizheng, et al.
Publicado: (2025)
por: Wang, Huizheng, et al.
Publicado: (2025)
Late Breaking Results: Leveraging Approximate Computing for Carbon-Aware DNN Accelerators
por: Panteleaki, Aikaterini Maria, et al.
Publicado: (2025)
por: Panteleaki, Aikaterini Maria, et al.
Publicado: (2025)
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
por: Kim, Dong Eun, et al.
Publicado: (2025)
por: Kim, Dong Eun, et al.
Publicado: (2025)
Apple vs. Oranges: Evaluating the Apple Silicon M-Series SoCs for HPC Performance and Efficiency
por: Hübner, Paul, et al.
Publicado: (2025)
por: Hübner, Paul, et al.
Publicado: (2025)
FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
por: Wang, Zeke, et al.
Publicado: (2025)
por: Wang, Zeke, et al.
Publicado: (2025)
Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
por: Zhu, Zhantong, et al.
Publicado: (2025)
por: Zhu, Zhantong, et al.
Publicado: (2025)
Efficient and Accurate Graph Classification with Hyperdimensional Computing on FPGA
por: Arockiaraj, Jebacyril, et al.
Publicado: (2025)
por: Arockiaraj, Jebacyril, et al.
Publicado: (2025)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
por: Mazzola, Sergio, et al.
Publicado: (2024)
por: Mazzola, Sergio, et al.
Publicado: (2024)
Spatz: Clustering Compact RISC-V-Based Vector Units to Maximize Computing Efficiency
por: Perotti, Matteo, et al.
Publicado: (2023)
por: Perotti, Matteo, et al.
Publicado: (2023)
PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration
por: Chong, Yue Jiet, et al.
Publicado: (2025)
por: Chong, Yue Jiet, et al.
Publicado: (2025)
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
por: Liu, Lian, et al.
Publicado: (2025)
por: Liu, Lian, et al.
Publicado: (2025)
SpikeStream: Accelerating Spiking Neural Network Inference on RISC-V Clusters with Sparse Computation Extensions
por: Manoni, Simone, et al.
Publicado: (2025)
por: Manoni, Simone, et al.
Publicado: (2025)
ACS: Concurrent Kernel Execution on Irregular, Input-Dependent Computational Graphs
por: Durvasula, Sankeerth, et al.
Publicado: (2024)
por: Durvasula, Sankeerth, et al.
Publicado: (2024)
FuseMax: Leveraging Extended Einsums to Optimize Attention Accelerator Design
por: Nayak, Nandeeka, et al.
Publicado: (2024)
por: Nayak, Nandeeka, et al.
Publicado: (2024)
ADS-IMC: Accelerating Data Sorting with In-Memory Computation
por: Dhakad, Narendra Singh, et al.
Publicado: (2026)
por: Dhakad, Narendra Singh, et al.
Publicado: (2026)
The Tiny Median Filter: A Small Size, Flexible Arbitrary Percentile Finder Scheme Suitable for FPGA Implementation
por: Wu, Jinyuan
Publicado: (2024)
por: Wu, Jinyuan
Publicado: (2024)
ControlPULP: A RISC-V On-Chip Parallel Power Controller for Many-Core HPC Processors with FPGA-Based Hardware-In-The-Loop Power and Thermal Emulation
por: Ottaviano, Alessandro, et al.
Publicado: (2023)
por: Ottaviano, Alessandro, et al.
Publicado: (2023)
Automated Physical Design Watermarking Leveraging Graph Neural Networks
por: Zhang, Ruisi, et al.
Publicado: (2024)
por: Zhang, Ruisi, et al.
Publicado: (2024)
EN-T: Optimizing Tensor Computing Engines Performance via Encoder-Based Methodology
por: Wu, Qizhe, et al.
Publicado: (2024)
por: Wu, Qizhe, et al.
Publicado: (2024)
NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data Processing
por: Wang, Yitu, et al.
Publicado: (2023)
por: Wang, Yitu, et al.
Publicado: (2023)
Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAM
por: Mamdouh, Ahmed, et al.
Publicado: (2024)
por: Mamdouh, Ahmed, et al.
Publicado: (2024)
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
por: Wu, Haibin, et al.
Publicado: (2024)
por: Wu, Haibin, et al.
Publicado: (2024)
SiHGNN: Leveraging Properties of Semantic Graphs for Efficient HGNN Acceleration
por: Xue, Runzhen, et al.
Publicado: (2024)
por: Xue, Runzhen, et al.
Publicado: (2024)
CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
por: Dong, Jiale, et al.
Publicado: (2025)
por: Dong, Jiale, et al.
Publicado: (2025)
RecFlash: Fast Recommendation System on In-Storage Computing with Frequency-Based Data Mapping
por: Baik, Jangho, et al.
Publicado: (2026)
por: Baik, Jangho, et al.
Publicado: (2026)
A Computing-in-Memory-based One-Class Hyperdimensional Computing Model for Outlier Detection
por: Wang, Ruixuan, et al.
Publicado: (2023)
por: Wang, Ruixuan, et al.
Publicado: (2023)
The Data Conversion Bottleneck in Analog Computing Accelerators
por: Meech, James T., et al.
Publicado: (2023)
por: Meech, James T., et al.
Publicado: (2023)
Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation
por: Wang, Yanjing, et al.
Publicado: (2025)
por: Wang, Yanjing, et al.
Publicado: (2025)
Ejemplares similares
-
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
por: Zou, Jiaxiang, et al.
Publicado: (2026) -
A System Development Kit for Big Data Applications on FPGA-based Clusters: The EVEREST Approach
por: Pilato, Christian, et al.
Publicado: (2024) -
When Small Variations Become Big Failures: Reliability Challenges in Compute-in-Memory Neural Accelerators
por: Qin, Yifan, et al.
Publicado: (2026) -
An Affordable Experimental Technique for SRAM Write Margin Characterization for Nanometer CMOS Technologies
por: Alorda, Bartomeu, et al.
Publicado: (2024) -
Leveraging Recurrent Patterns in Graph Accelerators
por: Rahimi, Masoud, et al.
Publicado: (2025)