Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Qiao, Alomairy, Rabab, Wang, Dali, Gu, Zhuowei, Cao, Qinglei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
Hierarchical Recursive Precision for Accelerating Symmetric Linear Solves on MXUs
by: Carrica, Vicki, et al.
Published: (2026)
by: Carrica, Vicki, et al.
Published: (2026)
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
by: Zheng, Size, et al.
Published: (2025)
by: Zheng, Size, et al.
Published: (2025)
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
by: Carrica, Vicki, et al.
Published: (2025)
by: Carrica, Vicki, et al.
Published: (2025)
W4A16 Mixed-Precision Matrix Multiplication on Decoupled Architecture: Kernel Design and Memory Bottleneck Analysis for Ascend NPUs
by: He, Yuanhong, et al.
Published: (2026)
by: He, Yuanhong, et al.
Published: (2026)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
by: Chang, Dali, et al.
Published: (2026)
by: Chang, Dali, et al.
Published: (2026)
Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
by: Shubham, et al.
Published: (2024)
by: Shubham, et al.
Published: (2024)
Xorbits: Automating Operator Tiling for Distributed Data Science
by: Lu, Weizheng, et al.
Published: (2023)
by: Lu, Weizheng, et al.
Published: (2023)
Sparsity-Aware Roofline Models for Sparse Matrix-Matrix Multiplication
by: Qian, Matthew, et al.
Published: (2026)
by: Qian, Matthew, et al.
Published: (2026)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
by: Carl, Natalie, et al.
Published: (2025)
by: Carl, Natalie, et al.
Published: (2025)
Accelerating Sparse DNNs Based on Tiled GEMM
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
An Efficient and Adaptive Watermark Detection System with Tile-based Error Correction
by: Zhong, Xinrui, et al.
Published: (2025)
by: Zhong, Xinrui, et al.
Published: (2025)
Caching Aided Multi-Tenant Serverless Computing
by: Qiao, Chu, et al.
Published: (2024)
by: Qiao, Chu, et al.
Published: (2024)
Gaia: Hybrid Hardware Acceleration for Serverless AI in the 3D Compute Continuum
by: Reisecker, Maximilian, et al.
Published: (2025)
by: Reisecker, Maximilian, et al.
Published: (2025)
Accelerating Mixed-Precision Out-of-Core Cholesky Factorization with Static Task Scheduling
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
Stencil Matrixization
by: Zhao, Wenxuan, et al.
Published: (2023)
by: Zhao, Wenxuan, et al.
Published: (2023)
WANSpec: Leveraging Global Compute Capacity for LLM Inference
by: Martin, Noah, et al.
Published: (2026)
by: Martin, Noah, et al.
Published: (2026)
IsoSched: Preemptive Tile Cascaded Scheduling of Multi-DNN via Subgraph Isomorphism
by: Zhao, Boran, et al.
Published: (2025)
by: Zhao, Boran, et al.
Published: (2025)
Leveraging Teaching on Demand: Approaching HPC to Undergrads
by: Catalán, S., et al.
Published: (2026)
by: Catalán, S., et al.
Published: (2026)
Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
by: Uchino, Yuki, et al.
Published: (2026)
by: Uchino, Yuki, et al.
Published: (2026)
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
by: Fang, Jiahao, et al.
Published: (2024)
by: Fang, Jiahao, et al.
Published: (2024)
FalconGEMM: Surpassing Hardware Peaks with Lower-Complexity Matrix Multiplication
by: Zhu, Honglin, et al.
Published: (2026)
by: Zhu, Honglin, et al.
Published: (2026)
PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel
by: Yi, Jinjun, et al.
Published: (2025)
by: Yi, Jinjun, et al.
Published: (2025)
Architectural Foundations for Checkpointing and Restoration in Quantum HPC Systems
by: Guan, Qiang, et al.
Published: (2026)
by: Guan, Qiang, et al.
Published: (2026)
POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication
by: Rao, Yizhuo, et al.
Published: (2026)
by: Rao, Yizhuo, et al.
Published: (2026)
Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations
by: Daas, Hussam Al, et al.
Published: (2024)
by: Daas, Hussam Al, et al.
Published: (2024)
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
by: Jia, Wenqi, et al.
Published: (2025)
by: Jia, Wenqi, et al.
Published: (2025)
Data-Centric Design: Introducing An Informatics Domain Model And Core Data Ontology For Computational Systems
by: Knowles, Paul, et al.
Published: (2024)
by: Knowles, Paul, et al.
Published: (2024)
scaleTRIM: Scalable TRuncation-Based Integer Approximate Multiplier with Linearization and Compensation
by: Farahmand, Ebrahim, et al.
Published: (2023)
by: Farahmand, Ebrahim, et al.
Published: (2023)
Scalable Analysis of Urban Scaling Laws: Leveraging Cloud Computing to Analyze 21,280 Global Cities
by: Li, Zhenhui, et al.
Published: (2024)
by: Li, Zhenhui, et al.
Published: (2024)
GPU-Based Parallel Computing Methods for Medical Photoacoustic Image Reconstruction
by: Yi, Xinyao, et al.
Published: (2024)
by: Yi, Xinyao, et al.
Published: (2024)
Sparsity-Preserving Encodings for Straggler-Optimal Distributed Matrix Computations at the Edge
by: Das, Anindya Bijoy, et al.
Published: (2024)
by: Das, Anindya Bijoy, et al.
Published: (2024)
Harnessing Your DRAM and SSD for Sustainable and Accessible LLM Inference with Mixed-Precision and Multi-level Caching
by: Peng, Jie, et al.
Published: (2024)
by: Peng, Jie, et al.
Published: (2024)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
by: Qiang, Xinwei, et al.
Published: (2026)
by: Qiang, Xinwei, et al.
Published: (2026)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
by: Xue, Weicheng, et al.
Published: (2025)
by: Xue, Weicheng, et al.
Published: (2025)
A Comparative Review of Parallel Exact, Heuristic, Metaheuristic, and Hybrid Optimization Techniques for the Traveling Salesman Problem
by: Alkhalifa, Rabab, et al.
Published: (2025)
by: Alkhalifa, Rabab, et al.
Published: (2025)
TileLoom: Automatic Dataflow Planning for Tile-Based Languages on Spatial Dataflow Accelerators
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
by: Wang, Hansheng, et al.
Published: (2024)
by: Wang, Hansheng, et al.
Published: (2024)
Similar Items
-
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
by: Ringoot, Evelyne, et al.
Published: (2025) -
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
by: Ringoot, Evelyne, et al.
Published: (2025) -
Hierarchical Recursive Precision for Accelerating Symmetric Linear Solves on MXUs
by: Carrica, Vicki, et al.
Published: (2026) -
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
by: Zheng, Size, et al.
Published: (2025) -
Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM
by: Carrica, Vicki, et al.
Published: (2025)