LP-GEMM: Integrating Layout Propagation into GEMM Operations
Fuente:
arXiv
Saved in:
| Main Authors: | Carneiro, César Guedes, Alvarenga, Lucas, Araujo, Guido, Rigo, Sandro |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
by: Hu, Huanqi, et al.
Published: (2025)
by: Hu, Huanqi, et al.
Published: (2025)
Accelerating Sparse DNNs Based on Tiled GEMM
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
by: Wu, Shixun, et al.
Published: (2025)
by: Wu, Shixun, et al.
Published: (2025)
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
by: Swann, Ryan, et al.
Published: (2025)
by: Swann, Ryan, et al.
Published: (2025)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
by: McFarland, Thomas, et al.
Published: (2025)
by: McFarland, Thomas, et al.
Published: (2025)
DynLP: Parallel Dynamic Batch Update for Label Propagation in Semi-Supervised Learning
by: Shovan, S M, et al.
Published: (2026)
by: Shovan, S M, et al.
Published: (2026)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
by: Metere, Alfredo
Published: (2025)
by: Metere, Alfredo
Published: (2025)
CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference
by: Zou, Yulin, et al.
Published: (2026)
by: Zou, Yulin, et al.
Published: (2026)
DeepFedNAS: Efficient Hardware-Aware Architecture Adaptation for Heterogeneous IoT Federations via Pareto-Guided Supernet Training
by: Khan, Bostan, et al.
Published: (2026)
by: Khan, Bostan, et al.
Published: (2026)
Ask the Expert: Collaborative Inference for Vision Transformers with Near-Edge Accelerators
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
Constraint-Aware Execution Planning for Hybrid Space-Ground Compute Workloads
by: Mitra, Subhadip
Published: (2026)
by: Mitra, Subhadip
Published: (2026)
Consistent Point Matching
by: Yerebakan, Halid Ziya, et al.
Published: (2025)
by: Yerebakan, Halid Ziya, et al.
Published: (2025)
InvarDiff: Cross-Scale Invariance Caching for Accelerated Diffusion Models
by: Wu, Zihao
Published: (2025)
by: Wu, Zihao
Published: (2025)
Accelerating Distributed ML Training via Selective Synchronization
by: Tyagi, Sahil, et al.
Published: (2023)
by: Tyagi, Sahil, et al.
Published: (2023)
Synthetic data shuffling accelerates the convergence of federated learning under data heterogeneity
by: Li, Bo, et al.
Published: (2023)
by: Li, Bo, et al.
Published: (2023)
Rapid Distributed Fine-tuning of a Segmentation Model Onboard Satellites
by: Plumridge, Meghan, et al.
Published: (2024)
by: Plumridge, Meghan, et al.
Published: (2024)
I-SplitEE: Image classification in Split Computing DNNs with Early Exits
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
Rate-My-LoRA: Efficient and Adaptive Federated Model Tuning for Cardiac MRI Segmentation
by: He, Xiaoxiao, et al.
Published: (2025)
by: He, Xiaoxiao, et al.
Published: (2025)
Progressive Neural Compression for Adaptive Image Offloading under Timing Constraints
by: Wang, Ruiqi, et al.
Published: (2023)
by: Wang, Ruiqi, et al.
Published: (2023)
Cross-Silo Federated Learning Across Divergent Domains with Iterative Parameter Alignment
by: Gorbett, Matt, et al.
Published: (2023)
by: Gorbett, Matt, et al.
Published: (2023)
Federated Low-Rank Tensor Estimation for Multimodal Image Reconstruction
by: Van Nguyen, Anh, et al.
Published: (2025)
by: Van Nguyen, Anh, et al.
Published: (2025)
Federated Learning for Diffusion Models
by: Peng, Zihao, et al.
Published: (2025)
by: Peng, Zihao, et al.
Published: (2025)
Decentralized Diffusion Models
by: McAllister, David, et al.
Published: (2025)
by: McAllister, David, et al.
Published: (2025)
Enhancing Split Computing and Early Exit Applications through Predefined Sparsity
by: Capogrosso, Luigi, et al.
Published: (2024)
by: Capogrosso, Luigi, et al.
Published: (2024)
FedPylot: Navigating Federated Learning for Real-Time Object Detection in Internet of Vehicles
by: Quéméneur, Cyprien, et al.
Published: (2024)
by: Quéméneur, Cyprien, et al.
Published: (2024)
EdgeOL: Efficient in-situ Online Learning on Edge Devices
by: Li, Sheng, et al.
Published: (2024)
by: Li, Sheng, et al.
Published: (2024)
Data-Free Federated Class Incremental Learning with Diffusion-Based Generative Memory
by: Wang, Naibo, et al.
Published: (2024)
by: Wang, Naibo, et al.
Published: (2024)
PARDON: Privacy-Aware and Robust Federated Domain Generalization
by: Nguyen, Dung Thuy, et al.
Published: (2024)
by: Nguyen, Dung Thuy, et al.
Published: (2024)
MTL-Split: Multi-Task Learning for Edge Devices using Split Computing
by: Capogrosso, Luigi, et al.
Published: (2024)
by: Capogrosso, Luigi, et al.
Published: (2024)
Self-supervised Cross-silo Federated Neural Architecture Search
by: Liang, Xinle, et al.
Published: (2021)
by: Liang, Xinle, et al.
Published: (2021)
One-Shot Sequential Federated Learning for Non-IID Data by Enhancing Local Model Diversity
by: Wang, Naibo, et al.
Published: (2024)
by: Wang, Naibo, et al.
Published: (2024)
FedDM: Enhancing Communication Efficiency and Handling Data Heterogeneity in Federated Diffusion Models
by: Vora, Jayneel, et al.
Published: (2024)
by: Vora, Jayneel, et al.
Published: (2024)
Federated Self-supervised Domain Generalization for Label-efficient Polyp Segmentation
by: Tan, Xinyi, et al.
Published: (2025)
by: Tan, Xinyi, et al.
Published: (2025)
FedRepOpt: Gradient Re-parametrized Optimizers in Federated Learning
by: Lau, Kin Wai, et al.
Published: (2024)
by: Lau, Kin Wai, et al.
Published: (2024)
Entity Augmentation for Efficient Classification of Vertically Partitioned Data with Limited Overlap
by: Amalanshu, Avi, et al.
Published: (2024)
by: Amalanshu, Avi, et al.
Published: (2024)
Recurrent Early Exits for Federated Learning with Heterogeneous Clients
by: Lee, Royson, et al.
Published: (2024)
by: Lee, Royson, et al.
Published: (2024)
A Converting Autoencoder Toward Low-latency and Energy-efficient DNN Inference at the Edge
by: Mahmud, Hasanul, et al.
Published: (2024)
by: Mahmud, Hasanul, et al.
Published: (2024)
A Parallel Workflow for Polar Sea-Ice Classification using Auto-labeling of Sentinel-2 Imagery
by: Iqrah, Jurdana Masuma, et al.
Published: (2024)
by: Iqrah, Jurdana Masuma, et al.
Published: (2024)
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge
by: Grailoo, M., et al.
Published: (2026)
by: Grailoo, M., et al.
Published: (2026)
Similar Items
-
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
by: Hu, Huanqi, et al.
Published: (2025) -
Accelerating Sparse DNNs Based on Tiled GEMM
by: Guo, Cong, et al.
Published: (2024) -
TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
by: Wu, Shixun, et al.
Published: (2025) -
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025) -
tritonBLAS: Triton-based Analytical Approach for GEMM Kernel Parameter Selection
by: Swann, Ryan, et al.
Published: (2025)