Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
Fuente:
arXiv
Saved in:
| Main Author: | Metere, Alfredo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prompt-Aware Scheduling for Low-Latency LLM Serving
by: Tao, Yiheng, et al.
Published: (2025)
by: Tao, Yiheng, et al.
Published: (2025)
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
by: Jung, Victor J. B., et al.
Published: (2024)
by: Jung, Victor J. B., et al.
Published: (2024)
Communication-Efficient Personalized Federal Graph Learning via Low-Rank Decomposition
by: Liu, Ruyue, et al.
Published: (2024)
by: Liu, Ruyue, et al.
Published: (2024)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
by: Yang, Hanmei, et al.
Published: (2024)
by: Yang, Hanmei, et al.
Published: (2024)
FastPersist: Accelerating Model Checkpointing in Deep Learning
by: Wang, Guanhua, et al.
Published: (2024)
by: Wang, Guanhua, et al.
Published: (2024)
Preventing Rank Collapse in Federated Low-Rank Adaptation with Client Heterogeneity
by: Wu, Fei, et al.
Published: (2026)
by: Wu, Fei, et al.
Published: (2026)
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
by: Yuan, Ziqi, et al.
Published: (2025)
by: Yuan, Ziqi, et al.
Published: (2025)
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
by: Hu, Huanqi, et al.
Published: (2025)
by: Hu, Huanqi, et al.
Published: (2025)
Towards Robust and Efficient Federated Low-Rank Adaptation with Heterogeneous Clients
by: Koo, Jabin, et al.
Published: (2024)
by: Koo, Jabin, et al.
Published: (2024)
AB-Training: A Communication-Efficient Approach for Distributed Low-Rank Learning
by: Coquelin, Daniel, et al.
Published: (2024)
by: Coquelin, Daniel, et al.
Published: (2024)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
by: Yu, Shan, et al.
Published: (2025)
by: Yu, Shan, et al.
Published: (2025)
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
by: Daghero, Francesco, et al.
Published: (2025)
by: Daghero, Francesco, et al.
Published: (2025)
MQ-GNN: A Multi-Queue Pipelined Architecture for Scalable and Efficient GNN Training
by: Ullah, Irfan, et al.
Published: (2026)
by: Ullah, Irfan, et al.
Published: (2026)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
by: Miao, Xupeng, et al.
Published: (2023)
by: Miao, Xupeng, et al.
Published: (2023)
Reliable Microservice Tail Latency Prediction via Decoupled Dual-Stream Learning and Gradient Modulation
by: Qian, Wenzhuo, et al.
Published: (2025)
by: Qian, Wenzhuo, et al.
Published: (2025)
Efficient Chromosome Parallelization for Precision Medicine Genomic Workflows
by: Montserrat, Daniel Mas, et al.
Published: (2025)
by: Montserrat, Daniel Mas, et al.
Published: (2025)
Multi-Dimensional Autoscaling of Stream Processing Services on Edge Devices
by: Sedlak, Boris, et al.
Published: (2025)
by: Sedlak, Boris, et al.
Published: (2025)
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques
by: Tomczak, Nathaniel, et al.
Published: (2025)
by: Tomczak, Nathaniel, et al.
Published: (2025)
Sometimes Painful but Certainly Promising: Feasibility and Trade-offs of Language Model Inference at the Edge
by: Abstreiter, Maximilian, et al.
Published: (2025)
by: Abstreiter, Maximilian, et al.
Published: (2025)
GPU Kernel Optimization Beyond Full Builds: An LLM Framework with Minimal Executable Programs
by: Chu, Ruifan, et al.
Published: (2025)
by: Chu, Ruifan, et al.
Published: (2025)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
by: Li, Xiangchen, et al.
Published: (2025)
by: Li, Xiangchen, et al.
Published: (2025)
A Comparative Study of OpenMP Scheduling Algorithm Selection Strategies
by: Korndörfer, Jonas H. Müller, et al.
Published: (2025)
by: Korndörfer, Jonas H. Müller, et al.
Published: (2025)
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
by: Huang, Zixiao, et al.
Published: (2025)
by: Huang, Zixiao, et al.
Published: (2025)
Compiler-First State Space Duality and Portable $O(1)$ Autoregressive Caching for Inference
by: Santoni, Cosmo
Published: (2026)
by: Santoni, Cosmo
Published: (2026)
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
by: Nichols, Daniel, et al.
Published: (2026)
by: Nichols, Daniel, et al.
Published: (2026)
Syno: Structured Synthesis for Neural Operators
by: Zhuo, Yongqi, et al.
Published: (2024)
by: Zhuo, Yongqi, et al.
Published: (2024)
TrainMover: An Interruption-Resilient Runtime for ML Training
by: Lao, ChonLam, et al.
Published: (2024)
by: Lao, ChonLam, et al.
Published: (2024)
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service
by: Imani, HamidReza, et al.
Published: (2024)
by: Imani, HamidReza, et al.
Published: (2024)
Training Time Prediction for Mixed Precision-based Distributed Training
by: Kang, Minchul, et al.
Published: (2026)
by: Kang, Minchul, et al.
Published: (2026)
Democratizing AI: A Comparative Study in Deep Learning Efficiency and Future Trends in Computational Processing
by: Amin, Lisan Al, et al.
Published: (2026)
by: Amin, Lisan Al, et al.
Published: (2026)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
by: Singh, Siddharth, et al.
Published: (2023)
by: Singh, Siddharth, et al.
Published: (2023)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Performance and Power: Systematic Evaluation of AI Workloads on Accelerators with CARAML
by: John, Chelsea Maria, et al.
Published: (2024)
by: John, Chelsea Maria, et al.
Published: (2024)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
by: Yang, Shang, et al.
Published: (2025)
by: Yang, Shang, et al.
Published: (2025)
SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module Accelerators
by: Odema, Mohanad, et al.
Published: (2024)
by: Odema, Mohanad, et al.
Published: (2024)
Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
by: Zhang, Qizheng, et al.
Published: (2025)
by: Zhang, Qizheng, et al.
Published: (2025)
Distributed Matrix-Based Sampling for Graph Neural Network Training
by: Tripathy, Alok, et al.
Published: (2023)
by: Tripathy, Alok, et al.
Published: (2023)
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
by: Sarkar, Aishwarya, et al.
Published: (2024)
by: Sarkar, Aishwarya, et al.
Published: (2024)
Similar Items
-
Prompt-Aware Scheduling for Low-Latency LLM Serving
by: Tao, Yiheng, et al.
Published: (2025) -
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
by: Jung, Victor J. B., et al.
Published: (2024) -
Communication-Efficient Personalized Federal Graph Learning via Low-Rank Decomposition
by: Liu, Ruyue, et al.
Published: (2024) -
ProTrain: Efficient LLM Training via Memory-Aware Techniques
by: Yang, Hanmei, et al.
Published: (2024) -
FastPersist: Accelerating Model Checkpointing in Deep Learning
by: Wang, Guanhua, et al.
Published: (2024)