FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Wenhao, Shao, Zishan, Cui, Kangning, Kim, Jinhee, Wang, Yixiao, Ye, Hancheng, Zhuo, Danyang, Chen, Yiran |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025)
by: Shao, Zishan, et al.
Published: (2025)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
by: Gupta, Ahan, et al.
Published: (2023)
by: Gupta, Ahan, et al.
Published: (2023)
Efficient GPU implementation of randomized SVD and its applications
by: Struski, Łukasz, et al.
Published: (2021)
by: Struski, Łukasz, et al.
Published: (2021)
Automotive Middleware Performance: Comparison of FastDDS, Zenoh and vSomeIP
by: Klüner, David Philipp, et al.
Published: (2025)
by: Klüner, David Philipp, et al.
Published: (2025)
A Zoned Storage Optimized Flash Cache on ZNS SSDs
by: Yang, Chongzhuo, et al.
Published: (2024)
by: Yang, Chongzhuo, et al.
Published: (2024)
PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters
by: Thapa, Krishu K, et al.
Published: (2025)
by: Thapa, Krishu K, et al.
Published: (2025)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
by: Qiao, Liang, et al.
Published: (2025)
by: Qiao, Liang, et al.
Published: (2025)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
by: Shin, Jiho, et al.
Published: (2024)
by: Shin, Jiho, et al.
Published: (2024)
Scalable Binary CUR Low-Rank Approximation Algorithm
by: Su, Bowen
Published: (2025)
by: Su, Bowen
Published: (2025)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
by: Zheng, Wenhao, et al.
Published: (2025)
by: Zheng, Wenhao, et al.
Published: (2025)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
by: An, Zihao, et al.
Published: (2025)
by: An, Zihao, et al.
Published: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
by: Fu, Zizhuo, et al.
Published: (2025)
by: Fu, Zizhuo, et al.
Published: (2025)
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
by: Holmes, Connor, et al.
Published: (2024)
by: Holmes, Connor, et al.
Published: (2024)
Fairness in Serving Large Language Models
by: Sheng, Ying, et al.
Published: (2023)
by: Sheng, Ying, et al.
Published: (2023)
Fast Entropy Decoding for Sparse MVM on GPUs
by: Schätzle, Emil, et al.
Published: (2026)
by: Schätzle, Emil, et al.
Published: (2026)
Dissecting Embedding Bag Performance in DLRM Inference
by: Ambati, Chandrish, et al.
Published: (2025)
by: Ambati, Chandrish, et al.
Published: (2025)
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
by: Ray, Kaustabha, et al.
Published: (2025)
by: Ray, Kaustabha, et al.
Published: (2025)
ZERNIPAX: A Fast and Accurate Zernike Polynomial Calculator in Python
by: Elmacioglu, Yigit Gunsur, et al.
Published: (2024)
by: Elmacioglu, Yigit Gunsur, et al.
Published: (2024)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
by: Metere, Alfredo
Published: (2025)
by: Metere, Alfredo
Published: (2025)
Analysis of Stable Vertex Values: Fast Query Evaluation Over An Evolving Graph
by: Afarin, Mahbod, et al.
Published: (2025)
by: Afarin, Mahbod, et al.
Published: (2025)
CAPSim: A Fast CPU Performance Simulator Using Attention-based Predictor
by: Xu, Buqing, et al.
Published: (2025)
by: Xu, Buqing, et al.
Published: (2025)
Tuning Fast Memory Size based on Modeling of Page Migration for Tiered Memory
by: Chen, Shangye, et al.
Published: (2024)
by: Chen, Shangye, et al.
Published: (2024)
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
by: Ren, Jie, et al.
Published: (2025)
by: Ren, Jie, et al.
Published: (2025)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
by: Ding, Jiabiao, et al.
Published: (2026)
by: Ding, Jiabiao, et al.
Published: (2026)
Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
by: Salaria, Shweta, et al.
Published: (2025)
by: Salaria, Shweta, et al.
Published: (2025)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
by: Kermani, Arshia, et al.
Published: (2025)
by: Kermani, Arshia, et al.
Published: (2025)
Explainable Port Mapping Inference with Sparse Performance Counters for AMD's Zen Architectures
by: Ritter, Fabian, et al.
Published: (2024)
by: Ritter, Fabian, et al.
Published: (2024)
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
by: Benazir, Afsara, et al.
Published: (2025)
by: Benazir, Afsara, et al.
Published: (2025)
Ranking with Ties based on Noisy Performance Data
by: Sankaran, Aravind, et al.
Published: (2024)
by: Sankaran, Aravind, et al.
Published: (2024)
Efficiently Ranking Software Variants with Minimal Benchmarks
by: Matricon, Théo, et al.
Published: (2025)
by: Matricon, Théo, et al.
Published: (2025)
Updates on the Low-Level Abstraction of Memory Access
by: Gruber, Bernhard Manfred
Published: (2023)
by: Gruber, Bernhard Manfred
Published: (2023)
Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
by: Wang, Tuowei, et al.
Published: (2024)
by: Wang, Tuowei, et al.
Published: (2024)
Fast solvers for Tokamak fluid models with PETSC
by: Adams, Mark F., et al.
Published: (2025)
by: Adams, Mark F., et al.
Published: (2025)
L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
by: Singh, Raul, et al.
Published: (2025)
by: Singh, Raul, et al.
Published: (2025)
A Study on Inference Latency for Vision Transformers on Mobile Devices
by: Li, Zhuojin, et al.
Published: (2025)
by: Li, Zhuojin, et al.
Published: (2025)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
Iterative Layer Pruning for Efficient Translation Inference
by: Moslem, Yasmin, et al.
Published: (2025)
by: Moslem, Yasmin, et al.
Published: (2025)
Similar Items
-
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
by: Shao, Zishan, et al.
Published: (2025) -
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
by: Gupta, Ahan, et al.
Published: (2023) -
Efficient GPU implementation of randomized SVD and its applications
by: Struski, Łukasz, et al.
Published: (2021) -
Automotive Middleware Performance: Comparison of FastDDS, Zenoh and vSomeIP
by: Klüner, David Philipp, et al.
Published: (2025) -
A Zoned Storage Optimized Flash Cache on ZNS SSDs
by: Yang, Chongzhuo, et al.
Published: (2024)